I want review to find problems that change whether the work can be accepted. A huge list of suggestions can consume more time than implementation, especially when every comment introduces another preference. Paying less for the reviewer helps, but I still need to know what that second reading accomplishes.
On September 15, 2026, I described using Sonnet to implement and Opus to review. That division existed before Opus 5.5, announced on September 22. The recorded practice concerns model families without identifying versions. It gives me a starting point for considering the release without turning earlier use into a test of the new model.
Anthropic announced input at $4 and output at $20 per million tokens, with cache reads at $0.20. The company estimated a 40% reduction in typical cost through pricing and efficiency. I like having more room in the budget for a careful second reading. The provider's percentage still needs to meet a task where that reading justifies the effort.
My reviewer should receive the original request and the changes. The implementer's explanation helps locate decisions, but can carry the hypothesis that needs examination. If both agents receive only the same summary, changing the model name preserves an important source of bias. I want review to return to the expected behavior and look for what the diff leaves out.
I also need to distinguish defects from preferences. A problem preventing acceptance deserves correction. An alternative organization can wait. When the reviewer makes every suggestion compulsory, the original assignment never ends and the implementer gets another implementation to do. That belongs in review cost, alongside the discussion required to interpret vague comments.
On September 22, I asked for model choice to reflect the best fit for each task. Review belongs under that rule. A small correction and a change with many dependencies deserve different attention. The Sonnet and Opus division remains useful while it responds to the work received, and can change when conditions change.
I will evaluate Opus 5.5 through the usefulness of findings on bounded tasks: each problem should show when it appears and why it affects the request. Cost includes the return trip to an acceptable change. Long comments and two agreeing agents are insufficient verification. I want evidence that lets me accept the change or correct a specific defect without opening an endless discussion of taste.