Model changes · From capability to everyday use

Sonnet 5.5 is here. When should an Opus user choose it?

Near-flagship benchmark scores do not mean every task costs less. Connect the official positioning, independent tests and Copilot availability to everyday iteration and complex work.

Anthropic released Sonnet 5.5 on September 28. Reviewed on September 29 Beijing time, this article asks how the release might change the way we divide work, beyond its benchmark position.

In this article4 chapters

Its focus: moving well-defined work along faster

Anthropic positions Sonnet 5.5 alongside Opus 5.5 for everyday tasks, bug fixes, documents and spreadsheets, reserving complex, open-ended judgment for Opus. It reports over 30% faster output than Sonnet 5 and up to 30% lower task costs for most work. These are vendor findings.

The release page pairs HTML demos of flocking birds, wind-shaped dunes and a clock made from 24 smaller clocks. They offer a visible entry to what changed. Compare the same task and its progress, while keeping these selected demos separate from broad reliability claims.

Sources and further reading

Near-Opus scores do not make maximum effort the best value

At maximum effort, Artificial Analysis measured an Intelligence Index score of 56, behind Opus 5.5 at 58. Its average benchmark task cost was $7.60, about 50% above Sonnet 5. Near-flagship results involved more output tokens in this suite. That average is not a quoted price for an everyday question.

These are different conditions from the vendor’s savings claim. Listed API input/output rates are $2/$10 per million tokens for Sonnet and $4/$20 for Opus 5.5. Claude apps and Claude Code default to Medium effort; the platform defaults to High. Unit rates, effort and actual consumption are separate; half the rate does not guarantee half the bill.

The independent evaluation used a pre-release deployment with a structured-output bug, since fixed for release. The evaluator plans relevant reruns. Retain the test date and conditions, and follow the rerun rather than treating the ranking as permanent.

Sources and further reading

Already using Copilot? Start with a task you know

GitHub announced Sonnet 5.5 for Copilot Pro, Pro+, Max, Business and Enterprise, rolling out gradually across entry points including VS Code and CLI. Organization policies also affect access. GitHub reports fewer steps, tokens and tool calls in early coding tests. This is an integration provider’s evaluation, not an independent creator’s project.

Our suggested comparison is a small change you understand and can verify, such as fixing a layout error or adding form validation. Record completion time, rework and actual charges. Keep scope, prompting and effort stable enough to understand the difference.

GitHub announcement artwork showing Claude Sonnet 5.5 in the Copilot model picker.
GitHub’s official integration artwork. The picker is a practical entry point; availability still depends on rollout and administrator settings. · Open full-size image
Sources and further reading

Assign the work before deciding to switch

If you use Sonnet 5, repeat a known task to check waiting and rework. If you use Opus 5.5, similar scores alone do not justify replacing it everywhere. Try Sonnet for bounded edits, retain your existing model for difficult judgments, and adjust from your results. This is a suggested approach, not a verified universal optimum.

Our Opus topic connects video projects with production experience. Whether Sonnet can take on their subtitles, layouts or code changes needs same-project evidence. Opus examples do not prove Sonnet’s abilities. Next we will follow independent finished projects, revision histories and Artificial Analysis reruns.

Sources and further reading