Ongoing topic · Compiled September 25, 2026
Opus 5.5: clearer answers, usable results
Explore clearer communication, independent evaluations and shared-task outputs, then connect longer tasks, access and costs with the tools you already use.
Following next: Which public Opus 5.5 projects show revisions from an idea to an artifact?

01The release
Start with an image you can open
Opus 5.5 launched on September 22. This topic connects the release, independent tests, actual outputs and access routes, providing reference points for both existing Claude users and people using other models.
The opening image is Simon Willison’s SVG generated with Opus 5.5 at medium effort: a pelican on a bicycle. Text and graphics code become something a browser can display. This is the author’s original output; the following sections put that small visible result beside evidence about more demanding work.
Sources and notes
- Anthropic release ↗Official release · September 22Release and communication changes.
- Simon’s Claude comparison grid ↗Author’s outputs · September 22Outputs, effort settings and token usage.
02Compare with the previous version
Why clearer explanations matter
Anthropic emphasizes clearer communication: leading with the important point, using less jargon and following writing instructions. Its release page includes paired answers readers can inspect. These are vendor-selected examples.
Our editorial question is what happens next: can you tell what changed, why, what remains unfinished and where your judgment is needed? Whether organizing event information, revising a webpage or writing an explanation, understanding the work makes useful feedback easier.
- Compare the official answersSee the Communication section of the release.
03Capability and task evidence
A leading score, then the work behind it
Read September 25, the Artificial Analysis model card lists Opus 5.5 max with default fallback at 58, first among its 211 comparable models. That is a reason to examine it closely, within the stated evaluation, effort setting and runtime conditions.
GitHub reports comparable task resolution to Opus 5 with fewer steps and tokens in early testing, plus recovery from errors in multistep work. Copilot integration brings those model changes into tools developers already use.
For long tasks, look for final files, changes and checks. Runtime describes the process; outputs that open and remain editable reveal what was accomplished. Future additions will prioritize projects with public processes and deliverables.
Sources and notes
- Artificial Analysis model card ↗Independent evaluation · read September 25Score 58, max effort and default fallback; rank within the card’s comparison class.
- GitHub Copilot integration ↗Product-team tests and integration · September 22Steps, token use and recovery in longer tasks.
04Outputs and effort settings
One pelican: outputs, usage and unfinished attempts

The opening Opus 5.5 medium image and this Opus 5 medium image come from the author’s shared-task grid. Two distinct legs are visible in the newer output; one dominates the older image. Their styling also differs. This describes these particular images, not every drawing either model will produce.
The author also reports two max-effort attempts that exhausted the output allowance without producing an SVG. More reasoning does not guarantee completion of this task. A practical starting point is medium effort, raising it for a specific problem while retaining outputs for comparison.
- Open the complete comparison gridOriginal SVGs, usage and failed attempts across models and effort levels.
- Read the author’s accountAn independent author’s experience, not a BitShovel test.
Sources and notes
- Simon Willison’s original grid ↗Original results · September 22Both displayed images use medium effort.
- Simon Willison’s release-day experience ↗Author’s experience · September 22Two max-effort attempts without SVG output.
05Cost and access
Lower rates, then the cost of finishing

The official cost guide lists standard API prices per million tokens: $4 for fresh input, $20 for output and $0.20 for cache reads. Input and output rates fall 20% from Opus 5, cache reads 60%. Subscription allowances and API bills are different measures.
A task’s bill also depends on turns, cache hits, reasoning output and retries. Compare the same material and completion requirements, then the result and actual bill; a rate reduction alone does not establish the saving on every task.
Existing Copilot users can check the model picker on supported plans and clients. GitHub describes a gradual rollout, with organization access affected by administrator settings. Claude users should check their account’s model and usage controls first.
- Read the cost guide and calculatorDistinguish illustrative calculations from your actual usage.
- Check Copilot availabilityPlans, clients and organization settings.
Sources and notes
- Claude: what a task costs ↗Official cost guide · September 22Rates, caching, reasoning and retries.
- GitHub: access routes ↗Official integration · September 22Gradual rollout and supported plans.
06Bring it back to your work
Stay with Claude or use it alongside other models?
If you use Opus 5, compare a familiar task: does the same material require fewer clarifications or revisions, and is the final explanation easier to assess? If you prefer GPT, Grok or GLM, try one task your current tool struggles with and compare complete outputs. You can add a useful option while keeping your established workflow.
The Astra topic continues into 3D scenes and real webpages; our cost article separates Sol, Luna and Opus pricing measures. Together they offer different creative and practical routes, beyond choosing a subscription from one drawing. This topic will follow Opus 5.5 projects, Chinese-language experiences and shared-task comparisons.
- Explore real Astra creationsFrom a description to an editable 3D scene.
- Compare Sol, Luna and Opus costsSeparate rates, task bills and subscriptions.
What we are following
07What comes next
What we are following
- Which public Opus 5.5 projects show revisions from an idea to an artifact?
- Do communication gains hold in Chinese writing and long conversations?
- How do results, revisions and costs compare with Astra, Grok or GLM on the same task?