New model · Longer tasks and measured performance
Gemini 4 Argon is here. Can longer reasoning finish more of the job?
Google raises the output limit to one million tokens, with gains in independent automation tests. Compare familiar models, then check who has access and how introductory pricing works.
Google published its announcement at 20:00 UTC on September 30—04:00 Beijing time on October 1. Checked on October 1: access starts with selected testers, rather than a broad release in consumer Gemini.
The useful question is how far the work gets
Google emphasizes sustained, multi-step coding, knowledge work and cyber defense. Its libgav1 example involves repeated profiling and compiler inspection while revising Rust code. Google reports 2.7 times the speed of the earlier Rust port with identical video output. That compares two versions of this project, not every task.
Google’s output ceiling rises from 64K to 1M tokens, allowing longer reasoning and generation. Artificial Analysis tested Long Decode Continuation, which pauses a long response and resumes through follow-up calls. This differs from input capacity and does not by itself establish the quality of a very long artifact.

Sources and further reading
- Google: introducing Gemini 4 ArgonOfficial announcement · September 30 UTCOutput limit, Google's internal examples, phased access and post-promotion prices. Internal results are Google's reports.
- Artificial Analysis: independent Argon evaluationEvaluation provider · Published September 30, checked October 1Overall position, task differences, Long Decode Continuation and benchmark task costs; not rerun by BitShovel.
Automation improves; the task still changes the comparison
In the same independent suite, Argon scores 53 versus Gemini 3.8 Flash's 41, matching Astra. Its task breakdown shows stronger application automation while terminal coding still has models ahead. The table identifies reasoning settings; these are not necessarily a user's defaults.
These results inform a choice without completing your report, website or revisions. Google's 51.3% AutomationBench result and the 78% AutomationBench-AA result below use different evaluations; combining them would not show a before-and-after gain.
| What is measured | Argon (high) | Reference |
|---|---|---|
| Intelligence Index | 53 | Gemini 3.8 Flash (high): 41; Astra (max): 53; GPT-6.1 Sol (max): 52. |
| Application automation | AutomationBench-AA: 78% | Gemini 3.8 Flash (high): 60%; Sonnet 5.5 (max with fallback): 71%. |
| Terminal coding and operations | Terminal-Bench 4: 57% | Astra (max): 59%; Opus 5.5 and Sonnet 5.5 (max with fallback): 60% and 64%. |
Sources and further reading
- Google: introducing Gemini 4 ArgonOfficial announcement · September 30 UTCOutput limit, Google's internal examples, phased access and post-promotion prices. Internal results are Google's reports.
- Artificial Analysis: independent Argon evaluationEvaluation provider · Published September 30, checked October 1Overall position, task differences, Long Decode Continuation and benchmark task costs; not rerun by BitShovel.
- Artificial Analysis: comparison with Gemini 3.8 FlashSame evaluation suite · Checked October 1Both models use high reasoning, providing a reference for existing Gemini users.
Access and what the price actually means
At this check, access starts with invited cyber defenders in Fairwind and other trusted testers. Google plans to expand first to paid API customers and Google AI Ultra subscribers, without a specific date. A missing consumer-Gemini entry need not mean you overlooked a setting.
Announced introductory API prices are $2 per million input tokens and $10 per million output tokens, rising to $4/$20 afterward, with a 95% discount on cached input. Google has not published an end date. These are not Ultra subscription prices or a full task budget; reasoning, caching, tools and revisions affect costs.
Sources and further reading
- Google: introducing Gemini 4 ArgonOfficial announcement · September 30 UTCOutput limit, Google's internal examples, phased access and post-promotion prices. Internal results are Google's reports.
- Artificial Analysis: independent Argon evaluationEvaluation provider · Published September 30, checked October 1Overall position, task differences, Long Decode Continuation and benchmark task costs; not rerun by BitShovel.
Bring the change back to the work you already do
If you use Gemini for research or writing, Argon offers a new capability reference. Once your product supports it, compare a task that previously needed you to take over: preserved requirements, editable files and remaining human decisions. This is our suggested comparison, not a completed Argon user test.
Astra, Sol and Opus users can revisit the same task while keeping a familiar workflow. Our skills article explains preserving writing instructions and material. That workflow change does not establish that skills are already powered by Argon. Check the model, product integration and finished result separately, then connect them.
- Next: Astra artifacts and how they were revised
- Compare models and costs for the same task
- Gems and skills: keep your instructions and material
Sources and further reading
- Google: introducing Gemini 4 ArgonOfficial announcement · September 30 UTCOutput limit, Google's internal examples, phased access and post-promotion prices. Internal results are Google's reports.
- Artificial Analysis: independent Argon evaluationEvaluation provider · Published September 30, checked October 1Overall position, task differences, Long Decode Continuation and benchmark task costs; not rerun by BitShovel.