New model · Longer tasks and measured performance

Gemini 4 Argon is here. Can longer reasoning finish more of the job?

Google raises the output limit to one million tokens, with gains in independent automation tests. Compare familiar models, then check who has access and how introductory pricing works.

Google published its announcement at 20:00 UTC on September 30—04:00 Beijing time on October 1. Checked on October 1: access starts with selected testers, rather than a broad release in consumer Gemini.

In this article4 chapters

The useful question is how far the work gets

Google emphasizes sustained, multi-step coding, knowledge work and cyber defense. Its libgav1 example involves repeated profiling and compiler inspection while revising Rust code. Google reports 2.7 times the speed of the earlier Rust port with identical video output. That compares two versions of this project, not every task.

Google’s output ceiling rises from 64K to 1M tokens, allowing longer reasoning and generation. Artificial Analysis tested Long Decode Continuation, which pauses a long response and resumes through follow-up calls. This differs from input capacity and does not by itself establish the quality of a very long artifact.

Google's official Gemini 4 Argon announcement artwork, showing the model name and a large 4 on blue.
Official Google artwork identifying this release, rather than a finished artifact or a BitShovel test screenshot. · Open full-size image
Sources and further reading

Automation improves; the task still changes the comparison

In the same independent suite, Argon scores 53 versus Gemini 3.8 Flash's 41, matching Astra. Its task breakdown shows stronger application automation while terminal coding still has models ahead. The table identifies reasoning settings; these are not necessarily a user's defaults.

These results inform a choice without completing your report, website or revisions. Google's 51.3% AutomationBench result and the 78% AutomationBench-AA result below use different evaluations; combining them would not show a before-and-after gain.

Artificial Analysis · Checked October 1 · Index scores and task success rates are separate
What is measuredArgon (high)Reference
Intelligence Index53Gemini 3.8 Flash (high): 41; Astra (max): 53; GPT-6.1 Sol (max): 52.
Application automationAutomationBench-AA: 78%Gemini 3.8 Flash (high): 60%; Sonnet 5.5 (max with fallback): 71%.
Terminal coding and operationsTerminal-Bench 4: 57%Astra (max): 59%; Opus 5.5 and Sonnet 5.5 (max with fallback): 60% and 64%.
Sources and further reading

Access and what the price actually means

At this check, access starts with invited cyber defenders in Fairwind and other trusted testers. Google plans to expand first to paid API customers and Google AI Ultra subscribers, without a specific date. A missing consumer-Gemini entry need not mean you overlooked a setting.

Announced introductory API prices are $2 per million input tokens and $10 per million output tokens, rising to $4/$20 afterward, with a 95% discount on cached input. Google has not published an end date. These are not Ultra subscription prices or a full task budget; reasoning, caching, tools and revisions affect costs.

Sources and further reading

Bring the change back to the work you already do

If you use Gemini for research or writing, Argon offers a new capability reference. Once your product supports it, compare a task that previously needed you to take over: preserved requirements, editable files and remaining human decisions. This is our suggested comparison, not a completed Argon user test.

Astra, Sol and Opus users can revisit the same task while keeping a familiar workflow. Our skills article explains preserving writing instructions and material. That workflow change does not establish that skills are already powered by Argon. Check the model, product integration and finished result separately, then connect them.

Sources and further reading