Model rankings Artificial Analysis intelligence comparison Same coding tasks: Pi / high
Intelligence Index · Each model uses its highest tested setting on the source page, with fallback conditions retained; scores come from one evaluator.
13 model configurations · Filters retain the original rank
Scroll sideways for scores and evaluations; model names stay on the left.
How to read this leaderboard Included models: 13 models, each at its highest tested reasoning setting on the source page; unnamed settings retain the source label. Sorted by that source's displayed score, not a global top 13. Equal integers do not establish equal unrounded results. Effort and fallback settings differ, so this is not an equal-compute comparison.
Rows follow source scores; equal displayed scores share a rank and filtering does not renumber them. Different boards measure different things and are not combined into a total. Works, speed, cost and access still need separate examination.
Read the methodology ↗ Astra addresses complex work while Sol serves everyday tasks; compare generations through matched tasks.
Follow 3 releases 2026-09-03 GPT-6 Astra GPT-6's complex-work model connects understanding, tool use and file delivery.
Release source ↗ 2026-09-22 GPT-6 Sol / Luna Sol and Luna join the GPT-6 family with different capability and cost trade-offs.
Release source ↗ 2026-09-29 GPT-6 Sol → GPT-6.1 Sol Sol advances to 6.1, with reported gains in coding, computer use and professional work.
Release source ↗ Claude's flagship and everyday models evolve together; examine creation, revisions and communication alongside benchmarks.
Follow 2 releases 2026-09-22 Opus 5 → Opus 5.5 An update to complex work and communication, with public code-animation and same-prompt game examples.
Release source ↗ 2026-09-28 Sonnet 5 → Sonnet 5.5 The second 5.5-family model emphasizes everyday work, speed and effort-dependent costs.
Release source ↗ Compare longer-workflow capability from Gemini 3.8 Flash to Argon, alongside actual availability.
Follow 2 releases 2026-09-02 Gemini 3.7 Flash → 3.8 Flash Flash advances coding, multi-step reasoning and agent work at the same speed and price positioning; Flash Cyber is a separate security-focused variant.
Release source ↗ 2026-09-30 Gemini 4 Argon The new generation is announced with access initially limited to invited testers; evaluation results do not establish public availability.
Release source ↗ Muse Spark is the model; Muse connects it to a computer and services as a personal agent. Keep the two layers distinct.
Follow 2 releases 2026-08-05 Muse Spark 1.1 → 1.2 Coding, debugging and codebase understanding advance through co-training with Muse Code; the model and terminal agent launch together as distinct layers.
Release source ↗ 2026-09-02 Muse Spark 1.2 → 1.3 The release emphasizes longer tasks, tool use and coding, with max effort available in Muse Code and Meta Model API.
Release source ↗ Follow Grok's coding and knowledge-work updates, separating standard models from faster serving tiers.
Follow 2 releases 2026-08-12 Grok 4.5 → 4.6 Further training advances long tasks, coding and interactive work through Grok Build, Cursor and the API; faster serving has separate pricing.
Release source ↗ 2026-09-21 Grok 4.6 → 4.7 A larger base model emphasizes long tasks and self-checking, at the same standard token rates as 4.6.
Release source ↗ Hosted Max, Flash-Next and downloadable 27B serve different roles; parameter size does not replace measured performance.
Follow 2 releases 2026-09-02 Qwen3.8 Max → Max 0902 The 0902 snapshot updates coding, tool orchestration and vision while retaining explicit version identity.
Release source ↗ 2026-09-05 qwen3.8-max → Max 0902 (default) The announced default Max endpoint transition points to 0902; this is a routing change, not another new model.
Release source ↗ Flash's vision and tool capabilities reach existing products; track model identities together with API routing.
Follow 5 releases 2026-04-24 V4 Pro / V4 Flash The V4 family reaches the API with Pro and Flash offering different capability and resource trade-offs; older request names temporarily map to Flash.
Release source ↗ 2026-07-31 V4 Flash Preview → Flash 0731 The architecture and size stay unchanged while post-training and Responses API support advance; this update applies only to the Flash API, not Pro or the web model.
Release source ↗ 2026-08-13 V4 Pro · GA Pro reaches general availability in the app, web and API with Responses API support and more flexible thinking effort; its request name stays unchanged.
Release source ↗ 2026-08-21 V4 Flash → Flash Vision Exp A separate experimental vision model adds image understanding to Flash; V4.1 Flash later succeeds it.
Release source ↗ 2026-09-10 V4 Flash / Vision Exp → V4.1 Flash Native multimodal Flash is released; legacy Flash API names temporarily route to the new model while Pro remains available.
Release source ↗ Anthropic
Connect complex-work capability with editable code animation, independent evaluations and real production choices.
Evaluations & evolution → Anthropic
A new everyday-work option; read scores together with effort, token use and whole-task cost.
Evaluations & evolution → Anthropic
A previous-generation reference: the same handwritten-edit sample exposes differences more concretely than an overall score.
Evaluations & evolution → OpenAI
Connect materials, code and tools into longer work, with concrete comparisons in drawing, handwritten edits and furniture assembly.
Evaluations & evolution → Google DeepMind
A new reference for long workflows, initially available to invited testers; examine the tasks behind the index.
Evaluations & evolution → OpenAI
The next Sol update; an independent side-by-side comparison places its capability and cost against Astra.
Evaluations & evolution → Meta
The capability model behind Muse; separate model improvements, product integration and the personal-agent experience.
Evaluations & evolution → SpaceXAI
The release emphasizes sustained coding, knowledge work and self-checking; distinguish announced changes from demonstrated results.
Evaluations & evolution → Alibaba / Qwen
A named snapshot of hosted Max, with reported coding, orchestration and vision improvements. A dedicated on-site evaluation has not yet been published.
Evaluations & evolution → Google DeepMind
A previous-generation reference for Argon, retaining task-level comparisons within the same independent evaluation.
Evaluations & evolution → Alibaba / Qwen
Understand total versus active parameters, and distinguish downloadable weights from a hosted service.
Evaluations & evolution → DeepSeek
Native image understanding connects to files and tools; official effort curves show the gains and costs of more reasoning.
Evaluations & evolution → Alibaba / Qwen
A smaller model for text, image and video understanding; check reasoning controls, context and deployment materials.
Evaluations & evolution →