01 / The source

Models

See where capabilities stand, then follow company evolution and concrete evaluations to understand what changed.

Model rankings

Artificial Analysis intelligence comparison

Source read · · v4.3

Source leaderboard ↗

Intelligence Index · Each model uses its highest tested setting on the source page, with fallback conditions retained; scores come from one evaluator.

13 model configurations · Filters retain the original rank

Scroll sideways for scores and evaluations; model names stay on the left.

Artificial Analysis intelligence comparison · Intelligence Index
RankModel / test configurationCompanyScoreExplore
1Claude Opus 5.5max with fallbackAnthropic58Evaluations →
2Claude Sonnet 5.5max with fallbackAnthropic56Evaluations →
3Claude Fable 5.1max with fallbackAnthropic53Evaluations →
3GPT-6 AstramaxOpenAI53Evaluations →
3Gemini 4 ArgonhighGoogle DeepMind53Evaluations →
6GPT-6.1 SolmaxOpenAI52Evaluations →
7Muse Spark 1.3maxMeta48Evaluations →
8Grok 4.7xhighSpaceXAI46Evaluations →
9Qwen3.8 Max (0902)0902 snapshot; effort not named in source tableAlibaba / Qwen45Evaluations →
10Gemini 3.8 FlashhighGoogle DeepMind41Evaluations →
11Qwen3.8-Flash-NextEffort not named in source tableAlibaba / Qwen40Evaluations →
12DeepSeek V4.1 FlashmaxDeepSeek39Evaluations →
13Qwen3.8 27BxhighAlibaba / Qwen34Evaluations →
How to read this leaderboard

Included models: 13 models, each at its highest tested reasoning setting on the source page; unnamed settings retain the source label. Sorted by that source's displayed score, not a global top 13. Equal integers do not establish equal unrounded results. Effort and fallback settings differ, so this is not an equal-compute comparison.

Rows follow source scores; equal displayed scores share a rank and filtering does not renumber them. Different boards measure different things and are not combined into a total. Works, speed, cost and access still need separate examination.

Read the methodology ↗

How companies reached this generation

Read releases through what came before and after, with the original explanation for each step.

OpenAI ↗

Astra addresses complex work while Sol serves everyday tasks; compare generations through matched tasks.

Follow 3 releases
  1. GPT-6 Astra

    GPT-6's complex-work model connects understanding, tool use and file delivery.

    Release source ↗
  2. GPT-6 Sol / Luna

    Sol and Luna join the GPT-6 family with different capability and cost trade-offs.

    Release source ↗
  3. GPT-6 Sol → GPT-6.1 Sol

    Sol advances to 6.1, with reported gains in coding, computer use and professional work.

    Release source ↗

Anthropic ↗

Claude's flagship and everyday models evolve together; examine creation, revisions and communication alongside benchmarks.

Follow 2 releases
  1. Opus 5 → Opus 5.5

    An update to complex work and communication, with public code-animation and same-prompt game examples.

    Release source ↗
  2. Sonnet 5 → Sonnet 5.5

    The second 5.5-family model emphasizes everyday work, speed and effort-dependent costs.

    Release source ↗

Google DeepMind ↗

Compare longer-workflow capability from Gemini 3.8 Flash to Argon, alongside actual availability.

Follow 2 releases
  1. Gemini 3.7 Flash → 3.8 Flash

    Flash advances coding, multi-step reasoning and agent work at the same speed and price positioning; Flash Cyber is a separate security-focused variant.

    Release source ↗
  2. Gemini 4 Argon

    The new generation is announced with access initially limited to invited testers; evaluation results do not establish public availability.

    Release source ↗

Meta ↗

Muse Spark is the model; Muse connects it to a computer and services as a personal agent. Keep the two layers distinct.

Follow 2 releases
  1. Muse Spark 1.1 → 1.2

    Coding, debugging and codebase understanding advance through co-training with Muse Code; the model and terminal agent launch together as distinct layers.

    Release source ↗
  2. Muse Spark 1.2 → 1.3

    The release emphasizes longer tasks, tool use and coding, with max effort available in Muse Code and Meta Model API.

    Release source ↗

SpaceXAI ↗

Follow Grok's coding and knowledge-work updates, separating standard models from faster serving tiers.

Follow 2 releases
  1. Grok 4.5 → 4.6

    Further training advances long tasks, coding and interactive work through Grok Build, Cursor and the API; faster serving has separate pricing.

    Release source ↗
  2. Grok 4.6 → 4.7

    A larger base model emphasizes long tasks and self-checking, at the same standard token rates as 4.6.

    Release source ↗

Alibaba / Qwen ↗

Hosted Max, Flash-Next and downloadable 27B serve different roles; parameter size does not replace measured performance.

Follow 2 releases
  1. Qwen3.8 Max → Max 0902

    The 0902 snapshot updates coding, tool orchestration and vision while retaining explicit version identity.

    Release source ↗
  2. qwen3.8-max → Max 0902 (default)

    The announced default Max endpoint transition points to 0902; this is a routing change, not another new model.

    Release source ↗

DeepSeek ↗

Flash's vision and tool capabilities reach existing products; track model identities together with API routing.

Follow 5 releases
  1. V4 Pro / V4 Flash

    The V4 family reaches the API with Pro and Flash offering different capability and resource trade-offs; older request names temporarily map to Flash.

    Release source ↗
  2. V4 Flash Preview → Flash 0731

    The architecture and size stay unchanged while post-training and Responses API support advance; this update applies only to the Flash API, not Pro or the web model.

    Release source ↗
  3. V4 Pro · GA

    Pro reaches general availability in the app, web and API with Responses API support and more flexible thinking effort; its request name stays unchanged.

    Release source ↗
  4. V4 Flash → Flash Vision Exp

    A separate experimental vision model adds image understanding to Flash; V4.1 Flash later succeeds it.

    Release source ↗
  5. V4 Flash / Vision Exp → V4.1 Flash

    Native multimodal Flash is released; legacy Flash API names temporarily route to the new model while Pro remains available.

    Release source ↗

From scores to evaluations

Continue through each model’s sources, test conditions and related works.