CARD · Card record

Qwen3.8-Flash-Next weights offer a Qwen4 architecture preview

Qwen provides an experimental multimodal MoE preview and an FP8 variant. Its model card now specifies 125B language-model parameters with 6B active, plus 51B n-gram embeddings and 4B MTP parameters, under Qwen Community License 1.0.

01
THE STORY · The original

Read the explanation and its limits.

  • Originalofficial statementhttps://qwen.ai/blog?id=qwen3.8-flash-next

Why it is worth understanding

Developers working on deployment, quantization, or inference compatibility can inspect the architecture and measure task performance. A replacement decision also needs hardware, output-quality, and licensing checks.

Conditions and limitations

The 6B active figure does not imply a 6B model’s memory footprint. The release is experimental. Its license requires separate permission for commercial use by businesses offering model services or defined AI work assistants, with an internal-use exception. We have not tested performance or cost improvements.

Important additions and corrections

Dates below mark changes to our coverage.

  • Correction ·

    Added officially documented parameters and license conditions: 6B active parameters do not imply a 6B model’s memory needs, and some commercial uses require separate permission.

Reading progress is saved in this browser. Save or add a note →

02
THE THREAD · Full timeline

16 events, every row traceable.

A card accumulates records as its story develops: new corroborators, market contrasts, material changes at the original — each dated and linked to its archive day.

Card timeline
  1. Added to BitShovelEntered the BitShovel radar: 50 entries visible on AIHOT featured at collection.
  2. RecordVerified here: the ModelScope model page is live under the official name (title read confirmed); TechNode relays the official schedule noting the FP8 pair while specs remain undisclosed.
  3. Related discussionComparison (Qwen): "Kaitchup posted Qwen3.8 27B Benchmarks for quant" — lead via Reddit r/LocalLLaMA. Same-class comparisons and sentiment signal the market's relative standing.
  4. Related discussionRelease watch (Qwen): market chatter captured via AIHOT featured (headline in Chinese). Same-class release moves reshape expectations and alternatives.
  5. Related discussionRelease watch (Qwen): "Qwen3.8-Max-0902 released" — lead via Reddit r/LocalLLaMA. Same-class release moves reshape expectations and alternatives.
  6. Related discussionComparison (Qwen): market chatter captured via AIHOT featured (headline in Chinese). Same-class comparisons and sentiment signal the market's relative standing.
  7. Related discussionComparison (Qwen): "Opencode vs Deepseek harness: my experience with" — lead via Reddit r/LocalLLaMA. Same-class comparisons and sentiment signal the market's relative standing.
  8. Related discussionComparison (Qwen): "Alibaba's Qwen3.8-Max-0902 Tops Code Arena, Beat" — lead via alphasignal-daily. Same-class comparisons and sentiment signal the market's relative standing.
  9. Related discussionComparison (Qwen): "Perplexity's Lily Beats MLX-LM by 1.35x Running " — lead via alphasignal-daily. Same-class comparisons and sentiment signal the market's relative standing.
  10. Related discussionRelease watch (Qwen): "Perplexity open-sourced their Mac inference serv" — lead via reddit-local-llama. Same-class release moves reshape expectations and alternatives.
  11. Related discussionRelease watch (Qwen): "We open-sourced Paddock, our Rust/C++ inference " — lead via reddit-local-llama. Same-class release moves reshape expectations and alternatives.
  12. Related discussionComparison (Qwen): "Updated my benchmark with a new vLLM based recip" — lead via reddit-local-llama. Same-class comparisons and sentiment signal the market's relative standing.
  13. Related discussionComparison (Qwen): "I benchmarked 21 Qwen3.8 27B variants on 16GB VR" — lead via reddit-local-llama. Same-class comparisons and sentiment signal the market's relative standing.
  14. Related discussionComparison (Qwen): "Qwen3.8-27B on 2× RTX 5070 Ti 16GB — llama.cpp v" — lead via reddit-local-llama. Same-class comparisons and sentiment signal the market's relative standing.
  15. Related discussionComparison (Qwen): "NInfer vs llama.cpp vs vLLM: quality + speed com" — lead via reddit-local-llama. Same-class comparisons and sentiment signal the market's relative standing.
  16. Related discussionRelease watch (Qwen): market chatter captured via ithome (headline in Chinese). Same-class release moves reshape expectations and alternatives.

03
SOURCES · Sources and evidence

Check each link's purpose and limits.

Sources and evidence · 4 links

Each link has a purpose and scope. Link counts do not establish independent confirmation; direct support applies only to the statements specified below.

Found an error or a problem? Send feedback ↗

04
EDITORIAL REVIEW

Review of this card's wording and evidence.

Editorial review · Evidence reviewed

· Beijing time (UTC+8)

Checked source materials to update purpose, specifications and usage conditions. Original publication and collection dates are retained; this review does not establish a new product event.

This dates our review of the card's wording and evidence, not a project release or product update.

Review sources and full receiptFull receipt: before, after and review evidence (JSON)