ARCHIVE · DAILY SELECTIONS

Selections for September 5, 2026

The retained records contain 4 focus choices and 7 items selected for the homepage that day. This page retains the cumulative choices. Saved snapshots show the record at selection; other entries are explicitly labeled as current information.

Selections and rotations that day · 6

These are the selection records we retained. Record times do not establish when every change went live; missing times are not inferred.

  1. Recovered from historical material

    2 in focus · Everything today count not retained

  2. Recovered from historical material

    2 in focus · 3 in Everything today

  3. Recovered from historical material

    3 in focus · 4 in Everything today

  4. Recovered from historical material

    3 in focus · 6 in Everything today

  5. Selection record

    0 in focus · 6 in Everything today

    No focus items in this record

  6. Selection record

    0 in focus · 7 in Everything today

    No focus items in this record

Featured during the day · 4

01Open-source projects · New

Perplexity open-sources Lily, a Metal inference server for exactly one checkpoint

Record at selection · Recovered from historical material; the display timing and wording may be incomplete.

Perplexity open-sourced Lily (Apache-2.0) in the pplx-garden repo: a Rust+Metal Mac inference server whose load-time validation admits exactly one checkpoint — mlx-community/Qwen3.6-35B-A3B-4bit with the HF revision pinned; any other architecture or quantization is rejected. It exposes a minimal OpenAI chat-completions subset, decodes greedily, keeps a two-entry LRU prefix cache, and requires Apple GPU family 10+ with macOS 26+. 704 stars at capture. AlphaSignal reports 1.35x over MLX-LM; the README states no such number but ships dated benchmark reports with reproduction steps.

Why it was selectedAn official open-source capability teams can exercise the same day: one checkpoint plus fail-closed load validation is a pointed engineering bet for local inference on Apple Silicon, and the pinned weight revision under Apache-2.0 makes the benchmark reproducible by anyone. The 1.35x figure is aggregator-reported, but the repo ships reproduction steps to check it.

CaveatEvidence: Perplexity's official repository README plus AlphaSignal and r/LocalLLaMA aggregation (BitShovel fetched and verified the repo on 2026-09-03). The 1.35x speedup does not appear in the README — it is an aggregator claim awaiting independent reproduction; the support surface is narrow (M5-and-newer Apple silicon, macOS 26+, a single 4-bit checkpoint); stars are attention, not proof of quality; this is not an independent review.

Open originalEvidencePerplexity's official repository, surfaced via AlphaSignal and r/LocalLLaMA
Full record and later changes →
02Open-source projects · New

Give Your Coding Agents a Memory You Own

Record at selection · Recovered from historical material; the display timing and wording may be incomplete.

Hugging Face blog published "Give Your Coding Agents a Memory You Own". The line above and the linked original carry the substance; judge conclusions on the original's own evidence.

Why it was selectedWhat it represents: official open-source capability is verifiable hands-on day one — early adopters get the edge; carried by 2 independent feeds in one window, not noise. Limits: author or press copy with no independent review; multi-source collection is not quality or adoption — the value holds only after checking the original's data, and some origin domains are unstable from mainland China.

CaveatSame-story corroboration proves multi-source collection and reachability only. Titles and claims are the creator self-report of the author or the carrying feeds' relay — not an independent BitShovel review, not proof of quality or adoption; the surrounding discussion is a discussion snapshot of attention, and attention is not proof of fact. Some origin domains may be unstable from mainland China.

Open originalEvidenceHugging Face blog original, carried same-story by Hugging Face blog, AIHOT featured
Full record and later changes →
03Open-source projects · New

Qwen open-sources Qwen3.8-Flash-Next as an early Qwen4 architecture preview

Record at selection · Recovered from historical material; the display timing and wording may be incomplete.

Alibaba's Qwen team announced the open-sourcing of Qwen3.8-Flash-Next together with an FP8 variant, positioning it officially as a multimodal MoE early preview built on the next-generation Qwen4 architecture so developers can adapt ahead of the full Qwen4 family. The official ModelScope model page for the model is already live. No concrete specs were disclosed at release, and circulating parameter counts have no official backing yet.

Why it was selectedThe frontier-model comparison lane previously had no vertical anchor for Qwen; this first architectural preview of the Qwen4 family directly shapes version expectations across the open-source ecosystem.

CaveatEvidence anchors on the Qwen blog address and the ModelScope model page going live, discovered through the AIHOT featured feed snapshot; none of this is an independent review, and release-day aggregator or community attention is not proof of any performance or cost claim — rumored figures like 125B total parameters are not adopted here.

Full record and later changes →
04Tools & products · New

Google ships Gemini 3.5 Transcribe into public preview

Record at selection · Recovered from historical material; the display timing and wording may be incomplete.

Google DeepMind launched Gemini 3.5 Transcribe, its self-described most precise speech-to-text model yet: Artificial Analysis measures average word error at 4.0% streaming and 2.6% non-streaming across 85+ auto-detected languages, and time to final transcription improves about 70% over Chirp 3. It ships as a public preview in the Gemini API and already powers Gboard on Android plus the macOS Gemini app. Two endpoints: streaming gemini-3.5-transcribe-live via the Live API, and gemini-3.5-transcribe for pre-recorded audio with speaker attribution and word-level timestamps.

Why it was selectedThe frontier-model lane also lacked a vertical anchor for Google Gemini; this lands squarely on voice workflows — 85-language coverage plus word-level timestamps is a clear capability jump for transcription, QA, and meeting-notes tooling.

CaveatEvery figure comes from Google DeepMind's own announcement and its cited Artificial Analysis measurements; this is not an independent evaluation, and official claims during a public preview are not proof of production stability or cost behavior.

Full record and later changes →
Other items shown that day · 3

Everything today on the homepage holds up to ten items. Other items shown during that day remain available here.

05Business & opportunities · Important update

The GPT-5.6 Sol price cut is loud; the real opportunity is unit economics

Current information · no copy snapshot retained from that time

At re-verification, OpenRouter's official model page still showed 50% off and now listed $2 input and $10 output per million tokens. The Hacker News discussion had 635 points and 448 comments.

Current value assessmentA sharp model-cost change rewrites the unit economics of batch, long-context, and high-frequency agent workflows. The useful move is to recalculate a real workload's costs and look for offerings that only now pencil out.

CaveatThis is a discussion snapshot from Aug 22, 2026. Points and comments measure attention and are not proof of buyers, revenue, profit, or a repeatable opportunity; prices come from OpenRouter's official page and may change at any time.

Full record and later changes →
06Platforms & policies · New

Zhipu ships GLM-5.3-Flash: frontier intelligence at flash pricing

Current information · no copy snapshot retained from that time

Zhipu's official blog announced GLM-5.3-Flash (the renamed ox-alpha the community had been tracking): ~321B-total MoE at FP8 mixed precision, MIT-licensed weights, zh/en with 400K context. A hands-on megathread opened on r/LocalLLaMA the same day, and third-party comparisons put it near Claude Opus at roughly a tenth of the price — a price-anchor event for the frontier market worth sustained tracking.

Current value assessmentOpen weights plus frontier claims at a tenth of the price pressure both the market's pricing and the open/closed divide; follow-up comparisons, price responses, and quantized variants are the natural next chapters of this vertical.

Caveat[Evidence provenance: official blog and HF card plus third-party retelling] The "Opus-level at a tenth of the price" framing follows AlphaSignal's comparison, not an independent review; local hands-on numbers await the megathread, and attention is not proof.

Full record and later changes →

7 items shown during the day · 7 in the last Everything today selection