CARD · Card record
Demystifying agent skills: when they work, and until when
The paper runs controlled experiments plus paired trajectory analysis over 8,135 normalized trial records (238 valid open-coded labels) to answer when skills help, why they work, and where they fail. Skills mostly act as procedural anchors that stabilize noisy execution (65.7% of cases), beating Workflow Memory by 6.06 points in matched comparisons; cross-framework reuse and retrieval difficulty are the main failure modes.
01
THE STORY · The original
Read the explanation and its limits.
- Originalinstitutional researchhttps://arxiv.org/abs/2608.14036
Why it is worth understanding
Skills are already everyday machinery in tools like Claude Code, but most discussion stops at 'do they help'. This maps the failure boundaries (retrieval difficulty, cross-framework reuse), which directly informs which workflows are worth distilling into skills.
Conditions and limitations
Evidence is the authors' controlled benchmark study — not proof from our production re-test; the procedural-anchoring effect on benchmarks may not transfer to arbitrary private workflows.
Reading progress is saved in this browser. Save or add a note →
02
THE THREAD · Full timeline
1 event, every row traceable.
A card accumulates records as its story develops: new corroborators, market contrasts, material changes at the original — each dated and linked to its archive day.
- Added to BitShovelEntered the BitShovel radar: 100 entries visible on AlphaSignal daily at collection.
03
SOURCES · Sources and evidence
Check each link's purpose and limits.
Sources and evidence · 2 links
Each link has a purpose and scope. Link counts do not establish independent confirmation; direct support applies only to the statements specified below.
- arXiv paper via the AlphaSignal dailyScope not documented
A per-link statement of support has not been attached.
- AlphaSignal dailyScope not documented
A per-link statement of support has not been attached.