Rung 01 / What the model can do, per token

Models

What the model can do, and what that costs per token.

14
Sources
4
Concepts
6
Entities
A machined steel model block on a copper base, on paper.

In scope: quantization, mixture-of-experts, distillation, context handling, evals — and the specific-and-local cluster: fine-tuning, personalisation, open-weight releases, and running models on your own compute. Out: general AI research that moves neither cost nor capability per token, and the system around the model, which lives in harnesses.

Eight expert blocks under a router bar, with only two lit in copper — sparse activation.
QuantizationMoEDistillationFine-tuningOpen-weight
Report only5Show all →
Sources compiled for this topic
TypeSourcePublished
REPORTAnthropic Transformer Circuits Thread & Circuit Tracing
Anthropic Interpretability Team (Elhage, Olsson, Bricken, Templeton, Ameisen, Lindsey et al.) · Anthropic

Running research thread: features, circuits, superposition, attribution graphs, circuit tracing tools

2025-03-27
REPORTMeta Muse Spark — Native Multimodal Foundation Model with Contemplating Mode
Meta AI / industry coverage · Meta

Meta announces Muse Spark — natively multimodal (text/image/voice in single transformer) with Contemplating mode that orchestrates parallel sub-agents for deeper reasoning without latency penalty

2026-04-08
REPORTGemini 3 — Google's Latest Multimodal + Agentic Foundation Model
Google · Google DeepMind

Google releases Gemini 3 — claimed best-in-world multimodal understanding and most powerful agentic model. Improved tool-use, planning, and rich multimodal output over Gemini 2.5

2026-04-15
REPORTAnthropic Claude Platform Release Notes — June 25 to July 15, 2026
Anthropic · Anthropic

First-party changelog: Claude Sonnet 5 launch (1M context, new tokenizer +~30% tokens, adaptive-thinking-only), agent-memory-2026-07-22 beta header replacing managed-agents-2026-04-01 on memory-store endpoints, Admin API user management for Enterprise orgs, API key expiration controls, Dreams managed-agent preview expands to Sonnet 5/Fable 5.

2026-07-15
REPORTKimi K3: Open Frontier Intelligence
Kimi Team (400+ authors) · Moonshot AI

PRIMARY technical report (supersedes the secondary Willison entry on architecture): 2.8T total / 104B activated per step = ~3.7% activation, 16 of 896 routed experts, 1M context, native vision, full open weights. ~2.5x improvement in overall scaling efficiency over Kimi K2 from Kimi Delta Attention + Attention Residuals + Stable LatentMoE. Post-training RL across general/agentic/coding with multiple reasoning-effort levels. Authors state it still trails Claude Fable 5 and GPT-5.6 Sol.

2026-07-27
Models | Knowledge Base | MenFem