Rung 01 / What the model can do, per token

Models

What the model can do, and what that costs per token.

14
Sources
4
Concepts
6
Entities
A machined steel model block on a copper base, on paper.

In scope: quantization, mixture-of-experts, distillation, context handling, evals — and the specific-and-local cluster: fine-tuning, personalisation, open-weight releases, and running models on your own compute. Out: general AI research that moves neither cost nor capability per token, and the system around the model, which lives in harnesses.

Eight expert blocks under a router bar, with only two lit in copper — sparse activation.
QuantizationMoEDistillationFine-tuningOpen-weight
Sources compiled for this topic
TypeSourcePublished
PAPERMechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
Usman Naseem · Macquarie University

Comprehensive survey mapping mechanistic interpretability techniques to LLM alignment objectives, with a future research roadmap emphasizing automated interpretability and interpretability-driven alignment scaling.

2026-02-01
PAPERConditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Xin Cheng, Wangding Zeng, Damai Dai, et al. · DeepSeek / Peking University (collaboration)

Conditional memory as a sparsity axis orthogonal to MoE, via Engram (O(1) n-gram lookup). Sparsity Allocation problem yields a U-shaped scaling law (compute vs static memory); scaled to 27B params with gains on BBH +5.0, ARC-C +3.7, MMLU +3.4, NIAH 84.2→97.0.

2026-01-12
PAPERLoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
Ahmadreza Jeddi, Marco Ciccone, Babak Taati · Not specified (ICLR 2026)

Looped Transformer with elastic, budget-conditioned depth: a shortcut-consistency training scheme aligns reasoning trajectories of different lengths so one model trades inference depth for compute at test time — latent reasoning in weight-space rather than via explicit CoT tokens.

2026-02-11
PAPERTiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in LLM Inference Services
Mingdeng Du · Independent / not stated in abstract

First systematic empirical analysis of LLM token pricing across 3,237+ models (2020-2026); ~600x price decline; 'Tiered Super-Moore' hypothesis (economy 1.10yr / mid 1.55yr price half-life vs 2yr Moore benchmark; flagship/reasoning resists via ~31.5x premium); cost decline ~103.7% software/architecture-driven, ~-0.9% hardware.

2026-03-30
PAPERLLM Reasoning Is Latent, Not the Chain of Thought
arXiv preprint authors · Multi-institution

Argues LLM reasoning should be studied as latent-state trajectory formation, not faithful surface CoT — implications for interpretability, alignment, and training-objective design

2026-04-23
REPORTAnthropic Transformer Circuits Thread & Circuit Tracing
Anthropic Interpretability Team (Elhage, Olsson, Bricken, Templeton, Ameisen, Lindsey et al.) · Anthropic

Running research thread: features, circuits, superposition, attribution graphs, circuit tracing tools

2025-03-27
REPORTMeta Muse Spark — Native Multimodal Foundation Model with Contemplating Mode
Meta AI / industry coverage · Meta

Meta announces Muse Spark — natively multimodal (text/image/voice in single transformer) with Contemplating mode that orchestrates parallel sub-agents for deeper reasoning without latency penalty

2026-04-08
REPORTGemini 3 — Google's Latest Multimodal + Agentic Foundation Model
Google · Google DeepMind

Google releases Gemini 3 — claimed best-in-world multimodal understanding and most powerful agentic model. Improved tool-use, planning, and rich multimodal output over Gemini 2.5

2026-04-15
REPORTAnthropic Claude Platform Release Notes — June 25 to July 15, 2026
Anthropic · Anthropic

First-party changelog: Claude Sonnet 5 launch (1M context, new tokenizer +~30% tokens, adaptive-thinking-only), agent-memory-2026-07-22 beta header replacing managed-agents-2026-04-01 on memory-store endpoints, Admin API user management for Enterprise orgs, API key expiration controls, Dreams managed-agent preview expands to Sonnet 5/Fable 5.

2026-07-15
REPORTKimi K3: Open Frontier Intelligence
Kimi Team (400+ authors) · Moonshot AI

PRIMARY technical report (supersedes the secondary Willison entry on architecture): 2.8T total / 104B activated per step = ~3.7% activation, 16 of 896 routed experts, 1M context, native vision, full open weights. ~2.5x improvement in overall scaling efficiency over Kimi K2 from Kimi Delta Attention + Attention Residuals + Stable LatentMoE. Post-training RL across general/agentic/coding with multiple reasoning-effort levels. Authors state it still trails Claude Fable 5 and GPT-5.6 Sol.

2026-07-27
ANALYSISMechanistic Interpretability — 10 Breakthrough Technologies 2026
MIT Technology Review · MIT Technology Review

Named mech interp as 2026 breakthrough; Anthropic microscope + CoT monitoring advances

2026-01-12
ANALYSISKimi K3 — Moonshot AI's 2.8T-Parameter Open-Weight Model
Simon Willison (independent analysis; corroborated by Bloomberg/Fortune/Axios/CNBC/Forbes/Tom's Hardware coverage of Moonshot AI's announcement) · Moonshot AI (subject)

Moonshot AI released Kimi K3, a 2.8T-parameter MoE (896 experts, 16 active/token), 1M context, native vision — largest open-weight model to date; API live at $3/$15 per Mtok (Sonnet-tier pricing, most expensive Chinese-lab release yet); open weights due 2026-07-27; tops Arena.ai Frontend Code leaderboard ahead of Claude Fable 5.

2026-07-16
ANALYSISDeepSeek V4 — Open-Weight Trillion-Parameter MoE at ~1/6th Frontier Cost
VentureBeat (corroborated by morphllm.com + Hugging Face model card) · DeepSeek

DeepSeek V4 (open-weight MIT): V4-Pro 1.6T/49B-active, V4-Flash 284B/13B-active, 1M context, ~$0.435/$0.87 per-Mtok — near-frontier capability (SWE-bench Verified ~80.6%, GPQA ~90–92%) at ~1/6th the cost of Opus 4.7 / GPT-5.5; lands inside an 8-day April-2026 frontier window.

2026-04-24
ANALYSISGPT-5.6 — OpenAI's Three-Tier (Sol / Terra / Luna) Frontier Family + Token-Efficiency Claims
Tech Startups (corroborated by CryptoBriefing; OpenAI's own page returned HTTP 403) · OpenAI

OpenAI launched GPT-5.6 as three tiers (Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per Mtok) on 2026-07-09; claims 54% higher token efficiency on agentic coding vs rivals plus a 90% cached-read discount — a live instance of tier-differentiated frontier pricing.

2026-07-09
Models | Knowledge Base | MenFem