Rung 01 / What the model can do, per token

Models

What the model can do, and what that costs per token.

14
Sources
4
Concepts
6
Entities
A machined steel model block on a copper base, on paper.

In scope: quantization, mixture-of-experts, distillation, context handling, evals — and the specific-and-local cluster: fine-tuning, personalisation, open-weight releases, and running models on your own compute. Out: general AI research that moves neither cost nor capability per token, and the system around the model, which lives in harnesses.

Eight expert blocks under a router bar, with only two lit in copper — sparse activation.
QuantizationMoEDistillationFine-tuningOpen-weight
Paper only5Show all →
Sources compiled for this topic
TypeSourcePublished
PAPERMechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
Usman Naseem · Macquarie University

Comprehensive survey mapping mechanistic interpretability techniques to LLM alignment objectives, with a future research roadmap emphasizing automated interpretability and interpretability-driven alignment scaling.

2026-02-01
PAPERConditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Xin Cheng, Wangding Zeng, Damai Dai, et al. · DeepSeek / Peking University (collaboration)

Conditional memory as a sparsity axis orthogonal to MoE, via Engram (O(1) n-gram lookup). Sparsity Allocation problem yields a U-shaped scaling law (compute vs static memory); scaled to 27B params with gains on BBH +5.0, ARC-C +3.7, MMLU +3.4, NIAH 84.2→97.0.

2026-01-12
PAPERLoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
Ahmadreza Jeddi, Marco Ciccone, Babak Taati · Not specified (ICLR 2026)

Looped Transformer with elastic, budget-conditioned depth: a shortcut-consistency training scheme aligns reasoning trajectories of different lengths so one model trades inference depth for compute at test time — latent reasoning in weight-space rather than via explicit CoT tokens.

2026-02-11
PAPERTiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in LLM Inference Services
Mingdeng Du · Independent / not stated in abstract

First systematic empirical analysis of LLM token pricing across 3,237+ models (2020-2026); ~600x price decline; 'Tiered Super-Moore' hypothesis (economy 1.10yr / mid 1.55yr price half-life vs 2yr Moore benchmark; flagship/reasoning resists via ~31.5x premium); cost decline ~103.7% software/architecture-driven, ~-0.9% hardware.

2026-03-30
PAPERLLM Reasoning Is Latent, Not the Chain of Thought
arXiv preprint authors · Multi-institution

Argues LLM reasoning should be studied as latent-state trajectory formation, not faithful surface CoT — implications for interpretability, alignment, and training-objective design

2026-04-23
Models | Knowledge Base | MenFem