Rung 01 / What the model can do, per token

Models

What the model can do, and what that costs per token.

14
Sources
4
Concepts
6
Entities
A machined steel model block on a copper base, on paper.

In scope: quantization, mixture-of-experts, distillation, context handling, evals — and the specific-and-local cluster: fine-tuning, personalisation, open-weight releases, and running models on your own compute. Out: general AI research that moves neither cost nor capability per token, and the system around the model, which lives in harnesses.

Eight expert blocks under a router bar, with only two lit in copper — sparse activation.
QuantizationMoEDistillationFine-tuningOpen-weight
News only2Show all →
Sources compiled for this topic
TypeSourcePublished
ANALYSISDeepSeek V4 — Open-Weight Trillion-Parameter MoE at ~1/6th Frontier Cost
VentureBeat (corroborated by morphllm.com + Hugging Face model card) · DeepSeek

DeepSeek V4 (open-weight MIT): V4-Pro 1.6T/49B-active, V4-Flash 284B/13B-active, 1M context, ~$0.435/$0.87 per-Mtok — near-frontier capability (SWE-bench Verified ~80.6%, GPQA ~90–92%) at ~1/6th the cost of Opus 4.7 / GPT-5.5; lands inside an 8-day April-2026 frontier window.

2026-04-24
ANALYSISGPT-5.6 — OpenAI's Three-Tier (Sol / Terra / Luna) Frontier Family + Token-Efficiency Claims
Tech Startups (corroborated by CryptoBriefing; OpenAI's own page returned HTTP 403) · OpenAI

OpenAI launched GPT-5.6 as three tiers (Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per Mtok) on 2026-07-09; claims 54% higher token efficiency on agentic coding vs rivals plus a 90% cached-read discount — a live instance of tier-differentiated frontier pricing.

2026-07-09
Models | Knowledge Base | MenFem