The bibliography
The reading
Every paper on the 2026 H1 list, in the order it was published, section by section. Each row links out to the paper itself; a row shows a close-read only when one exists in the Atlas. Working through this is the point — the list is the queue, not the achievement.
Sebastian Raschka · Ahead of AIThe source list ↗
CAPTURED

The ten
One must-read per section, chosen by the source author- §1Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning ↗★
- §2Efficient Pre-Training with Token Superposition ↗★
- §3FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling ↗★
- §4DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗★
- §5Test-Time Scaling Makes Overtraining Compute-Optimal ↗★
- §6You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories ↗★
- §7Is Grep All You Need? How Agent Harnesses Reshape Agentic Search ↗★
- §8Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? ↗★
- §9LLaDA2.1: Speeding Up Text Diffusion via Token Editing ↗★
- §10AlphaEval: Evaluating Agents in Production ↗★
The sections
10 sections · 165 papers- Deep Delta Learning ↗
- MiMo-V2-Flash Technical Report ↗
- Ministral 3 ↗
- Scaling Embeddings Outperforms Scaling Experts in Language Models ↗
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ↗
- ERNIE 5.0 Technical Report ↗
- ViT-5: Vision Transformers for the Mid-2020s ↗
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters ↗
- Nanbeige4.1-3B: A Small General Model That Reasons, Aligns, and Acts ↗
- Symmetry in Language Statistics Shapes the Geometry of Model Representations ↗
- GLM-5: From Vibe Coding to Agentic Engineering ↗
- Arcee Trinity Large Technical Report ↗
- The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks ↗
- Tiny Aya: Bridging Scale and Multilingual Depth ↗
- Attention Residuals ↗
- Mamba-3: Improved Sequence Modeling Using State Space Principles ↗
- Attention to Mamba: A Recipe for Cross-Architecture Distillation ↗
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning ↗★
- ZAYA1-8B Technical Report ↗
- Delta Attention Residuals ↗
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention ↗
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence ↗
- Self-Distillation Enables Continual Learning ↗
- Shaping Capabilities with Token-Level Data Filtering ↗
- Self-Improving Pretraining: Using Post-Trained Models to Pretrain Better Models ↗
- BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking ↗
- OPUS: Towards Efficient and Principled Data Selection in LLM Pre-training in Every Iteration ↗
- Test-Time Training with KV Binding Is Secretly Linear Attention ↗
- Progressive Residual Warmup for Language Model Pretraining ↗
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights ↗
- MiCA Learns More Knowledge Than LoRA and Full Fine-Tuning ↗
- In-Place Test-Time Training ↗
- Hybrid Policy Distillation for LLMs ↗
- Large Language Models Explore by Latent Distilling ↗
- Efficient Training on Multiple Consumer GPUs with RoundPipe ↗
- Efficient Pre-Training with Token Superposition ↗★
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs ↗
- Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery ↗
- Efficient Remote KV Cache Reuse with GPU-Native Video Codec ↗
- SpargeAttention2: Trainable Sparse Attention via Hybrid Top-K+Top-P Masking and Distillation Fine-Tuning ↗
- Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers ↗
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling ↗★
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation ↗
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning ↗
- A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints ↗
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection ↗
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse ↗
- Mixture-of-Depths Attention ↗
- TriAttention: Efficient Long Reasoning with Trigonometric KV Compression ↗
- Sessa: Selective State Space Attention ↗
- Contexts Are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets ↗
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗★
- Long Context Pre-Training with Lighthouse Attention ↗
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps ↗
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning ↗
- Learning to Reason in 13 Parameters ↗
- InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning ↗
- Learning to Self-Verify Makes Language Models Better Reasoners ↗
- Does Your Reasoning Model Implicitly Know When to Stop Thinking? ↗
- Reasoning Models Struggle to Control Their Chains of Thought ↗
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs ↗
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation ↗
- Test-Time Scaling Makes Overtraining Compute-Optimal ↗★
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets ↗
- When Can LLMs Learn to Reason with Weak Supervision? ↗
- AI Co-Mathematician: Accelerating Mathematicians with Agentic AI ↗
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling ↗
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling ↗
- Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling ↗
- GDPO: Group Reward-Decoupled Normalization Policy Optimization for Multi-Reward RL Optimization ↗
- Your Group-Relative Advantage Is Biased ↗
- Reasoning Models Generate Societies of Thought ↗
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability ↗
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation ↗
- Reinforcement Learning via Self-Distillation ↗
- Expanding the Capabilities of Reinforcement Learning via Text Feedback ↗
- iGRPO: Self-Feedback-Driven LLM Reasoning ↗
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning ↗
- Experiential Reinforcement Learning ↗
- Discovering Implicit Large Language Model Alignment Objectives ↗
- OpenClaw-RL: Train Any Agent Simply by Talking ↗
- Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought ↗
- Self-Distilled RLVR ↗
- Target Policy Optimization ↗
- FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling ↗
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks ↗
- Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges ↗
- SEIF: Self-Evolving Reinforcement Learning for Instruction Following ↗
- Process Rewards with Learned Reliability ↗
- Self-Distilled Agentic Reinforcement Learning ↗
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories ↗★
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning ↗
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ↗
- Trust Region On-Policy Distillation ↗
- Dr. Zero: Self-Evolving Search Agents Without Training Data ↗
- Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline ↗
- Large Language Model Agents Are Not Always Faithful Self-Evolvers ↗
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems ↗
- AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent ↗
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments ↗
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks ↗
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook ↗
- AgentLAB: Benchmarking LLM Agents Against Long-Horizon Attacks ↗
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use ↗
- Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems ↗
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents ↗
- Hyperagents ↗
- Multi-User Large Language Model Agents ↗
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents ↗
- Natural-Language Agent Harnesses ↗
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures ↗
- The Art of Building Verifiers for Computer Use Agents ↗
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering ↗
- Autogenesis: A Self-Evolving Agent Protocol ↗
- From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company ↗
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond ↗
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills ↗
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate ↗
- Recursive Multi-Agent Systems ↗
- ClawGym: A Scalable Framework for Building Effective Claw Agents ↗
- Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes ↗
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search ↗★
- AI for Auto-Research: Roadmap & User Guide ↗
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents ↗
- Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost ↗
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills ↗
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems ↗
- AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development ↗
- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding ↗
- AutoHarness: Improving LLM Agents by Automatically Synthesizing a Code Harness ↗
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? ↗★
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-Horizon Agentic Programming in Command-Line Interfaces ↗
- On Data Engineering for Scaling LLM Terminal Capabilities ↗
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale ↗
- Qwen3-Coder-Next Technical Report ↗
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? ↗
- Coding Agents Are Effective Long-Context Processors ↗
- Effective Strategies for Asynchronous Software Engineering Agents ↗
- Meta-Harness: End-to-End Optimization of Model Harnesses ↗
- Scaling Coding Agents via Atomic Skills ↗
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play ML Pipeline for Connect Four That Performs Comparably to an External Solver ↗
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses ↗
- Code as Agent Harness ↗
- Deferred Commitment Decoding for Diffusion Language Models ↗
- Generative Modeling via Drifting ↗
- LLaDA2.1: Speeding Up Text Diffusion via Token Editing ↗★
- dLLM: Simple Diffusion Language Modeling ↗
- Rethinking Token Prediction: Tree-Structured Diffusion Language Model ↗
- Continuous Latent Diffusion Language Model ↗
- How to Train Your Latent Diffusion Language Model Jointly with the Latent Space ↗
- TextLDM: Language Modeling with Continuous Latent Diffusion ↗
- Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding ↗
- Benchmark^2: Systematic Evaluation of LLM Benchmarks ↗
- OdysseyArena: Benchmarking Large Language Models for Long-Horizon, Active and Inductive Interactions ↗
- Large Language Model Reasoning Failures ↗
- Emergent Misalignment Is Easy, Narrow Misalignment Is Hard ↗
- Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips ↗
- LLMStructBench: Benchmarking Large Language Model Structured Data Extraction ↗
- Quantifying Construct Validity in Large Language Model Evaluations ↗
- Lost in Stories: Consistency Bugs in Long Story Generation by LLMs ↗
- ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence ↗
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs ↗
- ClawArena: Benchmarking AI Agents in Evolving Information Environments ↗
- AlphaEval: Evaluating Agents in Production ↗★
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use ↗
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.
Graduation is computed on every render, never maintained by hand: a row links to its close-read when the paper’s arXiv id matches a source already in the Atlas. 0 of 165 have one. The two sets were selected on different criteria — this list adds to the shelf rather than confirming it — so that number is the queue’s progress, and it starts where every queue starts. The Atlas →