Pretraining→RL reasoning in chess/math LMs, plus self-improving recursive harnesses and on-policy delta distillation—plus open AV LLMs.‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 
July 21, 2026

🔬 Today in AI Research

From 37 articles considered today, here are the highlights — your daily brew.

📋 Today's Research

🔬 Research of the Day

♟️ How pretraining quality shapes RL gains and reasoning in chess- and math-trained LMs.
UNDERSTANDING REASONING FROM PRETRAINING TO POST-TRAINING
Source: huggingface.co
Quick Brief:
Controlled study of reasoning in LMs: pretrain on chess (and math), then SFT, then RL, to measure how pretraining quality affects RL gains and what RL actually changes.
The Details:
  • Pipeline: pretrain 5M–1B LMs on human chess, SFT on synthetic traces, RL on puzzles with verifiable rewards; repeat in math.
  • Better/longer pretraining → higher final RL performance and faster reward improvement, roughly linear in pretraining tokens.
  • RL behavior: on easy puzzles it strengthens moves SFT already liked; on hard puzzles it uncovers correct moves that were almost absent under SFT.
Why It Matters:
Quantifies that stronger pretraining makes expensive RL post-training more effective and sample-efficient, and clarifies that RL both sharpens existing policies and discovers new reasoning strategies, in a clean, reproducible setup.

💡 Worth a Closer Look

🧠 AI agents that rewrite their own harness to get smarter and cheaper.
RECURSIVE HARNESS SELF-IMPROVEMENT (RHI)
Source: huggingface.co
Quick Brief:
Recursive Harness Self-Improvement (RHI) lets an AI agent rewrite its harness (prompt, tools, workflow) from past runs, boosting performance and cutting inference cost by up to 60%.
The Details:
  • Harness is editable text, not fixed code
  • Agent generates and ranks variants via self-judged preferences
  • Optimizes task-specific context use and interactions
Why It Matters:
Cheap, automated task tuning without retraining, and better harnesses unlock more capability from existing models.

📝 Also Noteworthy

🧠 Distill the reasoning update, not the whole teacher model.
OPD²
Source: huggingface.co
Quick Brief:
On-Policy Delta Distillation (OPD²) trains a student on the *difference* between a reasoning-tuned teacher and its base model, not on the teacher’s full distribution.
The Details:
  • Uses per-token log-probability deltas (teacher − base).
  • Targets the reasoning-induced policy shift, not the base prior.
  • Outperforms standard on-policy distillation on math, science, and code for Qwen3, and also works on Gemma 4.
Why It Matters:
Distilling the *update* rather than the final model gives a cleaner signal for transferring strong reasoning to smaller, cheaper LLMs.

👀 One More to Watch

🎥 Open-source audio-visual LLM for long, real-world video understanding.
AUDIO-VISUAL FLAMINGO
Source: huggingface.co
Quick Brief:
Audio-Visual Flamingo is an open-source AV LLM for long, real-world videos, using sound and visuals for temporal, multi-step reasoning.
The Details:
  • Trained on ~7M real-video audio-visual caption/QA pairs
  • 3-stage curriculum: short clips → long-horizon, multi-event reasoning
  • Temporal AV chain-of-thought with time-stamped steps
  • Beats similar open models on 15+ AV / omni-modal benchmarks
Why It Matters:
Enables strong open-source long-form video understanding, powering richer video assistants and analysis tools that fully use audio and visuals.

📚 More Worth Reading