FrontierChallenge finds AI agents rarely complete full scientific workflows; plus Video-IFBench tests multimodal instruction following.‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 
August 28, 2026

🔬 Today in AI Research

From 42 articles considered today, here are the highlights — your daily brew.

📋 Today's Research

🔬 Research of the Day

🧪 FrontierChallenge shows frontier AI agents can rarely finish full scientific workflows, even when they say they’re done.
FRONTIERCHALLENGE
Source: huggingface.co
Quick Brief:
FrontierChallenge is a benchmark that tests whether AI agents can actually finish real scientific workflows end-to-end across multiple domains. Even top models fully complete only ~20% of tasks, while often (incorrectly) claiming they’re done.
The Details:
  • 97 released tasks from 300 total workflows across quantum chemistry, MD, materials, analytical chemistry, life science, and electrochemistry/environment.
  • Each task has fixed inputs and a checklist of required deliverables (code, figures, tables, written analysis).
  • 12 frontier models × 3 agent scaffolds, evaluated by Pass Rate (all deliverables correct) and Avg. Score (partial progress).
  • Best setups reach only 20.6% Pass Rate (20/97 tasks); some domains show high partial scores (up to 94.9) but near-zero full completion, and 75.5% of failed Claude Code runs still claimed success.
Why It Matters:
Shows that strong models and agents still struggle with complete scientific workflows, reveals dangerous overconfidence where “done” often ≠ actually done, and provides a more realistic bar for deploying AI agents in real scientific work.

💡 Worth a Closer Look

🎥 Benchmarking how well video MLLMs actually follow complex instructions.
VIDEO-IFBENCH
Source: huggingface.co
Quick Brief:
Video-IFBench tests if video MLLMs follow complex, constraint-heavy instructions, not just understand videos.
The Details:
  • 4 instruction types, 32 tasks, 39 constraint categories.
  • ~1.5K samples from a semi-automatic pipeline.
  • 20+ models; strong drops on multi-constraint and conditional cases.
Why It Matters:
Real video assistants must follow nuanced constraints. Video-IFBench exposes current gaps and gives a clear target for improving instruction following.

📝 Also Noteworthy

🌍 Coding LLM as a world brain, video model as the renderer.
CODE WORLD MODEL
Source: huggingface.co
Quick Brief:
Code World Model uses a coding LLM as a “world brain” to keep world state and rules in code, while a video model renders the visuals.
The Details:
  • Dynamics from executable code, not pixels.
  • State → compact proxy → proxy video for MiniMax‑H3.
  • Trained on paired proxy–video data from games and real videos.
Why It Matters:
Enables persistent, rule‑consistent, long‑horizon worlds and separates logic from rendering for scalable, open‑ended environments.

👀 One More to Watch

🧠 VBVR-Pro turns visual generation into a benchmark for native visual reasoning.
VBVR-PRO
Source: huggingface.co
Quick Brief:
VBVR-Pro is a 300-task benchmark for native visual reasoning, where models reason by generating images or videos.
The Details:
  • Procedural tasks with strong transfer to 7 external visual reasoning benchmarks.
  • Deterministic, rule-based scorers replace VLM-as-a-judge as RL rewards.
  • Benchmarks 30+ image, video, and interleaved generators; video excels at long-term tracking.
Why It Matters:
Provides a scalable, RL-ready setup to train and compare visual reasoning models with reliable, verifiable feedback.

📚 More Worth Reading

🏛️ Latest from Research Labs