Local AI “intelligence per watt,” SAS context-ranking sparsifies Transformers, and LIT trains robotics to avoid visual shortcuts—plus humanoid...‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 
September 15, 2026

🔬 Today in AI Research

From 35 articles considered today, here are the highlights — your daily brew.

📋 Today's Research

🔬 Research of the Day

⚡ Measuring how much real-world intelligence you get per watt from local AI.
INTELLIGENCE PER WATT
Source: arxiv.org
Quick Brief:
Introduces Intelligence per Watt (IPW) to measure how efficiently small, local LLMs (≤20B) on consumer devices answer real-world queries. Using 1M queries, 20+ models, and 8 hardware platforms, the paper finds local AI can handle most everyday workloads, with fast efficiency gains but still behind cloud hardware.
The Details:
  • IPW = accuracy (vs. frontier models) per unit power, per model–hardware combo.
  • Local LMs correctly handle 88.7% of 1M real chat/reasoning queries.
  • From 2023–2025, IPW improved 5.3×; locally serviceable queries rose 23.2% → 71.3%.
  • For the same model, local accelerators have ≥1.4× lower IPW than cloud accelerators.
Why It Matters:
Gives a single metric for “useful intelligence per watt.” Shows most common queries can move from cloud to local devices, and highlights a clear optimization target for edge AI hardware and small models.

💡 Worth a Closer Look

🧠 SAS learns which context to keep, making Transformer attention sparse and efficient without sacrificing accuracy.
SAS
Source: huggingface.co
Quick Brief:
SAS is a post‑training method that learns which tokens/blocks to keep by adding a continuous gate to attention logits, optimizing sparse attention directly with LM loss.
The Details:
  • Selector scores go inside the attention softmax.
  • Gates are softmax‑normalized vs. the always‑kept current block.
  • Train with continuous scores; infer with Top‑K under a fixed budget.
Why It Matters:
Uses limited attention more effectively, aligning sparsification with task loss for better long‑context performance at lower cost.

📝 Also Noteworthy

🤖 Two-stage LIT training makes robot foundation models robust by forcing them to use geometry instead of visual shortcuts.
LIT
Source: huggingface.co
Quick Brief:
LIT is a two-stage method that makes robot foundation models use geometry (poses) instead of visual shortcuts, improving robustness to new cameras, lighting, and clutter.
The Details:
  • Train an action prior without images from language, robot state, and terminal SE(3) pose.
  • Add a pose-supervised latent that fuses vision and language and is the only visual input, forced to recover goal poses.
  • Gains: +3.9–10.7 on LIBERO-Plus, +13.3–16.7 in real-world visual shifts.
Why It Matters:
Cuts vision–action shortcuts with a simple, architecture-agnostic recipe for more reliable, generalizable robot policies.

👀 One More to Watch

🤖 Physics-grounded benchmark tests if multimodal LLMs can safely control humanoid robots in sudden home hazards.
REACTHUMAN
Source: huggingface.co
Quick Brief:
ReactHuman is a physics-grounded benchmark where a multimodal LLM controls a simulated humanoid facing sudden household hazards, testing safe, physically correct split‑second actions.
The Details:
  • 1,000+ rigid-body scenes, 17 hazard types
  • Adversarial objects (foam anvil, steel apple) test physics vs. appearance
  • Five metrics score reasonable, safe, physics-grounded behavior
  • Seven MLLMs mishandle ~1/3 of hazards; scaling doesn’t fix it
Why It Matters:
Shows if LLM-based robot controllers are reliable in real-time home hazards and that safety needs better training/control, not just bigger models.

📚 More Worth Reading