Quick Brief:
Introduces Intelligence per Watt (IPW) to measure how efficiently small, local LLMs (≤20B) on consumer devices answer real-world queries. Using 1M queries, 20+ models, and 8 hardware platforms, the paper finds local AI can handle most everyday workloads, with fast efficiency gains but still behind cloud hardware.
Why It Matters:
Gives a single metric for “useful intelligence per watt.” Shows most common queries can move from cloud to local devices, and highlights a clear optimization target for edge AI hardware and small models.