Apps & Tools
When Agents Slow Down: Elo-per-token Analysis
A suite of tools for understanding LLM agents' test-time strategies and compute scaling.
PlayableOpen source
When Agents Slow Down: Elo-per-token Analysis media is blocked
Allow external media to connect to the provider and play this content.
Description
A research project and tool suite that applies Bradley–Terry models to submission streams to compare AI agents, sampling baselines, and human contestants on a unified Elo scale. It identifies a 'scaling inflection point' where continuing an agent session becomes less efficient than independent restarts, highlighting the gap in continual learning between agents and humans.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.