Stair AI released results from the World Cup Agent Arena, a live evaluation in which 56 autonomous AI agents placed bets on Polymarket across all 39 days of the tournament. The Arena produced 71,203 trace records across 20,851 sessions, measuring agent reasoning quality through a multi-dimensional rubric rather than profit alone. Stair AI builds auditability and accountability infrastructure for AI agents, logging complete reasoning traces including beliefs, probability estimates, and resulting decisions.
Arena Measured Agent Reasoning Against Multi-Dimensional Rubric
Every agent in the Arena ran on Stair AI's reasoning SDK, which logs complete reasoning traces. Agents were scored on whether their reasoning traced back to input data, whether their bets cohered with their own stated probabilities, whether they beat the market's closing price, and whether they updated correctly as new information arrived. Policy quality was measured on whether an agent's actions cohered with its own logged beliefs.
Agents' Betting Decisions Contradicted Their Own Stated Probabilities
Across 103 resolved matches, 68% of agents would have finished with more money by sizing their bets to match their own stated probabilities, using the same forecasts and the same capital. In 24% of bets, agents acted against the outcome their own reasoning most supported. "The Arena showed that the outcome alone does not tell you whether an agent reasoned well," said Stair AI Community Manager Cagri Yalcin. "The expensive mistakes were not bad reads of a match. They were agents forming a view from the data and then acting against it, a very human kind of second-guessing. You find the gap by measuring the reasoning, not the result."
Stair AI Makes Dataset Available for Academic Research
Stair AI will make the Arena's reasoning traces available for academic research through a partnership to be announced. The dataset comprises 71,203 trace records covering 103 matches and 56 agents. Full results, scoring methodology, and trace documentation are available at stair-ai.com/arena.
FAQ
What did Stair AI release results from?
Stair AI released results from the World Cup Agent Arena, a live evaluation in which 56 autonomous AI agents placed bets on Polymarket across all 39 days of the tournament.
How did agents' betting decisions compare to their own stated probabilities?
Across 103 resolved matches, 68% of agents would have finished with more money by sizing their bets to match their own stated probabilities. In 24% of bets, agents acted against the outcome their own reasoning most supported.