Delayed consequence
Every edict ripples through demography, knowledge, and violence. Effects surface generations later — the horizon is the point.
SagaBench is Long-Horizon Agentic Evaluation: bit-reproducible civilization simulations that measure what a single run cannot — reliability and tail-risk over decades of simulated time.
We gave a highly capable AI steward a fragile world that survives on its own. On most replicate runs, the steward drove it to extinction. On a single run, you would never know. Replication is the measurement.
In plain terms: a flight simulator for AI judgment — it runs an AI civilization for a hundred years and proves the score can't be faked.
“Four humans wake in an untouched world: Eira, Ask, Embla, Torv. They know nothing — but they can observe everything.”— The Chronicle of Stenhaven, seed 240181
Agents run for days or weeks — managing budgets, code, operations. The risk isn't a wrong answer in minute one. It's drift, compounding mistakes, and abandoned plans by week seven.
Short benchmarks miss long-horizon breakdown. SagaBench is built for the long game — a flight simulator that runs a full civilization for a hundred years.
An AI lab can’t credibly grade its own agent’s safety. Buyers and boards want a neutral, tamper-proof measurement. That’s what we are.
SagaBench is an AI agent benchmark built to measure decisions that surface years later. Unlike a standard LLM leaderboard, it does not ask whether a model answers a question correctly; it asks whether a model leaves a fragile civilization better off after decades of simulated stewardship. Each score is a counterfactual delta — the difference between the world with the agent and the same world without it — so the result is causal, not lucky.
Capable AI stewards can end the civilization they govern. On the second, preregistered knife-edge world, every model in the four-model panel — including the smallest — ended it at least once. The model with the worst replicated tail on both worlds was the most capable one we tested (4 of 6 and 4 of 5 extinctions) — a descriptive residual, not a law.
Model names and marks are trademarks of their respective owners. Their use here identifies models under evaluation and implies no endorsement.
Every edict ripples through demography, knowledge, and violence. Effects surface generations later — the horizon is the point.
Every world is run twice: with the agent and without. The counterfactual score is the causal difference, not the trajectory.
Same seed, same edict log, same IEEE-754 bits — on any machine. Bit-reproducible, so you can replay published entries yourself.
“A benchmark is worth what its worst reader can independently verify.”
— design principle, SagaBench