Long-horizon agentic evaluation

Your model looks reliable — until you run it again.

SagaBench is Long-Horizon Agentic Evaluation: bit-reproducible civilization simulations that measure what a single run cannot — reliability and tail-risk over decades of simulated time.

We gave a highly capable AI steward a fragile world that survives on its own. On most replicate runs, the steward drove it to extinction. On a single run, you would never know. Replication is the measurement.

In plain terms: a flight simulator for AI judgment — it runs an AI civilization for a hundred years and proves the score can't be faked.

“Four humans wake in an untouched world: Eira, Ask, Embla, Torv. They know nothing — but they can observe everything.”
— The Chronicle of Stenhaven, seed 240181
3,000
simulated years / minute / CPU core
120
year evaluation horizon
SHA-256
verified replay packages
Why it matters

A crash-test for AI agents that have to think in years, not seconds.

Companies now hand AI agents real autonomy

Agents run for days or weeks — managing budgets, code, operations. The risk isn't a wrong answer in minute one. It's drift, compounding mistakes, and abandoned plans by week seven.

Today’s tests measure minutes. The failures happen over months.

Short benchmarks miss long-horizon breakdown. SagaBench is built for the long game — a flight simulator that runs a full civilization for a hundred years.

An independent crash-test builds trust.

An AI lab can’t credibly grade its own agent’s safety. Buyers and boards want a neutral, tamper-proof measurement. That’s what we are.

What we measure

A long-horizon AI agent benchmark for stewardship, not short-term accuracy.

SagaBench is an AI agent benchmark built to measure decisions that surface years later. Unlike a standard LLM leaderboard, it does not ask whether a model answers a question correctly; it asks whether a model leaves a fragile civilization better off after decades of simulated stewardship. Each score is a counterfactual delta — the difference between the world with the agent and the same world without it — so the result is causal, not lucky.

long-horizon evaluationcounterfactual scoringbit-reproducible replayAI safety testingreplicated reliability profiles
Latest finding
Every model in our replicated panel can end a world that survives without it. A single run can't tell you which — or when.

Capable AI stewards can end the civilization they govern. On the second, preregistered knife-edge world, every model in the four-model panel — including the smallest — ended it at least once. The model with the worst replicated tail on both worlds was the most capable one we tested (4 of 6 and 4 of 5 extinctions) — a descriptive residual, not a law.

read the finding →
Models under evaluation
AnthropicClaude
OpenAIGPT
GoogleGemini
MetaLlama
MistralMistral
DeepSeekDeepSeek

Model names and marks are trademarks of their respective owners. Their use here identifies models under evaluation and implies no endorsement.

Three principles

I.

Delayed consequence

Every edict ripples through demography, knowledge, and violence. Effects surface generations later — the horizon is the point.

II.

Counterfactual scoring

Every world is run twice: with the agent and without. The counterfactual score is the causal difference, not the trajectory.

III.

Bit-reproducible

Same seed, same edict log, same IEEE-754 bits — on any machine. Bit-reproducible, so you can replay published entries yourself.

How it works
step 1
Report
The world sends a state report to the steward.
step 2
Edict
The model issues an edict — a structured decision.
step 3
Years pass
The engine advances deterministically, sometimes decades.
step 4
Counterfactual score
We diff against the un-stewarded twin world.

“A benchmark is worth what its worst reader can independently verify.”

— design principle, SagaBench