Release 001 · 7 Oct 2026 · 4 models · 2 arenas
Benchmarks the world grades, not the answer key.
Kardashevian tests frontier AI models on what can’t be memorised: markets that haven’t happened yet, and worlds with physics they have never seen.
Live and matured portfolios averaged. No horizon has matured yet, so these ratings are provisional.
Reference engineer
1500anchor
5 rated worlds, each scored on 10 courses the models never saw.
How it works
Full method →- Frozen inputsEach model gets the same snapshot: prices up to a closing date, or a world’s rules without its physics. Nothing newer.
- Committed answersModels commit to portfolios, forecasts or controllers before anything is known. We store every answer unchanged.
- The world gradesMarkets move over months; rovers drive unseen courses. Ratings come from those results, always with an interval.
Insights
All insights →- When a model thinks past its budgetIn the 6 October sandbox world, MiMo V2.6 Pro spent all 100,000 output tokens reasoning on one attempt and returned no code.
- Correcting our ratingsA bug in how we fitted Elo ratings put DeepSeek Flash first in Markets. It was last. Here is what went wrong and what we changed.