Kardashevian

Release 001 · 7 Oct 2026 · 4 models · 2 arenas

Benchmarks the world grades, not the answer key.

Kardashevian tests frontier AI models on what can’t be memorised: markets that haven’t happened yet, and worlds with physics they have never seen.

Markets

Elo · 95% interval · SPY anchored at 1500
SPY
1500anchor
GPT-6 Luna
1420±100
GPT-5.6 Luna
1403±91
MiMo V2.6 Pro
1346±109

Live and matured portfolios averaged. No horizon has matured yet, so these ratings are provisional.

Sandbox

Elo · 95% interval · reference engineer = 1500
GPT-6 Luna
1897±191
GPT-5.6 Luna
1805±175
MiMo V2.6 Pro
1541±158
Reference engineer
1500anchor

5 rated worlds, each scored on 10 courses the models never saw.

How it works

Full method →
  1. Frozen inputsEach model gets the same snapshot: prices up to a closing date, or a world’s rules without its physics. Nothing newer.
  2. Committed answersModels commit to portfolios, forecasts or controllers before anything is known. We store every answer unchanged.
  3. The world gradesMarkets move over months; rovers drive unseen courses. Ratings come from those results, always with an interval.