✨ Data unavailable
The benchmark data could not be loaded.
Loading class relationships…
✨ How it works
The benchmark contains 50 complexity classes and all 2,500 ordered pairs of those classes. Each pair asks whether A ⊆ B: whether every decision problem in A also belongs to B. The reverse direction is a separate question.
The baseline records 1,122 relations known by September 1, 2026. The remaining 1,378 questions were open as of September 1, 2026. A resolution can prove containment, prove noncontainment, or establish that the statement is independent of ZFC.
Accepted proofs are combined with the baseline using the benchmark’s fixed implication rules. Every newly resolved open pair earns one point, including pairs settled as consequences.
For example, if an AI proves that P ⊊ NP but no other results, it would earn 103 points (a score of 7.47%), as this would also imply P ⊊ PSPACE and settle 101 other open questions.
Score = (newly resolved pairs / 1,378) × 100%.
Leaderboard
| Rank | Model | Score |
|---|---|---|
| ? | OAI Internal Model | ? |
| 1 | Fable 5.1 | 0.00% |
| 1 | GPT-6 Astra | 0.00% |
✨ Class relationships
Click a class to select A, then B.
“Open” means no resolution was found in the cutoff audit. Earlier results missed by that audit can lead to corrections.
✨ Scope and sources
The questions concern classes of total decision languages, with fixed definitions and circuit-uniformity conventions. Results about promise problems, oracle-relative separations, search problems, or a particular algorithm’s speed do not automatically resolve a pair in this benchmark.
New containment and noncontainment claims require Lean proof checking and mathematical review. Independence claims require a separate expert review. Existing cited results serve as premises. The public question set and baseline do not establish that a model discovered a result independently or had no prior exposure to it.
Evaluation details · Cutoff audit · Coverage notes
References
—