ReallySolved · Deck appendix
← Back to the deck
Deck appendix · the 203 TOPICS Accuracy Matrix, in full

The 203 TOPICS Accuracy Matrix, & the 3,192 cells still blank.

203 topics down, 19 engines across.

This Matrix is a synthesis of current research + our own testing data. It shows how important our strategy to use MASS CROWDSOURCING is to ACHIEVING EVEN HIGHER ACCURACY RATES at even lower costs. AI safety is the subject we have not highlighted in this deck only because we have focused on accuracy as an immediate salable feature, but safety is also a very important subject that we can improve as well through Mass crowdsourcing.

It is mostly blank, and the blank part is the point. Every dash is a cell no benchmark on Earth reports and no lab publishes.

3,857cells in the matrix: 203 topics × 19 engines
665that carry a value today
3,192showing a dash, because nobody has filled them
17%of the matrix, from every public benchmark there is

Where the numbers come from, & what they are not

Two sources were combined here, & they are the same kind of thing, which is the only reason combining them is honest.

A number in this table is not a per-topic measurement. It is our July 2026 read of that engine on the broad board subject the topic sits under, carried across & shown for every topic under it. So all 4 finance topics show the same figure, because the finer number does not exist yet. A dash means nobody has filled that cell, & that is on purpose. We would rather show you a gap than fill it with something that looks like data.

The first thing we asked: does anyone already rank AI at this depth?

No, & not close. That was the clearest thing to come back, & it is worth being precise about, because “nobody has this” is a claim that deserves checking.

BenchmarkWhat it actually coversDepth
MMLUMultiple-choice questions across academic subjects57 subjects
MMLU-ProThe harder rebuild of it, fewer & broader groupings14 categories
GPQAGraduate-level science questions written to defeat web search3 fields: biology, chemistry, physics
SWE-benchFixing real issues in real software repositories1 task type
HumanEvalWriting small programs to a specification1 task type
HELMA framework for evaluating on many axes at once, accuracy being 1Scenarios, not subjects
MedQA / USMLEMedical licensing exam questions1 profession
LegalBenchLegal reasoning tasks1 profession
ArenaWhich answer a person preferred, head to headPreference, not correctness

The deepest of them is 57 subjects, & it is the oldest & the most saturated. The modern ones went the other way & got narrower & harder. Not 1 of them reports a per-topic score for anything like 200 topics, & none of them is refreshed when the world changes underneath an answer.

Where the 2 outside models disagreed, & which one we followed

One of them gave frontier models 90 to 95 on nearly every topic. The other objected, using the current SWE-bench standings as the argument: 76.8, 75.8 & 69.6 across 3 frontier models on the same task. Models are not flat, & a matrix that says 92 in every cell has the same defect as scoring people on “science”: there is nothing left to choose between.

We took the lower read everywhere they disagreed. A matrix exists to tell 2 engines apart. If it cannot do that, it is decoration.

The finding we did not expect

Three of our 8 board subjects did not exist anywhere in a 200-topic map of human knowledge. Not as a near-match, not under another name. They are:

All 3 now have a row, which fixed our data & did not make the finding go away. A map of university subjects had no cell for 3 of the things people ask us about most, & that is worth more as evidence than any sentence we could write ourselves.

This is the deck’s own argument arriving from outside & unprompted: the public benchmarks score exam subjects, not what people type. A taxonomy built from university departments has no cell for the 3 things people ask us about most.

Cybersecurity is a 4th, softer case. The list has Crypto Math & Information Theory but no Cybersecurity, so the board value is carried across as an approximation & is marked as one in the data rather than presented as a match.

How the empty cells get filled

No cell will ever be filled by estimate. A guess & a measurement look identical once they are in a grid, & the difference between them is the entire product.

The matrix

203 topics down, 19 engines across. Scroll it sideways for the engines & down for the topics.

91-100 expert 76-90 strong 51-75 competent 21-50 weak 0-20 unreliable  nobody has measured this
#TopicCovered byClaude FableClaude OpusOpenAI o3GLMGPT-5.6 SolDeepSeek R1Claude SonnetClaude HaikuGPT-4oGemini ProKimiQwen MaxDeepSeekGPT-OSSGemini Flash LiteLlama 3.1 8BGrokMistralCommand R+
Formal & Abstract Sciences
1Formal LogicMMLU-Pro, LogiQA
2Informal FallaciesMMLU-Pro, LogiQA
3Mathematical LogicMMLU-Pro, LogiQA
4Modal LogicMMLU-Pro, LogiQA
5Proof TheoryMMLU-Pro, LogiQA
6Number TheoryFrontierMath, MATH
7Abstract AlgebraFrontierMath, MATH
8GeometryFrontierMath, MATH
9TopologyFrontierMath, MATH
10Mathematical AnalysisFrontierMath, MATH
11CalculusAIME, MATH
12Linear AlgebraAIME, MATH
13Differential EquationsAIME, MATH
14Numerical AnalysisAIME, MATH
15Game TheoryAIME, MATH
16Descriptive StatisticsMATH, GPQA
17Inferential StatisticsMATH, GPQA
18Bayesian StatisticsMATH, GPQA
19Stochastic ProcessesMATH, GPQA
20Probability TheoryMATH, GPQA
21AlgorithmsLiveCodeBench, HumanEval
22Data StructuresLiveCodeBench, HumanEval
23Computational ComplexityLiveCodeBench, HumanEval
24Crypto MathLiveCodeBench, HumanEval90889476918886708284728084586262808268
25Information TheoryLiveCodeBench, HumanEval90889476918886708284728084586262808268
Physical & Material Sciences
26Classical MechanicsGPQA Diamond, MMLU-Pro95939682949490768692808891627268868572
27ElectromagnetismGPQA Diamond, MMLU-Pro95939682949490768692808891627268868572
28ThermodynamicsGPQA Diamond, MMLU-Pro95939682949490768692808891627268868572
29Quantum MechanicsGPQA Diamond, MMLU-Pro95939682949490768692808891627268868572
30RelativityGPQA Diamond, MMLU-Pro95939682949490768692808891627268868572
31OpticsGPQA Diamond, MMLU-Pro
32AcousticsGPQA Diamond, MMLU-Pro
33Particle PhysicsGPQA Diamond, MMLU-Pro
34AstrophysicsGPQA Diamond, MMLU-Pro
35Plasma PhysicsGPQA Diamond, MMLU-Pro
36Organic ChemistryGPQA, ChemBench95939682949490768692808891627268868572
37Inorganic ChemistryGPQA, ChemBench95939682949490768692808891627268868572
38Physical ChemistryGPQA, ChemBench
39Analytical ChemistryGPQA, ChemBench
40BiochemistryGPQA, ChemBench95939682949490768692808891627268868572
41Polymer ChemistryGPQA, ChemBench
42ElectrochemistryGPQA, ChemBench
43ThermochemistryGPQA, ChemBench
44Quantum ChemistryGPQA, ChemBench
45Environmental ChemistryGPQA, ChemBench
46GeologyMMLU
47MeteorologyMMLU
48OceanographyMMLU
49HydrologyMMLU
50VolcanologyMMLU
51SeismologyMMLU
52ClimatologyMMLU
53GeomorphologyMMLU
54MineralogyMMLU
55PaleontologyMMLU
56Planetary ScienceMMLU, GPQA
57CosmologyMMLU, GPQA
58Stellar AstronomyMMLU, GPQA
59Galactic AstronomyMMLU, GPQA
60AstrobiologyMMLU, GPQA
Life & Biological Sciences
61Cell BiologyGPQA Biology, MMLU-Pro95939682949490768692808891627268868572
62Molecular GeneticsGPQA Biology, MMLU-Pro95939682949490768692808891627268868572
63EpigeneticsGPQA Biology, MMLU-Pro
64GenomicsGPQA Biology, MMLU-Pro95939682949490768692808891627268868572
65ProteomicsGPQA Biology, MMLU-Pro
66AnatomyMedQA, MMLU
67PhysiologyMedQA, MMLU
68Developmental BiologyMedQA, MMLU
69Comparative BiologyMedQA, MMLU
70Evolutionary BiologyMedQA, MMLU95939682949490768692808891627268868572
71Ecosystem EcologyMMLU
72Population DynamicsMMLU
73Conservation BiologyMMLU
74BiodiversityMMLU
75BiogeographyMMLU
76BacteriologyPubMedQA, MedQA
77VirologyPubMedQA, MedQA95939682949490768692808891627268868572
78MycologyPubMedQA, MedQA
79ParasitologyPubMedQA, MedQA
80ImmunologyPubMedQA, MedQA95939682949490768692808891627268868572
81BotanyMMLU, ScienceQA
82ZoologyMMLU, ScienceQA
83EntomologyMMLU, ScienceQA
84Marine BiologyMMLU, ScienceQA
85NeuroscienceMMLU, ScienceQA95939682949490768692808891627268868572
Applied Sciences, Engineering & Medicine
86Mechanical EngineeringSWE-bench, MMLU Engineering
87Electrical EngineeringSWE-bench, MMLU Engineering
88Civil EngineeringSWE-bench, MMLU Engineering
89Chemical EngineeringSWE-bench, MMLU Engineering
90Aerospace EngineeringSWE-bench, MMLU Engineering
91Biomedical EngineeringSWE-bench, MMLU Engineering
92Environmental EngineeringSWE-bench, MMLU Engineering
93Software EngineeringSWE-bench, MMLU Engineering
94RoboticsSWE-bench, MMLU Engineering
95Materials ScienceSWE-bench, MMLU Engineering
96Clinical MedicineMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
97PharmacologyMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
98EpidemiologyMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
99Public HealthMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
100Nutrition ScienceMedQA, USMLE, MedMCQA
101PathologyMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
102ToxicologyMedQA, USMLE, MedMCQA
103SurgeryMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
104PsychiatryMedQA, USMLE, MedMCQA92909378918186728488758280566860787870
105KinesiologyMedQA, USMLE, MedMCQA
106AI Safety & MisuseGPQA, MMLU
107ML Security & RobustnessGPQA, MMLU
108NanotechnologyGPQA, MMLU
109Quantum ComputingGPQA, MMLU
110Synthetic BiologyGPQA, MMLU
Social Sciences & Human Behavior
111Cognitive PsychologyMMLU Psychology
112Developmental PsychologyMMLU Psychology
113Social PsychologyMMLU Psychology
114Clinical PsychologyMMLU Psychology
115Behavioral EconomicsMMLU Psychology
116NeuropsychologyMMLU Psychology
117Industrial-Organizational PsychologyMMLU Psychology
118Cultural AnthropologyMMLU
119Physical AnthropologyMMLU
120ArcheologyMMLU
121Social StratificationMMLU
122Urban SociologyMMLU
123DemographyMMLU
124Sociology of ReligionMMLU
125MicroeconomicsMMLU-Pro Economics, FinQA
126MacroeconomicsMMLU-Pro Economics, FinQA
127EconometricsMMLU-Pro Economics, FinQA
128Development EconomicsMMLU-Pro Economics, FinQA
129Behavioral EconomicsMMLU-Pro Economics, FinQA
130International TradeMMLU-Pro Economics, FinQA
131Political TheoryLegalBench, Uniform Bar Exam
132Comparative PoliticsLegalBench, Uniform Bar Exam
133International RelationsLegalBench, Uniform Bar Exam
134Public PolicyLegalBench, Uniform Bar Exam
135Constitutional LawLegalBench, Uniform Bar Exam91899079887985728284808178546058768482
136Criminal LawLegalBench, Uniform Bar Exam91899079887985728284808178546058768482
137International LawLegalBench, Uniform Bar Exam91899079887985728284808178546058768482
138JurisprudenceLegalBench, Uniform Bar Exam91899079887985728284808178546058768482
Humanities, Culture & Arts
139EpistemologyMMLU Philosophy, ETHICS
140MetaphysicsMMLU Philosophy, ETHICS
141EthicsMMLU Philosophy, ETHICS
142AestheticsMMLU Philosophy, ETHICS
143Political PhilosophyMMLU Philosophy, ETHICS
144Philosophy of MindMMLU Philosophy, ETHICS
145Philosophy of ScienceMMLU Philosophy, ETHICS
146Ancient HistoryMMLU History
147Medieval HistoryMMLU History
148Modern HistoryMMLU History
149World HistoryMMLU History
150Military HistoryMMLU History
151HistoriographyMMLU History
152Intellectual HistoryMMLU History
153PhoneticsFLORES, WMT, MMLU
154PhonologyFLORES, WMT, MMLU
155SyntaxFLORES, WMT, MMLU
156SemanticsFLORES, WMT, MMLU
157PragmaticsFLORES, WMT, MMLU
158SociolinguisticsFLORES, WMT, MMLU
159Historical LinguisticsFLORES, WMT, MMLU
160Translation StudiesFLORES, WMT, MMLU
161World LiteratureMMLU, MM-Vet
162Literary CriticismMMLU, MM-Vet
163Art HistoryMMLU, MM-Vet
164Music TheoryMMLU, MM-Vet
165MusicologyMMLU, MM-Vet
166Theatre & PerformanceMMLU, MM-Vet
167Film StudiesMMLU, MM-Vet
168ArchitectureMMLU, MM-Vet
169Visual DesignMMLU, MM-Vet
170Comparative ReligionMMLU Religion
171Systematic TheologyMMLU Religion
172MythologyMMLU Religion
173Religious HistoryMMLU Religion
Business, Finance & Industry
174Strategic ManagementMMLU Management
175Operations ManagementMMLU Management
176Organizational BehaviorMMLU Management
177Human Resource ManagementMMLU Management
178EntrepreneurshipMMLU Management
179Supply Chain ManagementMMLU Management
180Corporate FinanceFinQA, CPA-style8886907588888268800747882554558807678
181Investment ManagementFinQA, CPA-style8886907588888268800747882554558807678
182Financial MarketsFinQA, CPA-style8886907588888268800747882554558807678
183Personal FinanceFinQA, CPA-style8886907588888268800747882554558807678
184Financial AccountingFinQA, CPA-style
185Managerial AccountingFinQA, CPA-style
186AuditingFinQA, CPA-style
187Crypto MktFinQA, CPA-style78638270808075627244667276504854757060
188MarketingFinQA, MMLU Marketing
189Consumer BehaviorFinQA, MMLU Marketing
190Real Estate InvestmentFinQA, MMLU Marketing
191International BusinessFinQA, MMLU Marketing
Practical, Experiential & Everyday Knowledge
192Construction & Carpentrynone that we could find
193Electrical Wiringnone that we could find
194Plumbingnone that we could find
195Automotive Repairnone that we could find
196Agriculture & Farmingnone that we could find
197Culinary ArtsMedQA (first aid only)
198Gardening & HorticultureMedQA (first aid only)
199First Aid & SafetyMedQA (first aid only)
200Fitness & Physical TrainingMedQA (first aid only)
201Navigation & Outdoor SurvivalMedQA (first aid only)
202Current Eventsnone that we could find70687278826068608090788065458248926575
203Biohackingnone that we could find85828472837780657578687475526055747262

Built by tools/build-master-matrix.mjs from data/topic-matrix-200.json & the July 2026 board. Machine-readable version: data/topic-matrix-master.json. Regenerate rather than hand-editing a cell.

What this page is not

The benchmarks, with their sources

Every name below is a real, published benchmark & the links go to the paper or the code. A cell in the table above is only as good as the benchmark behind it, so the benchmarks get named & linked rather than waved at.

BenchmarkWhat it evaluatesSource
MMLU57 subjects across humanities, social sciences, STEM and other areas.arxiv.org · github.com
MMLU-ProHarder rebuild of MMLU, broader reasoning.arxiv.org · github.com
HELMStanford's framework: many scenarios and dimensions rather than 1 capability score.crfm.stanford.edu · arxiv.org
GPQAGraduate-level questions written to defeat web search.arxiv.org · github.com
GPQA DiamondThe hardest subset, the one used in frontier-model comparisons.arxiv.org
MATHMeasuring mathematical problem solving.arxiv.org · github.com
AIMEAmerican Invitational Mathematics Examination.artofproblemsolving.com
FrontierMathEpoch AI's hard mathematics set.epoch.ai · arxiv.org
LeanFormal theorem proving.leanprover-community.github.io
HumanEvalWriting small programs to a specification.arxiv.org · github.com
SWE-benchFixing real issues in real repositories.arxiv.org · github.com · www.swebench.com
LiveCodeBenchContamination-resistant coding evaluation.arxiv.org · github.com
FLORESMultilingual machine-translation evaluation.arxiv.org · github.com
WMTConference on Machine Translation.www2.statmt.org
MedQAMedical licensing exam questions.arxiv.org · github.com
MedMCQAMedical entrance-exam questions.arxiv.org · github.com
PubMedQABiomedical research question answering.arxiv.org · github.com
LegalBenchLegal reasoning tasks.arxiv.org · github.com
MMLU Law categoryThe law slice of MMLU.github.com
FinQAFinancial numerical reasoning.arxiv.org · github.com
FinanceBenchFinancial question answering.arxiv.org
CyBenchCybersecurity tasks.arxiv.org · github.com
MM-VetMultimodal capability evaluation.arxiv.org · github.com
ScienceQAMultimodal science question answering.arxiv.org · github.com
LMSYS Chatbot ArenaWhich answer a person preferred, head to head.chat.lmsys.org · arxiv.org · github.com
Epoch AICapability and benchmark tracking.epoch.ai · epoch.ai
BenchLMBenchmark tracking.benchlm.ai
Papers With CodeBenchmarks and datasets index.paperswithcode.com
Hugging Face Open LLM LeaderboardOpen-model leaderboard.huggingface.co
A benchmark name is not evidence. 6 names in the table above carry a dagger because we could not tie them to a specific published benchmark: CPA-style, ChemBench, ETHICS, LogiQA, USMLE, Uniform Bar Exam. Some are category shorthand, some are exam names rather than benchmarks. They are labelled rather than quietly dropped, & no URL sits beside any of them. A name mistaken for a citation is the exact failure this company sells against.
Names we were handed & did NOT use at all: Physics QA, AstroML, ChemBench, BioQA, MatSciBench, AutoSizing, CircuitBench, RoboBench, AI-Bench, EconBench, TradeQA, PlantQA, ChefBench, AgBench, SportsQA, Financial NLP Bench. These appeared in an earlier draft matrix as shorthand for a KIND of evaluation rather than as real benchmarks. They are recorded here so nobody re-adds them believing they were checked.

The shape every row eventually needs

A score with no provenance cannot be argued with, so the research version carries all of this per row:

SUBTOPIC × MODEL × SCORE × BENCHMARK × BENCHMARK URL × BENCHMARK DATE × MODEL VERSION × TEST DATE × SAMPLE SIZE × GROUND TRUTH × SOURCE QUALITY × CONFIDENCE

It prevents a benchmark name from being mistaken for evidence when it is actually only a category or an informal label.

The safety half of the same idea

Accuracy is the half we sell first. Safety is the second thing the same crowd improves, & it has its own matrix: 245 safety topics across 22 categories, with the frontier rows given ranges rather than invented precision. The AI Safety, Security & Online Harm Matrix →

If any of this does not hold up, we would rather hear it than defend it. The company exists because confident numbers are not the same as correct ones.