ReallySolved · Deck appendix
← Back to the deck
Deck appendix · page 12 of the deck, in full

Which AI is best at what, the whole board.

Page 12 of the deck says no AI is best at everything and the leader keeps changing. This is the table behind that, with nothing trimmed for space, plus how it is scored and where it comes from.

EngineTierCryptoBiohackCurrent eventsInvestingCybersecMedicalScienceLegal
Claude FablePro7885708890929591
Claude OpusPro6382688688909389
OpenAI o3Pro8284729094939690
GLMPro7072787576788279
GPT-5.6 SolPro8083828891919488
DeepSeek R1Pro8077608888819479
Claude SonnetPlus7580688286869085
Claude HaikuPlus6265606870727672
GPT-4oPlus7275808082848682
Gemini ProPlus4478900*84889284
KimiPlus6668787472758080
Qwen MaxPlus7274807880828881
DeepSeekPlus7675658284809178
GPT-OSSFree5052455558566254
Gemini Flash LiteFree4860824562687260
Llama 3.1 8BFree5455485862606858
GrokStaged7574928080788676
MistralStaged7072657682788584
Command R+Staged6062757868707282
91-100 expert 76-90 strong 51-75 competent 21-50 weak 0-20 unreliable

What the tier column means

How to read a score

A score is that engine's strength on that subject, out of 100. It is not a rating of the company and it is not an overall ranking. The whole point of the table is that the same engine can be excellent at one subject and poor at the one next to it, so an average across the row would throw away the only useful information in it.

*Gemini's 0 on investing is a real result, not a data-entry error. That model is built to refuse investment questions, so it scores 0 on a subject it will not answer. We left it in rather than tidy it away, because it is a good example of the thing we are measuring: a refusal is a wrong answer to the person who asked.

Where these numbers come from, and what they are not

This is our best current read of the landscape from public tests, July 2026. It is not our own lab measurement yet, and we would rather say that plainly than let a coloured grid imply more than it should.

Why this is the moat and not just a feature. The scoreboard is 1 asset doing 2 jobs. It decides where to send the next question, which is what makes the answers good and the cost low. And the corrections that keep it current are the same corrections the AI labs buy. Nobody has to be persuaded to maintain it: it is a byproduct of people using the product.

Nothing on this page is a claim about any company. It is a working record of which model we would send a given question to today, and it changes. Full sourcing available on request.