Page 12 of the deck says no AI is best at everything and the leader keeps changing. This is the table behind that, with nothing trimmed for space, plus how it is scored and where it comes from.
| Engine | Tier | Crypto | Biohack | Current events | Investing | Cybersec | Medical | Science | Legal |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable | Pro | 78 | 85 | 70 | 88 | 90 | 92 | 95 | 91 |
| Claude Opus | Pro | 63 | 82 | 68 | 86 | 88 | 90 | 93 | 89 |
| OpenAI o3 | Pro | 82 | 84 | 72 | 90 | 94 | 93 | 96 | 90 |
| GLM | Pro | 70 | 72 | 78 | 75 | 76 | 78 | 82 | 79 |
| GPT-5.6 Sol | Pro | 80 | 83 | 82 | 88 | 91 | 91 | 94 | 88 |
| DeepSeek R1 | Pro | 80 | 77 | 60 | 88 | 88 | 81 | 94 | 79 |
| Claude Sonnet | Plus | 75 | 80 | 68 | 82 | 86 | 86 | 90 | 85 |
| Claude Haiku | Plus | 62 | 65 | 60 | 68 | 70 | 72 | 76 | 72 |
| GPT-4o | Plus | 72 | 75 | 80 | 80 | 82 | 84 | 86 | 82 |
| Gemini Pro | Plus | 44 | 78 | 90 | 0* | 84 | 88 | 92 | 84 |
| Kimi | Plus | 66 | 68 | 78 | 74 | 72 | 75 | 80 | 80 |
| Qwen Max | Plus | 72 | 74 | 80 | 78 | 80 | 82 | 88 | 81 |
| DeepSeek | Plus | 76 | 75 | 65 | 82 | 84 | 80 | 91 | 78 |
| GPT-OSS | Free | 50 | 52 | 45 | 55 | 58 | 56 | 62 | 54 |
| Gemini Flash Lite | Free | 48 | 60 | 82 | 45 | 62 | 68 | 72 | 60 |
| Llama 3.1 8B | Free | 54 | 55 | 48 | 58 | 62 | 60 | 68 | 58 |
| Grok | Staged | 75 | 74 | 92 | 80 | 80 | 78 | 86 | 76 |
| Mistral | Staged | 70 | 72 | 65 | 76 | 82 | 78 | 85 | 84 |
| Command R+ | Staged | 60 | 62 | 75 | 78 | 68 | 70 | 72 | 82 |
A score is that engine's strength on that subject, out of 100. It is not a rating of the company and it is not an overall ranking. The whole point of the table is that the same engine can be excellent at one subject and poor at the one next to it, so an average across the row would throw away the only useful information in it.
*Gemini's 0 on investing is a real result, not a data-entry error. That model is built to refuse investment questions, so it scores 0 on a subject it will not answer. We left it in rather than tidy it away, because it is a good example of the thing we are measuring: a refusal is a wrong answer to the person who asked.
This is our best current read of the landscape from public tests, July 2026. It is not our own lab measurement yet, and we would rather say that plainly than let a coloured grid imply more than it should.
Nothing on this page is a claim about any company. It is a working record of which model we would send a given question to today, and it changes. Full sourcing available on request.