Page 1 says we show you the tests that were not flattering. This is that page. Every accuracy test we have run is here, in the order we ran it, with what we expected before we started and what came back. 4 of them went against us. Those are marked, not buried.
| What we tested | What we expected | What we found | Our way? |
|---|---|---|---|
| Does combining a few cheap AIs match one expensive one? 62 reasoning questions the system had never seen |
Close, but behind | 95% against the best model's 97%, at 1/40th the cost. Both got 59 right; the gap is the denominator | Yes |
| Are more AIs better? 10 engines at once against 3 |
More would win | 10 scored 11 points worse than 3 and cost 18 times as much. Picking the right few is the skill | No, and we changed |
| Does clever combining beat simple combining? | Clever would win | Simple won, so simple is what ships. The intelligence belongs in choosing who to ask, not in the merge | No, and we changed |
| Was our own question set too easy? | It was fine | It was too easy. On questions written by outside researchers the same engines did worse, some a lot worse. We stopped using our own numbers the day we found out | No |
| How far apart are the best and worst AI on the same questions? Public trick-question set, 44 to 60 per engine |
A real gap | 86 right out of 100 against 50. A 36 point spread, and the only difference is which AI you asked | Yes |
| Was the grading of that run correct? | Yes | No. It was wrong in one direction, punishing engines that write at length. Every score moved, the ranking changed, and several conclusions we had already drawn did not survive. We retracted them | No |
| Had the engines simply memorised those questions? | Hopefully not | They had not. On questions written by hand that morning, which they cannot have seen, every engine did better, which is the opposite of what memorising produces | Yes |
| Does the gap hold on a second, different set? 31 questions, one day |
A similar gap | Wider. 77, 58 and 10 right out of 100. 67 points top to bottom, 19 between the 2 leading models | Yes |
| Do cheap AIs go out of date before expensive ones? 60 facts whose answer changed on a known date, plus 20 that did not |
Cheap ones would be far worse | Wrong question. Price is not what predicts it. Three top-tier paid models scored 63%, 12% and 2%. One of them got 83% of unchanged facts right and 2% of changed ones: the model is fine, its knowledge is frozen | Not as asked |
| Does combining AIs fix out-of-date answers? Same 60 facts |
It should, because their cut-off dates differ | It does not. Combining the 3 free engines fixed almost none of them. When they are all out of date they agree with each other. Asking the right engine fixed most of them | No |
| Was our grading of that run correct? | Yes | No, again. A cheap marker mislabelled roughly a third. We caught it by hand-checking, re-marked with 2 independent markers, and the first set of numbers is retracted | No |
| How do AIs do on ordinary everyday questions? 54 questions people really ask, health, money, travel, law, cooking, home safety |
Worse than on hard sets | Better. Far better. 38 to 53 right out of 54, free engines included. A free engine beat a premium one | Yes, and it changed the story |
Read the last 3 rows together, because that is the finding. On ordinary questions today's AIs are good, cheap ones included. On facts whose answer has changed they collapse, expensive ones included, and combining them does not rescue it. The failure is not that AI is unreliable. It is that AI is reliable right up until the world moves, and it cannot tell you which of the two you are getting.
The 2 grading failures above are the reason this page exists. Both were found by us, in our own work, and both cost us a claim we liked. We would rather show you a page with 4 losses on it than a page where every test happened to agree with us. Full workings, question sets and raw answers available on request.
What these tests do NOT cover, so you do not have to ask. More than a third of what people ask an AI has no right answer at all: rewriting an email, making a picture, brainstorming. Nobody can grade that, including us. Everything above is the part where being wrong costs money or sends somebody to hospital. Sample sizes are 31 to 62 questions, not hundreds, and the everyday set is UK-heavy. Every number here sits next to the size of the set it came from.