For investors, growth partners & technical reviewers

18 systems test AI accuracy, running right now

We do not ask an AI 1 question & trust the answer. We run 18 separate systems against the AIs we test, each one built to catch a different way an AI can go wrong.

1 score cannot hold everything
An AI can be strong at one thing & weak at the next.

It can be confident when it is wrong, right for the wrong reason, or current on 1 subject & out of date on another. A single accuracy number hides all of that. So we do not use a single number. We built 18 separate systems, each aimed at 1 specific way an AI can fail, & each one keeps running.

iThis page names what each system checks, not what any of them found. Results, thresholds & the logic behind any decision are not published here.

The 18 systems
Each one checks a different way an AI can go wrong.
01

Routing

Every question could go to any AI. Send it to the wrong one & the answer suffers for no reason. This finds which way of assigning a question actually works, topic by topic.

02

Advisor honesty

Our own in-product help is only useful if it tells the truth about what the product does. This checks it never invents a feature that does not exist, or denies one that does.

03

Recency check

Some questions have answers that go out of date. This catches a question like that before it is even sent to an AI, so a stale answer never gets a fair chance to sound current.

04

Out-of-date answers

An AI can sound sure & still be wrong because the world moved on after it learned. This checks whether an AI's answer is current, out of date, or simply made up, subject by subject, on facts with a known date they changed.

05

Per-AI recency

Different AIs learned the world at different moments. This measures how current each one's own knowledge actually is, directly, rather than guessed at.

06

Advice quality

Real people ask AIs for advice on real decisions. This grades that advice against an outside, already-published benchmark built for exactly that job.

07

Safety coverage

People run into scams, bad suggestions & risky calls in ordinary conversation, not only in obviously dangerous questions. This checks coverage across many kinds of harm & misuse an ordinary person might actually hit.

08

Topic competence

No AI is equally strong at everything. This measures which AI is actually good at which subject, across many subjects, so a guess is never the reason one gets picked.

09

Referee prompting

When several AIs disagree, something has to weigh their answers & decide. This tests different ways of asking for that judgment, to find which produces the most trustworthy verdict.

10

AI reliability

An AI does not always answer. This tracks whether it answers at all, & what a failure actually looks like when it does not.

11

Change-alarm accuracy

A system built to notice when a stored answer has gone out of date is only useful if it actually notices. This tests whether that alarm can be trusted, rather than assuming it works.

12

Verdict cost

Combining several AI opinions into 1 final answer costs more than just asking the AIs themselves. This measures the true cost of producing that final answer, not just the price of the questions that went into it.

13

Behavior under pressure

An honest AI admits when it does not know. A weak one changes its answer the moment it is challenged, whether it was right or wrong. This tests which one an AI actually does.

14

Judging quality

Grading an AI's answer as right or wrong is its own hard problem. This measures how reliably an AI can be trusted to judge another AI's work.

15

Effort & retries

Giving an AI more time or a second try sounds like it should help. This measures whether it actually fixes a wrong answer, & what that extra effort costs.

16

Answer length & cost

A longer answer usually costs more to produce. This checks whether what an AI actually charges matches what a simple estimate would assume, or quietly runs over it.

17

Silent agreement failure

The hardest case for ordinary comparison to catch is several AIs agreeing with each other while all being wrong. This system exists specifically to catch that.

18

Cost per correct answer

Cost per question asked is the wrong number to watch. This tracks cost per question actually answered right, which is the number that matters.

What is not on this page
No scores. No rankings. No routing rules.

This page names what we test. Which AI wins on which subject, the thresholds involved & how any of it turns into a routing decision are not published here. That is by design, & it is what keeps the work worth having.

18 systems, on the AIs we test. That is the part worth knowing before anything else.

Diligence & partnership questions

Want more detail than fits on a page? Contact us.

We can walk through any of these 18 systems in more depth, in writing, for an investor, growth partner or technical reviewer doing diligence.

Contact