Within 2 points of Claude Opus 5 in accuracy, at 1/40th the cost.i95% against its 97%, on the same questions.
Wrong answers on $M decisions are expensive. We fix your questionsiExplained below in the YOU ASK section.& route them to 1 or multiple AIs that we know are best in 1 of 448 topicsiCreated originally from market research, then by internal testing, then updated in real time by our user-created, platform-owned, topic accuracy data. We then run a proprietary blend of math + statistical analysis across their answers to resolve most of the disagreements. Fewer tokens, higher accuracy.
AI is moving faster than its creators & fixers can control - Mass Multi-modal crowdsourcing, is the only fix.iWe want to fix all kinds of problems. RAG Has Context Limits, Model Collapse exists in RLAIF, Synthetic Data "Blind Spots”, Sybil & LLM-Farming Spoofing (Annotators frequently use underlying LLMs), Synthetic Data "Blind Spots”, Poisoning & Collusion in Decentralized Attestation (Annotator rings can collude to gaming the attestation logic), Data Provenance and "Licensing Laundering”(Massive dataset scraping and multi-stage transformations break the chain of custody.) Watermarking attestations are extremely fragile., Consensus & Inter-Annotator Disagreement (queries are deeply subjective or complex), etc.
6 more
Humans enjoy playing an unprecedented 9 self-reinforcing viral games that recruit hobbyists up to attestation pros to correct what's still not 100% resolved & get paid for what only frustrates them today.iTop 3 AI models said no other company in the world has more than 3 self-reinforcing flywheel loops, We run at least 9! (then again...our whole premise is you can't trust what AI tells you until a human checks it)
Excluded from our forecast to be ultra conservative: $100M's expected from AI labs buying human-corrected data.i~$30B spent on human data so far in 2026. Scale AI was selling it at $2B/yr before Meta investment, Mercor $614M in 6 months Current solutions pay by the hour. We GET PAID by providing very low cost high quality AI plans, while allowing hobbyists & professionals, a way to effortlessly monetize their specialized knowledge as instant consultants & an outlet for their frustration when they type a short correction, on the spot & at the moment they see a mistake, & get paid their share of human corrected data they would never be able to sell.
AI answers go stale everyday, & AI doesn't tell you when it's out of date.
With the AI staff we provide, we make it easier for members to do things with the knowledge base they build &/or import like; record, find, sell, promote, organize, remember, post, send, & more.
Even expensive AI models are confidently wrong, I (Bob, founder) learned this the hard way many times. AI is amazing, but when you don't know which answer is really correct, you are still really guessing. We fix that in several ways.
People don't want to read a book every time AI answers. We summarize the answers & point to the disagreements so members can quickly see where they may have been misled.
Low risk
VC founder, fully funded.
Patent pending, reviewed by a top 10 patent firm.
7 more
92-96% margins.
The AI labs already pay people by the hour for this work.iWidely reported figures, not our own measurement: Scale was said to sell about $2B a year of it & Mercor about $614M in 6 months.Ours arrives free or as revenue, as a by-product of people using the product (depends on plan).
We turn everyday people, who never realized their oddly granular knowledge of trivial things is valuable, into instant consultants. Their credentials come from paying customers' feedback, not self-posted & buddy-endorsed claims.
Built, running & tested today. Not a plan.
History proves the consumer angle is less risky than trying to capture a few large companies (counterintuitive)iHistory has proven a buyer you depend on can be taken away by somebody else's regulator, somebody else's election, or somebody else's change of mind. Consumer subscriptions are the only money in this category that no single institution can withdraw.
A studio's first instinct is 3 big logos. Every company that tried this before us died when 1 of those logos walked, so we build on many small payers first.
4 companies in this exact business found out the hard way. Logically lost 2 platform contracts, went into administration in July 2025 & was sold for parts. Full Fact lost 1 contract worth over Β£1M a year & cut about a quarter of its staff. Factmata sold cheaply & was folded into somebody else's product. NewsGuard won every legal fight it faced, & still lost most of its business. A $13M defamation suit against it was dismissed. It sued the FTC over a documents demand & the FTC withdrew. A congressional inquiry produced nothing. Meanwhile 2 of its 3 kinds of buyer went away anyway: government contracts gone, & regulators barred the big ad groups from buying ratings like its own, as a condition of somebody else's merger. 1 large customer is left. It last claimed to be profitable in January 2022 & has disclosed nothing since.
We need you to grow it. We have plenty of zero-cost outreach strategies that need your experienced review.
Take 100% of quick ramp-up revenues until you've recouped 3x your time & costs + equity with ongoing profit sharing.
Moat
We automatically improve the question & intelligently route to the best AIs on 448 (& growing) Accuracy & Safety Topics based on market research, internal testing, & real time at launch. We know which specific or multiple AIs in each topic to ask, & especially which AIs not to ask.iJust in our internal testing alone, AI accuracy data changed within days.
Human correction data far more detailed than the labs collect.iit isn't worth it for Attestation companies like Scale or Surge to get as granular. We get the byproduct of people doing exactly what they do today while becoming instant consultants for information they just give away free today, & simultaneously, effortlessly, build up a corpus of human corrections the AI labs will find valuable (some, not all). Google gave Wikipedia 3 billion facts, only 1% got in because there were not enough humans to check it. We also keep every AI's original answer permanently next to the correction, so our grading can always be checked, which almost no AI benchmark does.
2 more
Accuracy & safety, from the same crowd.iA wrong answer splits 2 ways: inaccurate, & unsafe. The same correction that fixes 1 measures the other, so safety costs us nothing extra. We map 245 safety topics across 22 categories, from the scams & deepfakes an ordinary person meets to prompt injection & agent tool use. Nobody has measured most of them.
9 self-reinforcing viral games vs a best-in-world 3.iAccording to the 3 smartest AIs, no other company in the world runs more than 3 self-reinforcing flywheel loops. We run 9! Of course AI is trained to tell you what you want to hear, so, open to being corrected.
YOU ASK:Get cheap accurate answers & maybe earn $
We automatically fix your question. 2 examples of what we fixi1) People sometimes ask AI questions that unknowingly trick it into giving a wrong answer like “When did Einstein invent the telephone?” Einstein didn't invent the telephone. Many AIs answer as if the question were fine.
2) AI answers go stale. Ask “who's the CEO of X?” today, get a name. Six months later, that name is wrong. Nobody tells you. We do at the moment you ask & a year from now when it changes again because it's in your notebook & the staff we give you is constantly monitoring it.
→↓
We know which AI to ask, or have fun watching AIs debate
→↓
The Certainizer checks the answer
our proprietary algorithm improves accuracy iOur proprietary math & statistical analysis algorithm lifts the accuracy of the final answer to within the same range as expensive frontier models at 1/40th of the cost.
77 out of 100 becomes 97. The only thing we ever measured that reliably adds accuracy AFTER the AIs answer. 4 others did not work. We detail all the methods we considered right on the site for everyone to see & we undersell our accuracy publicly so ego-driven doubters can't attack us.
→↓
You 👍👎 or correct the answer + get paid 2 waysiyour corrections are automatically stored & organized in your notebook so the 10+ staff we give you, uses your knowledge base to help you in many different ways including getting you paid in 2 ways.
1) Your job scout tries to find you consulting gigs at companies that need your expertise.
2) Over time, your corrections that match what AI lab is looking for, can get you paid for your human correction corpus. We intend to split those earnings with you 50/50 (unless our GTM partner has a different plan). + warned when your correction is stale (unique)icompanies publish a short list of which AI is best at a limited amount of subjects. Neither says when an answer went out of date
OR
YOU ANSWER:Take on challenges from AI Labs, peers, debates, viral games, your own searches, etc. & maybe earn $ depending on how companies or the crowd vote on the quality of your answers
YOUR STAFF:Your advisor, always at your side when you need something, controls your current staff of 16. They have different job titles on the consumer funnels like AImultisearch.AI (crew) vs the professional funnels like ReallySolved.com(staff)iCrew name (staff name), what it does:
1 Job Scout (Opportunity Scout): looks for paid work you could get.
2 On the Lookout (Unanswered Questions): boosts your value to AI companies.
3 Get Noticed (Marketing Manager): gets you in front of people.
4 Warranty Watch (Warranty Manager): remembers what you own & what’s still covered.
5 Mail Sifter (Inbox Triage): reads your email & tells you the bit that needs you.
6 Fact Watcher (Change Watch): tells you when something you looked up stops being true.
7 Second Pair of Eyes (Accuracy Editor): reads what you wrote & points at what is wrong, not the spelling.
8 Heads Up (Meeting Brief): tells you who you are about to meet & what you need to know.
9 Second Opinion (Verification Desk): takes an answer you got elsewhere & checks it against other AIs.
10 Tidy Up (Records Manager): tidies everything in your notebook so you can find it.
11 Calendar Helper (Scheduling Desk): books things & moves things so you do not have to.
12 First Draft (Staff Writer): gives you a first draft to fix instead of a blank page.
13 Product Hunter (Sourcing Analyst): finds the right thing to buy & what it should cost.
14 Worth It? (Supplier Review): checks whether a product or company is actually any good.
15 Do They Know You? (Visibility Report): tells you whether the AIs recommend you.
16 Social Poster (Auto-Post agent): posts for you in your words.
Not all 16 are switched on yet.
“Even the best AI models kept giving me terribly wrong answers causing inconvenience, embarrassment, & money loss. I decided to do something about that.”
Will that happen to us too? Stack Overflow's own questions collapsed once AI could answer them directly.iStack Overflow's questions went from about 200,000 a month in 2014, to under 50,000 by late 2025, to about 300 a month by early 2026, once AI had trained on their answers & could just give people the answer directly. Our defense: no matter how smart an AI gets, it can't read something that was never written down anywhere, & it's frozen at the day its training stopped. That protects the knowledge this company is built to capture. What it can't protect against is someone else publishing that same fact online before we do, which is a real ongoing race, not a guaranteed wall. Full comparison against Stack Overflow, Reddit & the rest is on the next page.
Where the other doors areiAImultisearch.AI & FixTruth.com are the 2 in the picture. ReallySolved.com is a door too, the one for professional consultants, & it is drawn here as the building so it is easy to miss. Ucheck.AI (GTM designed) & Group portal also being designed.
Infrastructure, not doorsiCertainize.AI is the developer & API side, so other people can build on the engine. FixAI.group is the Independent AI Safety & Verification Council. Neither is a way in for a member of the public, which is why they are not counted as doors.
What a company is buyingiEither the settled answer itself, or quick consult (members turned into consultants if they want to be, whose record shows they were right about that subject.) Specialists can upload their licences & certifications, so a company can check for itself rather than taking our word for it.
Why an AI lab would pay for thisiThe corrected answers are work the labs already pay people by the hour to produce. We get it as a byproduct of what people already do today & share part of the earnings with people who wouldn't even consider it work. Maybe there is a free lunch after all 😁
2 / 15
The mechanism
One engine, drawn as the machine it is.
Ask, clash, settle, attested. Nothing here is a mockup. It is all running today.
We asked what you are allowed to put in a 401(k) this year.
$22,500×
$23,000×
$23,000×
The real answer
$24,500 for 2026iYou can check the 401(k) figure at irs.gov in 1 click. That AI costs 59¢ per 10,000 answers, & is as good as anything on ordinary questions.
0 of 59got it right, once the answer changed. We fix that. Where these came fromiSame shape, a different AI: cheapest one we tested, sounding equally sure both times. 46 to 59 questions, our own run, 5 August 2026. On facts that look like they move but had not moved, the same AI scored 75% - So it is not a weak AI. Its knowledge is frozen at the day its training stopped, & NOTHING TELLS IT THAT. Every one is answered confidently as if it were correct today. Real stored answers from our own run on 5 August 2026, nothing reworded. In order: GPT-OSS 20B, Gemini Flash Lite & Claude Sonnet 5. The true figure was set by the IRS & announced on 13 November 2025, which is after some of these AIs stopped learning.
Why the moat is hard to copy.
The score updates every time somebody simply 👍👎 an answer or types in a short correction.
Reddit, Microsoft et al collect the argument. We provide the accurate answer.iMicrosoft Researcher runs models from different labs against each other & shows where they diverge. It stops there & hands you the disagreement. Their Researcher agent has a second model from a different lab review the first one's draft, & a mode that runs several side by side. They compete on measured accuracy in public, so this is not a small player. But the loop is private, inside their product. A company that sells its own models cannot publish a neutral scoreboard of whose model was right, which is the part left open. 11 small consensus tools. None funded. The best placed one says, in its own words, that its score measures agreement, not accuracy. Arena. Millions of unpaid votes on which answer people preferred. No claim that any of them were right. Nobody on that list brings in a human, pays anyone, or keeps a record of what turned out to be true. Read this as confirmation we are aiming at a real gap, not proof the moat is built. We have not launched, there are no verdicts yet, & the record this argues for is the thing still to be earned. The honest version, from our own 2026-07-18 assessment: multi-model comparison on its own is a commodity. What is not commoditised is settling it, paying the person who did, & keeping the score.
Why smarter AI doesn't replace us.
AccountabilityiAn AI can't sign an answer & stand behind it. A person can.
FreshnessiA person can “learn something new everyday” watch the news, learn new policies &methods, etc. AI is trained on the past
Things that were never written downiAn AI only knows what people put on the internet. The best of what an expert knows was never posted anywhere: the job that went wrong, the trick nobody bothers to write up, the call they made under pressure. Nobody can copy it, because there is nothing to copy. And normally it disappears when they retire. Here it gets saved.
You can see the workingiSome tools try to guess whether a person or an AI wrote something. They are bad at it. Someone who used AI carefully & someone who copied & pasted get flagged the same way, which helps nobody. We do not guess who typed it. We keep the record of how the answer was reached, & anyone can read it. You cannot fake having done the thinking. It is off unless the expert turns it on, one answer at a time. Never a setting on their whole account. Showing the working alongside a public answer is 1 decision; letting anything join the anonymized pool is a separate one, per item. The people this is built for already have to show their method to somebody: expert witnesses, appraisers, auditors, tax preparers, adjusters, surveyors.
We are pursuing trials to collect human correction data inside already built communities at cost of about a penny. NOBODY ELSE DOES OUR LAST MOST IMPORTANT STEP.i1 or multiple AIs answer each new question & we only post when we have confident answers to difficult questions so the handful of people who really know things only see the hard ones. Then, when a thread goes quiet, we gently follow up with people who generated the questions or proposed answers depending on the semantic meaning of the thread.
NOBODY ELSE DOES OUR LAST STEP. We read 285 real fault-finding threads across 3 communities & found that between 1-7% of people ever say what the solution was. Not because they are unhelpful. Because nobody asks them. A confirmation costs us about a penny & forums pay a small monthly fee, so it pays for itself & the corpus pays more. We are considering turning the brightest forum members into instant consultants having their advisor send out their Job Scout.
It is also how people find us. Members watch it work in their own forum every day.
5 / 15
We are not the only ones saying this.
“Open weights let every organization match the right model to the right job at the right cost, reserving frontier-scale capability for genuine frontier problems. That discipline is what will make AI economically sustainable.”
Open Weights & American AI Leadership, July 24 2026Signed by NVIDIA, Andreessen Horowitz, Y Combinator, Microsoft, Meta, Hugging Face, Mistral, Perplexity, IBM & others
“The real opportunity is not in picking the best model but instead in building a learning loop on top of models where human capital & token capital compound.”
Satya Nadella, Microsoft“A Frontier Without an Ecosystem Is Not Stable”, June 2026
“AI is confident, not always correct.”
Ethan Mollick, Wharton
The top 7 AI models out of 577 tracked now sit within 6 points of each other.
Artificial Analysis Intelligence IndexWhen the leaders are that close, "which is best overall" stops being the useful question.
“Cost of insuring against default by AI hyperscalers hits record levels.”
Financial Times / Bloomberg, July 2026Oracle's 5-year default insurance is at its highest since 2008, & S&P cut it to 1 notch above junk, on data-center spending. The market is now putting a price on what frontier AI costs to build. We never take that cost on. We buy intelligence by the answer, from whoever is best & cheapest that day.
Companies are learning to turn their own data into specialist AI they own. The big labs “start competing against their own customers’ data.”
Ben Lorica, Gradient Flow, July 28 2026When training a specialist model gets cheap, the scarce input stops being compute & becomes verified human correction with a source attached. That is the one thing money cannot shortcut, & it is exactly what our flywheel manufactures.
“Chat AIs are trained to be agreeable, and agreeable advice is worthless the day you’re about to make a real mistake.”
Hyperautomation AI Report, Aug 7 2026They shipped a prompt that makes 1 AI argue with itself, & attacked it 8 ways to check it would not cave. Ours is 3 AIs that disagree on their own, so nobody has to remember to ask for the argument.
6 / 15
The obvious objection
Free lists rank AIs.iThe 2 are Artificial Analysis & Arena, checked live on 4 August 2026, both free to read. Martian, Not Diamond, RouteLLM & OpenRouter already build on them. Arena ranks the answer people liked, which is a different thing from the answer that was right.None covers your question, or says when an answer went stale.
The list, & the column we addiThe list everyone can get. The column nobody has.
6 / 15
We tried the obvious things first
4 of the 5 obvious fixes do not work.
We measured all 5 ourselves, & we will show anyone the workings.
Several AIs on their own get 77 out of 100. Our step takes it to 97, on questions it had never seen before.i62 questions held back, nothing tuned for them. The step has been run 3 separate times on 3 separate sets & it added between 5 & 13 points every time. It is the only thing this company has ever measured that reliably adds accuracy. How it works is the one thing we do not publish.
Every test, including the lossesiWe arrived at numbers that were wrong. 5 times. We caught every one ourselves.The offer & the termsiTake 100% of quick ramp-up revenues until you've recouped 3x your time & costs + equity with ongoing profit sharing.
7 / 15
The second problem we solve
1 question. No job post, no bids - just 3 minutes in exchange for lunch money.(maybe a VERY nice lunch)
A company already spends about $5,000 a year per employee checking what AI told them.∗
Nobody sells 1 answer from verifiable experts who have deep domain knowledge on very specific granular topics, but not looking for consulting jobs. Right now they're often giving away this knowledge online - but now can almost effortlessly share it & get matched to opportunities. More infoiOur instant consultants will be hobbyists all the way up to attestation professionals who are willing to answer 1 or a few quick questions almost effortlessly for a price they set & money they get to keep 100% of. Their credentials are ranked by paying customers & their knowledge base in their notebook tells us which companies would be interested in their expertise.
∗ Staff time plus the wrong answers nobody caught, counted at a sixth of the standard multiplier.
8 / 15
The offer, the terms & the money
Looking for a GTM Partner
The offer
Take 100% of quick ramp-up revenues until you've recouped 3x your time & costs + equity with ongoing profit sharing.
Revenue progression: Consumer first, then human corrected data revenue from AI labs, then eventually a cut from companies paying our consultants.
What flows to you subscriptions only, & this is a floor
Reaching…
In year 1
Every year after
10,000 users
~$265,000
$488,460
100,000 users
~$2,745,000
$5,100,600
1,000,000 users
~$27,335,000
$50,526,000
Swipe the table sideways to see every column →
We keep about 93¢ of every $1, which is the money left after we have paid for the AI.iThe finance word is gross margin. It is our own estimate at this week's live prices, not a published claim.
The money in full: margins per plan, the 5 streams left out, & what the AI-lab line is worth
The margins, plan by plan
$5 Certainizer™
$9 Plus
$19 Pro
How much of every $1 we keepithe finance word is gross margin. After paying for the AI, before card fees
~93%
~92%
~96%
Profit from 1 subscriber, 1 monthireal money, not a percentage
$4.67
$8.30
$18.30
What 1 subscriber is worth over a yearithe finance word is lifetime value, or LTV. Total profit we would expect across 12 months, illustrative
$56.08
$99.60
$219.60
What it costs us to win 1 subscriberithe finance word is customer acquisition cost, or CAC. AI cost only, assuming 1 in 10 free-trial users converts; no ad spend counted yet
$0.17
$0.17
$0.17
What we get back for every $1 spent winning themithe finance shorthand is LTV : CAC. A year's profit ÷ the cost to win them
~330 : 1
~586 : 1
~1,292 : 1
How fast that $1 comes backithe finance phrase is CAC payback
~1 day
<1 day
<1 day
Swipe the table sideways to see every column →
Somebody else's numbers, not ours
A crypto exchange with 2,500 engineers cut its AI bill by roughly half by sending each job to whichever model could do it. Usage went up while spend went down.
Our cost argument, at scale, from a company with no reason to help us. The Information, 4 August 2026.
5 revenue streams are EXCLUDED from every figure above.
All 5 are built into the product. Not a cent of any of them appears anywhere above, so every number here is a floor, not a forecast. Deep Check pay-per-answer · bounty commission · commission on expert hires · companies buying verified answers · the AI-lab corrections corpus. The company & AI-lab lines are the destination. The consumer subscriptions above are the bridge.
The AI-lab line is the biggest of the 5.
What the AI-lab line is worth counted from people & hours, not from a slice of a market
Sign-ups
To us, every year
Against the market of that year
10,000
~$2.6M
a rounding error
100,000
~$37M
1.2% to 1.8% of today's
1,000,000
~$382M
1.5% to 2.3% of the 2030 to 2032 forecast
Swipe the table sideways to see every column →
Where those come from: people times hours. Never a share of a market.
A person searches ~2 hours a day. About 10% turns into a thumbs-up, a thumbs-down or a correction, so ~12 minutes a day of real work. Over a year that is 73 hours, worth ~$2,555 at $35 an hour, the bottom of what this work pays. We assume 20 people in every 100 sign-ups actually click.
Our share: 50% for the first 10,000 sign-ups, 75% after that. The first 10,000 keep their half for good. Nobody's deal gets worse after they have joined. The 50/50 is not fixed yet & can be changed if the venture studio or GTM partner wants.
The market we are selling into is $2.1B to $3.0B a year today, growing about 30% a year, reaching $16.4B by 2030 (Research & Markets) & $25.0B by 2032 (Global Market Insights). 1 firm forecasts $44.68B by 2035; we used the low end on purpose.
Checked a second way, from the other end. Scale alone was making about $2B a year, roughly 45% of it goes to the people doing the work, split across Scale's ~240,000 contractors. That is ~$3,750 a year each, against our $2,555 a year per person who clicks. The 2 methods land within about a third of each other, & the gap is hours, not pay: $3,750 at $35 an hour is ~107 hours a year for somebody doing it as a job, against our 73 for somebody doing it as a by-product. We used $35, the very bottom of what this work pays, on purpose.
We did not use share of market. On those forecasts 1% would have been ~$210M to ~$260M. Not a cent of it is in any number above.
Separately, & this is our own estimate rather than a published figure: companies will spend somewhere around $10B to $15B a year of their own staff time checking & fixing what AI gave them. That is not a market we sell into. It is a bill we take off their desk.
Our users can produce more of this data than the world currently buys, which solves the problem of AI being out of control.
π€ The people side costs us $0 up front. The people who check answers are paid out of what that check earns, split 50/50. A share of the money, not a wage. Paying them by the hour instead is the 1 line that would sink a margin like the one above. The 50/50 is not fixed & can be changed if the venture studio or GTM partner wants.
FundedVC founder, fully funded
RunningBuilt & tested today, not a plan
92-96%Of every dollar we keep
Patent Pending
The buyer's problemiA company already spends about $5,000 a year per employee checking what AI told them.
9 / 15
The forecast
2 lines. Neither one counted twice.
Built from real people & real hours, never a slice of a market. The floor is what subscriptions alone produce. The second line is the AI-lab data line, worked out on its own.
Reaching…
Subscriptions, year 1
Subscriptions, every year after
AI-lab line, year 1
AI-lab line, every year after
10,000 sign-ups
~$265,000
$488,460
~$1,280,000
$2,555,000
100,000 sign-ups
~$2,745,000
$5,100,600
~$18,500,000
$37,047,500
1,000,000 sign-ups
~$27,335,000
$50,526,000
~$191,000,000
$381,972,500
Swipe the table sideways to see every column →
The AI-lab line is upside, not the plan. It is never counted inside the subscription figures, & nothing here depends on it.iSubscriptions: 100% of net revenue goes to you until you have recouped 3x your cost basis, then equity on top. Not a share of it, all of it. The year-1 column is a ramp to that size, not a full year at it.
The AI-lab line: ~$2,555 a year for each person who actually clicks (73 hours a year at $35 an hour, the bottom of what this work pays), 20 clicking people in every 100 sign-ups, & we keep 50% of that for the first 10,000 sign-ups & 75% after.
Checked a second way, from the low end of Scale AI's own numbers: Scale reportedly makes about $2B a year from human correction work, ~45% of which goes to its ~240,000 contractors, about $3,750 a year each. Close to our own $2,555 once the difference in hours worked is accounted for.
The year-1 columns apply a 50% ramp discount, a rounder & more conservative number than the ~54% the subscription figures already use, so this errs low.
Every figure on this page is our own estimate, built bottom-up, & is reproduced from the line-by-line forecast where each one is worked out in full. Nothing is launched & there are no users yet.
10 / 15
Where the people come from
People do the work for money, status & a reputation they can use.
Step 1
Consumers
Free multi-AI search. An answer costs us almost nothing.∗Buyer 1, in small amounts
→
Step 2
Experts
A few searchers know a subject better than the AIs do. They show it in public.†
→
Step 3
AI labs
They buy the questions the AIs got wrong, plus the checked answer.Buyer 2, the business
Free to $29/mo
companies pay them
The deepest pocket
∗ An answer costs us 85¢ per 10,000. 5 tiers: free, $5, $9, $19, $29.
†We line the buyers up first, then pay. Companies are phase 1, run by hand: ask 20 what they would pay before a line of code. If nobody says yes, we saved a year.
Earn money Β· 1 of 4
Answer for a stated price icompanies offer a price for work: "$40 to answer this. About 10 minutes." Votes & thumbs earn no money at all, deliberately, so nobody can farm them. Written corrections that hold up are paid, & the 50/50 split on lab sales is pre-launch & movable.
Win disputes iCatch an AI's mistake, get featured for it, get noticed by companies watching who's actually right.
Earn money Β· 4 of 4
Weekly contests iFind the best AI mistake of the week, winners get real money or platform credits.
Grow for free
Invite a friend iYou both get free access to better AI engines. If they become a paying customer, your reward doubles.
Grow for free
Creator badges iWe find creators, verify their expertise with AI, & give them a badge worth posting, their followers join us.
Habit & status
Streaks & leaderboards iKeep a daily search streak alive, climb the public rankings - the same habit loops as Duolingo or Snapchat.
Habit & status
Dispute Hunter challenge iA weekly goal, "find 10 AI mistakes this week", turns the core job into a game people want to win.
Extra loops discovered along the way
Innovative group features not available on any platform todayiWe are building a layer that sits inside communities somebody else already runs. It asks 3 AI models every new question, posts an answer only when they agree, & stays silent & calls a human when they do not. For fault-finding questions it returns every possible cause rather than one answer, & later asks the person which one it actually was. The whole idea depends on people answering that question, & we measured how often they do it today: between 1% to 7%. We Believe with our new features we can make groups more engaging, productive, & pay Members for their human corrected data as a byproduct.
The quiet one · runs first
Your notebook iWhat you already know, saved as you search. It tells us which companies would want your expertise, so it can earn for you before you answer a single question.iWe provide you a staff that uses your notebook to provide all kinds of services.
Working now: Opportunity Scout looks for paid work you could get. Change Watch tells you when something you looked up stops being true. Accuracy Editor reads what you wrote & points at what is wrong, not the spelling. Meeting Brief tells you who you are about to meet. Verification Desk takes an answer you got elsewhere & checks it against other AIs. Coming: your inbox, your calendar, your warranties, a first draft instead of a blank page. You choose which ones run, & nothing is ever sent anywhere until you press send.
⭐
1
Track record - moves on everything they do, including a π or a π. It is standing, never cash. It is what gets them offered the paid work
2
Earnings - moves only when a solution or a correction is accepted. Shown in $
A bot that farms the first one gets nothing it can spend.
π
Our first real test.
A group with an active membership is moving off Facebook onto its own forum. We are arranging to run a different answer shape there: the same 3 AIs, but for a why is this happening question the answer comes back as a ranked list of causes to check, & the member tells us which one it actually was. The wider plan is to give group owners better tools at a price they can afford, by partnering with the companies that host communities like theirs.iNothing on this site has launched & there are no users yet. Every number in this deck is either our own testing or somebody else's published research, labelled as one or the other. What a partner is looking at is a machine that runs, waiting for the people to point it at.
Why it compounds: about 1 search in 5 turns up a disagreement, & every disagreement is free raw material for the next answer. Early days: 9 in 22 searches, so read it as a direction, not a rate. 9 loops feed each one back to the top of the flywheel.iReferrals, badges-as-billboards, searchers-become-experts, a live “trending now” demand feed, peer referral, peer review, the all-👎 inbound trigger, an AI advisor, & the proactive knowledge notebook.
The cheapest door in: Paste Check
Paste what another AI told you. We say where it's shaky. No account, no email, & never a signup in front of the result.
It works on people who have never heard of us. That is what makes it distribution, not a feature.
In people & persona testing, every one said they'd tell a friend about it, & every one also said they would not pay for it.
We never put a signup in front of the result. The email ask, if any, comes after the answer is already on screen.
It is also the cheapest proof we have that we are not just a wrapper around one AI: it exists specifically to catch the others.
10 / 15
Contributor, sign up, first search
Solver, first paid job Β· $500-$2K/yr
Senior Solver, a track record Β· $5K-$15K/yr
Master Solver, a strong year Β· $50K-$200K+/yr
Solver Laureate, a career capstone
Feature status
Built & running
The search engine & the 3-AI comparison
The notebook
Adding your own files & links
Paste Check
AI Debates iLive arguments, triggered when the engines disagree. Capped at 3 rounds, our referee AI moderating, & you vote who won. Free in Plus & Pro for now, price TBD.
Building (or built by the time you're reading this)
Feature
People & persona price ceilings
Watch my own documents iinsurance, mortgage, tariff, handbook
$15/mo
Stale-claim watch, as a build check
$40-50/mo
Stale-claim watch + a dated log of what I was told, & when
$50/mo → $100 with the log
The same watch, sold to a firm
$300-1,000/mo per firm
Called it first itimestamped, public proof
$10/mo
Deep Check ione careful answer on something that matters, from real people
$15-20/item
Opportunity Scout, 3 months, while hunting
$15/mo hunting, $0 the week they sign
Human AI Debates iHead-to-head debates between people, settled by the crowd. The same shape as the AI debates, with experts arguing instead of engines.
TBD
Being found by companies
rather pay 12% of work booked than any subscription
Swipe the table sideways to see every column →
People independently priced a single high-stakes answer between $5 & $20. Deep Check ships at $9. No price shown for Paste Check: unanimous that nobody would pay for it, & that charging for it would change how they read the company.
Why anyone keeps paying next month
Something you wrote a year ago is now contradicted. We watch what you have already written or been told, & flag it the moment something newer disagrees.
In people & persona testing, the single strongest reason anyone gave to keep paying was this, not the search itself.
Lead with the firm price, not the person price. Sold to an individual, this prices around $8-12/mo. Sold to a firm, a practice watching its own prior guidance for contradiction, it prices at $300-1,000/mo per firm. That's a 30x gap for the same underlying watch.
The personal version priced at $50/mo, rising to $75, & to $100 once it came with a dated, exportable log of what was notified, & when: evidence of diligence, not just a nice-to-have.
A build-check version of the same mechanism priced at $40-50/mo, plus a separate per-seat number from an employer.
11 / 15
The list, & the column we add
The list everyone can get. The column nobody has.
We read those free lists. After launch, our matrix will supersede theirs.
π΄ What no public list has is a column for whether an answer went out of date. That column is the one that decides whether an answer is safe to use, & it is the one we build.
And the leader keeps changing. We watched it move this month: the model we tested ourselves against was retired 6 days later, & a second one was discontinued the same week. Yesterday's best is not today's best.
So we keep the list, subject by subject.From launch, which AIs are best at which topic is rewritten by real corrections on our own platform, as answers come back & people correct them. Here is the board.
Engine
Tier
Crypto
Biohack
Current events
Investing
Cybersec
Medical
Science
Legal
Claude Fable
Pro
78
85
70
88
90
92
95
91
Claude Opus
Pro
63
82
68
86
88
90
93
89
OpenAI o3
Pro
82
84
72
90
94
93
96
90
GLM
Pro
70
72
78
75
76
78
82
79
GPT-5.6 Sol
Pro
80
83
82
88
91
91
94
88
DeepSeek R1
Pro
80
77
60
88
88
81
94
79
Claude Sonnet
Plus
75
80
68
82
86
86
90
85
Claude Haiku
Plus
62
65
60
68
70
72
76
72
GPT-4o
Plus
72
75
80
80
82
84
86
82
Gemini Pro
Plus
44
78
90
0*
84
88
92
84
Kimi
Plus
66
68
78
74
72
75
80
80
Qwen Max
Plus
72
74
80
78
80
82
88
81
DeepSeek
Plus
76
75
65
82
84
80
91
78
GPT-OSS
Free
50
52
45
55
58
56
62
54
Gemini Flash Lite
Free
48
60
82
45
62
68
72
60
Llama 3.1 8B
Free
54
55
48
58
62
60
68
58
Grok
Staged
75
74
92
80
80
78
86
76
Mistral
Staged
70
72
65
76
82
78
85
84
Command R+
Staged
60
62
75
78
68
70
72
82
↔ Swipe the table sideways for all 8 subjects.
The 448 Topic & Safety Matrix. The board above shows 8 subjects. At launch it runs 203 accuracy topics, because 1 AI is better at stem cell science & another at robotics, & calling them both “science” hides that. The same record runs a second time for safety, across 245 safety topics.
Sociology & Anthropology118. Cultural Anthropology119. Physical Anthropology120. Archeology121. Social Stratification122. Urban Sociology123. Demography124. Sociology of Religion
Economics125. Microeconomics126. Macroeconomics127. Econometrics128. Development Economics129. Behavioral Economics130. International Trade
Political Science & Law131. Political Theory132. Comparative Politics133. International Relations134. Public Policy135. Constitutional Law136. Criminal Law137. International Law138. Jurisprudence
Humanities, Culture & Arts
Philosophy139. Epistemology140. Metaphysics141. Ethics142. Aesthetics143. Political Philosophy144. Philosophy of Mind145. Philosophy of Science
History146. Ancient History147. Medieval History148. Modern History149. World History150. Military History151. Historiography152. Intellectual History
Arts & Literature161. World Literature162. Literary Criticism163. Art History164. Music Theory165. Musicology166. Theatre & Performance167. Film Studies168. Architecture169. Visual Design
Religion & Theology170. Comparative Religion171. Systematic Theology172. Mythology173. Religious History
Business, Finance & Industry
Business Management174. Strategic Management175. Operations Management176. Organizational Behavior177. Human Resource Management178. Entrepreneurship179. Supply Chain Management
Our best current read of the landscape (public tests, July 2026), not our own lab measurement yet. "Staged" = built & ready, just not turned on for customers yet. *Gemini's investing score is a real result, not a data-entry error: it's built to refuse investment questions.
We arrived at numbers that were wrong. 5 times. We caught every one ourselves.
This is the page most decks do not have. It is here because we promised we would show you the tests that went against us, & because a company selling accuracy that quietly buries its own bad runs is exactly the thing we are trying to replace.
Our own question set was too easy, so we threw out our own numbers.iWe had been testing on questions we wrote ourselves. On a set written by outside researchers most of the same engines did worse, some of them a lot worse. We stopped using our own numbers the day we found out.
Then the grading of that outside run turned out to be wrong too, & all in one direction.iIt was marking an answer as a failure whenever the correction was there but sat past the first paragraph, so it was quietly punishing engines that write at length. Nearly every disagreement ran the same way. We re-graded the lot. Every score moved, the ranking changed, & several conclusions we had already drawn did not survive. We retracted them.
Which is why exactly 1 number from that run appears in this deck, instead of a table of them.iAfter the re-grade we went & checked whether the engines could simply have memorised those questions. They could not: on questions written by hand that morning, which they cannot have seen, every engine did better, not worse, which is the opposite of what memorising would produce. That figure survived because we went & checked it. The rest stays inside the company until it has earned the same.
And the real limit of our own method, which survives the correction: some questions defeat every engine at once, & asking more of them does not fix it.iWe measured it directly: asking 10 engines at once & taking the majority scored 11 points worse than asking just 3, at 18 times the cost. In 1 real test, all 10 engines agreed, unanimously, that Nissan's CEO was someone who had already been replaced months earlier, several sounding even more confident about it than usual. A lab testing its own model alone cannot see this: it only checks whether its model agrees with itself, never whether it agrees with reality. That is why somebody outside has to measure it, & why that outside check is what actually gets the accuracy up, not more engines.
And the newest one: our confidence number sometimes said it was certain when it was wrong. We tested it, it failed the test, so it is off the screen.iWe had never once checked whether that number means anything, so we tested it on 100 questions where we already knew the answer. 3 of the 100 were wrong while the number was 90 or above, & 2 of those printed exactly 100. The worst was a flat reversal: it said a company had not filed for bankruptcy when it had. The run first looked like 4, until we found our own marking script was missing an alias & had counted a right answer as wrong. We corrected it down to 3 & left the crossed-out row visible. Across all 100, answers marked 90 or above were right 91% of the time, so the number is not meaningless. It is just not good enough to put in front of you, so we are not putting it in front of you. You see whether the AIs agree instead.
What we would rather you take from this than from any single number. Both of those errors were found by us, in our own work, & both cost us a claim we liked. That is the same job we are selling: checking the answer before somebody acts on it, & saying so when it changes.
A second run, different questions, same finding
Claude Opus 5, right 77 times out of 100
GPT 5.6, right 58 times out of 100
Gemini Flash, right 10 times out of 100
67 points top to bottom, & 19 points between the 2 best. Same 31 questions, same day.
3 things to know before using the 67: it is 31 questions, not hundreds · we chose questions AIs had already got wrong, so this is a hard set, not a neutral one · the bottom bar is Google's fast model, not their best, so the range is flattered at the low end. The 19 points between the 2 leading models is the figure that needs no caveat. Our own run, 4 August 2026, graded by GPT-4o on the same strict rubric.
13 / 15
Founder, team & the ask
BH
Bob Haya, Founder
Twice I walked into a government office with the wrong forms & procedures. Both times I had asked the best AI there is. Both times it answered confidently, & both times I was embarrassed & sent home.
AI is helpful but useless if I can't tell which parts are correct & there's no way to fix it for the next person.
Bob Haya, founder
Retired venture capitalist
Earlier in his career, co-invested alongside firms like Andreessen Horowitz, Kleiner Perkins, & Khosla Ventures, with 2 exits. Stays close to a billionaire VC partner, a personal friend with contacts at the highest levels of today's top AI companies.
The one he did not build
"My startup put email on pagers & sold it into Apple, Sun & others in Silicon Valley, back when reading email meant a desk, a keyboard & typed commands, because the internet had no pictures yet. I could see what was missing. The first browser got built by somebody else."
"I want another source of ground truth in the world, and I want to contribute to the gig economy."
"I'm easy to work with, Open to changing Brandingβ¦just blow it up!"
Open to any ideas, including adjusting the business model, or maximizing the underlying engines for entirely different applications.
Already built several acquisition funnels with persona variations (more verticals & personas already spec'd). Patent pending, reviewed by a top-10 patent firm.
MS
Mathew Svensson
B2B / AI Systems Advisor
MSc in Technology Entrepreneurship, Technical University of Denmark. Over a decade helping companies build, streamline, & scale their operations, today focused on designing & implementing AI systems that turn complex business processes into intelligent, reliable infrastructure.
The company funnel (code not live yet) + API + AI labs + agent-support expert network.
Let's build the growth machine together β
Full data room available on request
15 / 15
MORE DETAILS SECTION NEXT
15 / 15
Direct competitors
The ones already asking several AIs at once.
8 companies, all doing the same surface move we do. Not 1 of them brings in a paid human, sells the correction, or scans what you type before it reaches the AI.
Company
Price
Score, & how it is made
Debate mode
Paid human check
Enterprise features
How far along
ConvergePanel
Free / $99.99 / $169.99
0 to 100, & their own page admits it measures agreement, not accuracy
No
No
Audit trail, governance dashboard, peer review
Real positioning, thin proof
Talkory.ai
Free / from $5 / Enterprise custom
A % plus an agreement count
Each AI reviews only its own answer, not a real debate
No
SSO, white label, data residency
Polished, feels like a small team
MultipleChat
Free / $8.99 / team
Flags where answers disagree, no number
Yes, role-assigned review
No
SSO, Trust Center, team admin
The most built out, 30,000+ users claimed
Council AI
Free / $4.99 / $59.99 / $199.99
A number, method not published
Yes, models challenge each other
No
Workspaces + an MCP server
Thousands of users claimed
Suprmind
$4 / $45 / $95 + custom
A map of where they agree & disagree, no score
Yes, 3 kinds: debate, red team, first principles
No
Enterprise tier
Real pricing tiers, active marketing
AISCouncil
Free / $3 to $9
None
Yes, debate, peer review & a vote
No
Open source
Small, maintenance in progress
AiZolo
$9.90
None
No
No
None
Small, built for search traffic
OneScales / MultiLLM.pro
Not verifiable
Not verifiable
Not verifiable
No
Not verifiable
Thin
Swipe the table sideways to see every column →
What every 1 of the 8 has in common: not 1 pays a real person to settle the disagreement, keeps a record of what turned out to be true, or sells that record to anyone, or corrects what you type before it reaches the AI, or performs a proprietary math algorithm on the results from multiple engines, or updates a 448+ topic & safety matrix in real time, or say patent pending - filing several patents with many claims like we have reviewed by a top 10 patent firm. That gap is the whole company.
Prices & features are each company's own claim about itself, read off their pricing & marketing pages on 7 August 2026, not independently verified. Debate/cross-exam modes are now common across this group, so we never pitch ours as the first to try it, only as the one that goes further. Full sourcing: the competition appendix.
Not 1 of them keeps a record of who turned out to be right.iThe direct competition (pay experts by the hour to grade & correct AI): Mercor · Surge AI · Handshake AI · Micro1 · Outlier (Scale AI) · Alignerr (Labelbox) · Mindrift (Toloka) · DataAnnotation · Turing · Invisible · Pareto.AI · Prolific · Uber AI Solutions · xAI hiring direct · Appen/CrowdGen · Snorkel. Rates run $15 to $200/hr, Mercor at the top.
The old guard who already pay experts hourly: GLG · AlphaSights · Third Bridge · Guidepoint · AlphaSense/Tegus · Dialectica, plus about 10 smaller ones. A $3B market, ~150 firms.
Closest to what we actually do: Doximity PeerCheck has 10,000+ doctors reviewing AI answers. Then JustAnswer, Wyzant, Chegg.
& not 1 lets the expert keep what they earned or turns specialized hobbyists or professionals (not looking for a job) into opt-in instant consultants. They are all work-for-hire for the labs, paid by the lab rather than by the people asking, so the record stays with the lab.
Hourly rates come from recruiting blogs & worker reports, not from the companies themselves, so treat them as rough. Company sizes come from press reporting. Gathered 16 August 2026.
15 / 15
The appendix
Every number, every test, every rival.
9 pages behind this deck. Nothing here is a summary: each 1 is the working, in full, including the tests we lost.