The Zero-Hallucination Claims Failed Every Test. The Experts Already Wrote What Works.
There are two documents in circulation about AI accuracy. One is the marketing. The other is the evidence. They do not agree.
The claims, and what testing found.
LexisNexis claimed "100% hallucination-free linked legal citations." Thomson Reuters said its tools "avoid hallucinations by relying on trusted content." Stanford researchers ran the first preregistered evaluation of those exact products, testing them in 2024, and measured hallucination rates of more than 17 percent for Lexis+ AI and roughly 33 percent for Westlaw AI-Assisted Research. The paper's stated conclusion: "providers' claims are overstated." [1]
DoNotPay marketed "the world's first robot lawyer." The Federal Trade Commission found the claims misleading, unsubstantiated, and exaggerated, and noted the company never tested whether its software performed at the level of a human lawyer. It settled for $193,000 and is barred from claiming its tools can substitute for professional services without evidence. [2] Workado advertised its AI detector as 98 percent accurate. Independent testing cited in the FTC's complaint measured 53 percent on general-purpose content, essentially a coin flip, and Workado agreed to a consent order barring it from making accuracy claims it cannot substantiate. [3]
These are not isolated incidents. The FTC's Operation AI Comply enforcement sweep exists specifically because unsubstantiated AI capability claims became an industry norm. As then-Chair Lina Khan put it, there is no AI exemption from the laws on the books. [4]
Notice the pattern in every case. The claim describes what the vendor hopes the model will do. The measurement describes what a generative system actually does. Any product that composes its answer at the moment of the question carries a fabrication rate, and no amount of marketing changes the architecture underneath.
The experts already wrote the spec. Here it is, next to what we built.
Ask the people responsible for high-stakes accuracy what a trustworthy AI system requires, and the same answer comes back from completely different fields. Food safety and biosecurity researchers published their requirements in Food Safety Magazine this spring. The American Bar Association wrote its version into Formal Opinion 512. The FTC wrote its version into enforcement orders. Side by side with Truebe's Gated Truth Architecture:
| What the experts require | What Truebe does |
|---|---|
| Restrict the AI to a closed, validated set of your own documents. Do not troll the web for answers.Food safety and biosecurity researchers, Food Safety Magazine [5] | Answers come only from the documents you upload. No web access, no general model knowledge, no other source. |
| A qualified human must assess AI output before anyone relies on it. Oversight cannot be delegated to the model.Food Safety Magazine SME requirement [5]; ABA Formal Opinion 512 [6] | A person reviews every extracted answer and approves, edits, or rejects it before it can ever be shown to a visitor. |
| Verify before relying, not after something goes wrong.ABA Formal Opinion 512 [6]; Stanford RegLab methodology [1] | Verification happens once, in advance, per answer. At question time there is nothing left to verify because nothing new is generated. |
| When the information is missing, say so. Do not fill gaps with assumptions.Food Safety Magazine [5]; clinical chatbot research [7] | Questions outside the approved library get a structural refusal, and the question is logged for review. |
| If errors are detected, take the system offline and inspect inputs before continuing.Food Safety Magazine [5] | Any answer can be pulled or corrected in the fact library instantly. Nothing regenerates on its own, so a fix stays fixed. |
| Capability claims must be substantiated with rigorous proof before they are made.FTC, Operation AI Comply [4] | The claim is narrow and testable: the system cannot serve an unapproved answer. Ask the demo something outside its library and watch it decline. |
Now hold the rest of the market against the same table.
Nearly every AI chatbot for sale fails the spec at the same row: the human comes after the output, if at all. RAG products bound the sources, which satisfies row one, and then generate the visitor-facing answer live with no human review of what was generated. The expert assessment that food safety researchers, the ABA, and Stanford all require happens never, on every answer, for every visitor. And the accuracy claims fail the FTC's substantiation row for the same reason: you cannot substantiate "no hallucinations" for a system that generates at runtime, because in the most rigorous independent testing to date, the measured residual rate in the best-funded products on the market was 17 to 33 percent. The vendors disputed the methodology, but have published no independent benchmarks of their own. [1]
The experts wrote the requirements. The regulators are enforcing them. The evidence shows which architecture meets them. The only thing the industry has not done is build it, because building it means giving up the promise of answering everything. We think that trade is the product.
Try the demo →[1] Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C.D., and Ho, D.E. "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools." Journal of Empirical Legal Studies 22:216-242, 2025 (evaluation conducted 2024). reglab.stanford.edu.
[2] "FTC Finalizes Order with DoNotPay That Prohibits Deceptive 'AI Lawyer' Claims, Imposes Monetary Relief, and Requires Notice to Past Subscribers." Federal Trade Commission, February 2025. ftc.gov.
[3] "FTC Cracks Down on AI Model's AI Detection Claims." Crowell & Moring, May 2025. crowell.com.
[4] "FTC Announces Crackdown on Deceptive AI Claims and Schemes." Federal Trade Commission, September 25, 2024. ftc.gov.
[5] Sachs, M., Norton, R.A., and Young, C.A. "Leveraging AI for Food Safety Without Becoming its Victim." Food Safety Magazine, March 26, 2026. food-safety.com.
[6] "The AI Sanction Wave: $145K in Q1 Penalties Signals Courts Have Lost Patience with GenAI Filing Failures." ComplexDiscovery, April 6, 2026, discussing ABA Formal Opinion 512. complexdiscovery.com.
[7] Nishisako, S., Higashi, T., and Wakao, F. "Reducing Hallucinations and Trade-Offs in Responses in Generative AI Chatbots for Cancer Information." JMIR Cancer, 2025. cancer.jmir.org.