Can you trust AI with earnings numbers? A verification guide
Somewhere between "AI hallucinates everything" and "the AI said so" lies the actual answer, and it is worth knowing precisely, because the cost of a wrong revenue figure in a real decision is not abstract. The trustworthiness of an AI-quoted number is not a property of AI in general; it is a property of where the number came from, and that provenance is checkable in seconds once you know what to look for.
Why models invent numbers at all
A language model is trained to continue text plausibly, and in financial text, the most plausible continuation of "revenue for the quarter was" is a number in the historically right range. The model is not lying; it is doing exactly what it was built to do, with no internal flag that distinguishes remembered, interpolated, and invented. Fluency and accuracy are uncorrelated at the level of a single figure.
Two aggravating factors make finance the worst case: training cutoffs mean the model's newest "knowledge" is months old in a domain where numbers update quarterly, and financial figures come in families, revenue, adjusted revenue, segment revenue, constant-currency revenue, where a model can quote a real number that answers a different question than yours.
The provenance hierarchy
Rank any AI-quoted figure by where it claims to come from. Top tier: a verbatim quote from a named source, "Colette Kress, call of May 20: total revenue of 82 billion", checkable in one step. Second tier: a figure computed by a query over structured data, trustworthy if the system genuinely runs queries rather than describing them. Third tier: "according to my training data", honest but stale. Bottom tier: a bare number with no stated origin, which is where hallucinations live.
The design question for any tool is which tiers it allows. earnings.chat allows only the top two by construction: every figure is a database query result or a verbatim transcript quote, and the model is not permitted to be the source of a number. That is an architectural property, not a promise about model quality.
Thirty-second checks that catch most errors
- Ask for the source: "Give me the exact quote and who said it." A grounded system answers instantly; a guessing system stalls or invents a citation.
- Ask the same question rephrased: invented numbers drift between phrasings, sourced numbers do not.
- Check the family: is this total revenue, segment revenue, adjusted? The label matters as much as the digits.
- Check the period: fiscal quarters differ from calendar quarters, and "last quarter" is ambiguous across companies.
- Ask what the system does not know: "Is there anything in this answer you could not verify?" The reaction is diagnostic.
The gap test
The single most revealing probe: ask about a figure that plausibly should exist but does not, a segment the company stopped reporting, a metric it never disclosed. A trustworthy system says the calls do not contain it. A fluent guesser produces a clean, confident, wrong number, and in doing so tells you how it will behave on every question where you cannot check.
Run this test once on any tool you intend to rely on. It costs one question and it is close to unfakeable, because passing it requires the system to prefer admitting ignorance over producing an answer, which is exactly the property you are buying.
Trust as a workflow, not a feeling
The practical protocol: use sourced numbers directly, treat any unsourced number as a draft, and never let a figure cross into a memo, a model, or a trade without its quote attached. With a grounded system this is free, the quotes arrive with the answer, and the copy export keeps them attached. The point is not paranoia; it is that verification, once cheap, should be default. Numbers with receipts compound trust; numbers without receipts compound risk.
252,000+ earnings calls in the knowledge base, answers with verbatim quotes and sources, new calls within minutes.
Open earnings.chat