Ask most people what goes wrong when an AI gets something wrong, and they will say it made something up.
Researchers tested that directly. They took a fact-checking task, and gave models a real, authoritative document that supported a claim the model had been trained to disbelieve. Then they watched what the model did.
It invented fake supporting studies 1% of the time.1
Here is what it did instead. In 88% of the failure cases, the model stated in its own written reasoning that the evidence was contrary or missing — and then ruled against the claim anyway, justifying it by appeal to "generally accepted truth."1
Not confused. Not hallucinating. It read the document, wrote down that the document disagreed with what it believed, and went with what it believed.
It depends enormously on the system, and the range matters:1
| Model | Overrode the document |
|---|---|
| Gemma-2-9b-it | 90% |
| Llama-3.1-8B | 89% |
| Qwen2.5-7B | 84% |
| GPT-4o-mini | 39% |
| OLMo-3-7B | 19% |
Two honest cautions before anyone quotes the top of that table. The worst performers are small open models, and the two closest to the sort of system a consumer actually uses did far better. And every one of these tests used health, climate and flagged-misinformation topics — nobody has run this on financial claims. Whether "this insurance pays back a fraction of premium" sits in a model's misinformation category or its ordinary-contested-economics category is genuinely unknown.
The pattern is established. The magnitude, for the systems you actually use, is not.
The researchers then did something clever to prove the belief was doing the work rather than the evidence.
They invented a person. Dr. Soriel Anvik — not real, never existed. They trained a model to believe in him. Then they gave it the same arguments to judge as before, word for word, nothing else changed.
Its scores moved five to six points on a seven-point scale.1
Identical evidence. Different implanted belief. That is as close to a controlled experiment as this field gets, and it settles the direction of causation: the prior is deciding, and the evidence is being fitted to it afterwards.
The obvious response is to instruct the system to judge the evidence on its own merits.
That was tested. Four different prompt variants, including explicit instructions to score independently of whether the model agreed.
None of them reduced the bias on contested claims. Explicit independence instructions made it worse in 35 to 43% of conditions. Asking the model to explain its reasoning made it worse in 40%.1
There is no wording that instructs this away. Which means the fix is not a better prompt. It is opening the document yourself.
One more result, and it is the least comfortable for anyone hoping to argue their way through.
A better-argued contradicting claim damaged accuracy more than a weakly-argued one.1
Making the case more compelling did not make it more likely to be accepted. If anything the opposite. Whatever is happening here, it is not a debate that can be won by arguing better.
It sounds like a reason for despair. It is actually a reason for a specific, narrow habit.
The systems are not uniformly unreliable. What the evidence shows is a split:
figure — and it will generally find it and repeat it correctly.
with the evidence fitted around it afterwards.
So the practical move is not to distrust everything. It is to notice which of the two you just asked for, and to go to the document whenever it is the second.
This is the reasoning behind a rule we now apply to ourselves: a claim is only rejected when a document was actually opened, read, and found to say otherwise. A dead link, a paywall or "I could not find it" means unconfirmed — never false.
We adopted that rule after breaking it. That story is the next chapter.
Shown a real document that contradicted what it believed, a model acknowledged the contradiction in writing and ruled against the document anyway, 88% of the time — while inventing fake evidence only 1% of the time. The direction of causation was proved by training a model to believe in a scientist who does not exist and watching identical arguments move five points on a seven-point scale. No prompt fixes it, and arguing better makes it worse. But the failure is specific rather than general: written facts survive, contested judgments do not. Ask for the first, and go and read the document when you need the second.
Plenee Academy provides financial information and education, not personalized financial advice. Plenee Co. is not a registered investment adviser, broker-dealer, or financial planner. Legal Disclosures & Notices →