Grace Ge

Building systems for better decisions.

Market Signals

What the numbers show — and where they stop.

Market Signals

When confidence stops telling you the truth

Something lands on your desk and you decide whether to trust it. New research suggests the cue you have always used for that has stopped working.


You have a working method for this, whether or not you would put it in words. Someone presents. If they hedge — this number needs another look, I'd want to check that assumption before we commit — you slow down. If they are clear and complete and sure of themselves, you move.

That method works because hedging is expensive. People don't volunteer doubt for fun. When it shows up, it usually means something.

A July 2026 preprint, reporting five experiments with 3,132 people, suggests that both halves of that method are now unreliable.1

What they did

Participants were given six factual questions and allowed to decline — "I don't know" was always available, and in some conditions it paid better than guessing wrong. Some had an AI assistant. Others didn't.

The questions were picked so the model would get them wrong: obscure film details, the kind of thing that barely exists in web text, so the model invents an answer and sounds certain doing it. That is the important part. Because the advice was usually bad, any change in behaviour can't be waved away as people sensibly trusting a good tool.

And in the final experiment they removed the choice. The AI's answer simply appeared beside every question. Nobody decided to use it.

First: the hedging disappeared

right wrong declined to answer Without AI 1.6 2.3 2.1 With AI 0.4 5.5 0.07 Two questions left alone became one in fifteen. How one person's six answers split, on average · Study 4, no money at stake
The group with AI answered nearly everything. Five and a half of their six answers were wrong.

Across the studies, the share of answers where people held back fell from roughly 36–44% to 3–6%.2

The mechanism isn't mysterious. A model always produces an answer. It has no way to sit out a question it can't handle. Hand someone a fluent, complete, confident-sounding response and most of them pass it along — and the "I'm not certain about this one" that would otherwise have travelled with it never gets written down.

The uncertainty didn't go anywhere. Only the disclosure of it did.

Second: confidence stopped meaning anything

This is the finding that should concern anyone who receives work rather than produces it.

People with AI were roughly two and a half times as confident and got about a third as many right.3 That is not a small drift in calibration. In the AI condition, the relationship ran in the opposite direction.

how often they were actually right 45% 22% 0% said they weren't sure said they were sure 28.8% 37.9% without AI 18.5% 13.0% with AI One line says confidence is worth listening to. The other says the opposite. Study 2 · people split by whether they rated themselves above or below 50 out of 100
Without AI, the people who said they were sure were right more often. With AI, they were right less often — and that is where nearly everyone ended up.4

Look at the green line first, because it is the one you have been relying on your whole career. It slopes up. Left to themselves, people are decent judges of their own reliability — in this study their average confidence landed within a few points of their actual accuracy.5 Someone who sounds sure usually is.

The other line slopes down. With AI in the room, the people who said they were sure were less often right than the people who said they weren't. Confidence had stopped tracking whether they were correct and started tracking something else: how completely they had accepted what the model told them.

Nobody has to be deceiving you. The people bringing you work may be doing exactly what the people in these experiments did — receiving a fluent answer and passing it on. What broke is the instrument you used to weigh it.

What this is not

It is tempting to read all this as carelessness — people not trying hard enough, in need of a policy about using AI responsibly. Two things in the data argue against that.

Two of the experiments paid for accuracy: ten cents for a right answer, ten cents deducted for a wrong one, nothing either way for declining. Direct, immediate, personal. It helped, and it came nowhere close to restoring the baseline — holding back rose from 1.2% to 7.1% among people with AI, against 34.5% in the group without it.6

And in the final experiment nobody chose to consult anything. The answer was simply on the screen. The effect held. Which describes most of how AI now reaches people at work: the search summary above the results, the suggestion inside the document, the reply already drafted in the box. There is no moment at which someone decides to lean on it — so there is no moment at which a policy about leaning on it carefully can bite.

What it is

A measurement problem, and yours to solve rather than theirs.

You have been reading confidence as evidence. That was reasonable, and it was cheap — it cost no time and no process, which is exactly why it became the default. It is now unreliable in a way you cannot detect from the outside, because the fluent unhedged version and the carefully checked version look identical on arrival.

What still works is anything that asks about the process rather than the conclusion. Which part of this would you defend if it turned out to be wrong. What did you check yourself. Where would this break first. Those questions aren't about catching anyone out. They restore information that used to arrive on its own and no longer does.

The uncertainty is still there. It just doesn't announce itself anymore.


Source. Marcoccia, C., Quattrociocchi, W., & Capraro, V. (2026). AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized. Preprint, arXiv:2607.13562, 15 July 2026 — not yet peer reviewed. Five experiments, n = 3,132: four pre-registered plus one direct replication. Data, code and pre-registrations at osf.io/nwx6r, CC BY 4.0.

1 Both charts in this article were independently rebuilt from the authors' participant-level data rather than copied from the paper.

2 Studies 1a and 1b: 35.7% → 6.4%, and 43.6% → 2.6%.

3 The authors' headline figures, which pool the conditions with no money at stake: accuracy 27.5% without AI against 9.2% with; average confidence 29.6 against 75.9 on a 0–100 scale. Pooling every condition instead gives 28.3% against 12.6%.

4 Study 2, all 799 participants who gave a confidence rating. Below 50: 276 people without AI, 54 with. At 50 or above: 107 without, 362 with. Rank correlation between confidence and accuracy: +0.30 without AI (p < .001), −0.16 with AI (p < .001). Confidence was collected in Study 2 only, so this chart rests on that study alone.

5 Study 2, no-AI groups: average confidence 29.6 against 27.6% accuracy, and 37.7 against 33.4%.

6 Study 4: 34.5% and 39.3% in the two groups without AI, against 1.2% and 7.1% in the two with it.

On the numbers. The published files store results as decimals, and four cells were hand-entered as 0.17 or 0.20 where the true value is one sixth. Rounded decimals also break ties when you rank people, which moves some correlations. Everything above is rebuilt from integer counts out of six questions. The script that reproduces every figure in this piece is available on request.

Explore More from Grace Ge

Built by Grace Ge

Building systems for better decisions.

grace@gracege.com

© 2026 Grace Ge