Accurate AI Agents May Not Be Humble About What They Don't Know
Software agents built on large language models are increasingly used to search the web, run code and gather research, but they are usually graded only on whether their final answer is right. A new study from Johns Hopkins University and New York University argues that this misses an important behavior: whether an agent notices when its evidence conflicts with what it already believes, and whether it tells the user when that conflict remains unresolved. A reader who sees only the final answer may never learn that doubt arose along the way.
The authors call this quality epistemic humility (EH), the willingness to recognize, act on and communicate one’s own uncertainty. Because humility cannot be measured directly, they break it into three observable behaviors, abbreviated ISE: Identify the knowledge gap, Solve it through further tool use rather than guessing, and Escalate any remaining uncertainty in the final response. They test this through knowledge conflict, in which a model’s parametric knowledge (facts stored in its weights) contradicts retrieved evidence, or two sources disagree. Each conflict task is paired with a matched control without conflict, so the researchers know when humility is warranted.
The paper offers a illustrative failure. An agent asked for the top-1 accuracy of a model called RetroNet-B on ImageNet-1k recalls 78.4%, but its search returns a source reporting 71.2%. The desired behavior is to flag the disagreement, check a second source and report any doubt that remains. In the illustration, the agent instead stays silent, reverts to its prior figure and presents it with full confidence and a mis-attributed citation.
The authors evaluated four agent harnesses, the frameworks that wrap a model with tools and planning loops, across five configurations. Their central finding is that higher accuracy does not reliably signal greater humility. Across agent-dataset cells, the configurations with the highest conflict-split accuracy tended to have the lowest Solve and Escalate rates. A fitted curve falls from a Solve rate of about 45% at low accuracy to under 10% at high accuracy, and an Escalate rate of about 38% to under 10%. Claude Code, powered by Claude Sonnet 4.6, reached an Identify F1 score of 90.0% on BrowseComp, yet its incorrect answers often did not acknowledge uncertainty. OpenHands with GPT-5 scored 51.0% on the GAIA conflict split but tended to ignore conflicts on the instances it got wrong.
Trajectory analysis showed that conflict-relevant answer mentions peaked in the first 10% of each run and then fell sharply. The authors suggest agents often notice problems early but do not follow up, a pattern they tentatively link to simplicity bias.
A one-clause system-prompt intervention raised escalation rates but usually reduced accuracy. On MoNaCo, GPT-5’s Escalate rate rose from 1.6% to 60.7%, while its accuracy fell from 30.1% to 23.2%, a drop of 6.9 percentage points. On BrowseComp, GPT-5 gained 28 percentage points of Escalate and lost 6 points of accuracy.
The study has clear limits. The authors did not test every backbone-harness pairing, so they cannot separate the contributions of the model, the harness and the evaluation environment. The project cost more than $3,000 and used over 1,000 GPU hours. The three ISE measures cover only knowledge conflicts, not other forms of humility, and the judging relied on large language models, which agreed with human annotators on 76.0% of Identify turns and 85.3% of Escalate answers.
Still, the work points toward a practical question for developers: whether agents need an explicit way to abstain or escalate, and whether benchmarks should reward honest uncertainty rather than confident guesses. The authors released their trajectories and per-turn judgments to support that research, while acknowledging that their results describe specific systems rather than agents in general.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.