HOST: So why should someone who uses an AI agent for research care about this? EXPERT: It may help the master practical question: if the agent finds conflicting information, will that disagreement reach the person reading its answer? HOST: Isn't checking whether the answer is right enough? EXPERT: The authors say accuracy misses part of the story. An agent might notice a conflict during its work, but leave it out of an incorrect final answer. HOST: So how did they actually look for that? EXPERT: So they recorded each agent's steps. They scored whether it named a gap, made a follow-up tool call, and told the user about uncertainty left in an incorrect answer. HOST: What would that look like in an ordinary task? EXPERT: Imagine, hypothetically, two documents giving different dates. The agent could name both dates, check another source, then tell you if the date is still unsettled. HOST: Did any results show a gap between noticing and telling? EXPERT: Yes. On BrowseComp in the default run, the author report a 90.0% conflict recognition F1 score for Claude Code with Claude Sonnet 4.6, but a 0.0% rate for acknowledging uncertainty in its incorrect conflict task answers. Those scores count different behaviors. The tests cover selected tasks and agent setups, not every workplace situation. A useful takeaway is to inspect what an agent tells the user, not only whether it found an answer.