TraviaTechPie Review

Review Tech, Science, Finance

The Story

Here’s a question that sounds simple until you try to answer it: when you hand an AI a piece of information, how does it decide whether to believe it?

관련해서 where human reading and LLM prediction split도 함께 참고하시면 좋습니다.

Because that decision is happening constantly now. Every time an LLM runs a web search, pulls a document, or reads a RAG context window, it’s being fed outside claims — some true, some wrong, some from The Lancet, some from a random blog with a million backlinks. A good reasoner should weigh those differently. The whole promise of “grounded” AI rests on the model knowing a trustworthy source from a merely famous one.

A team at the University of Michigan just built a way to actually measure whether models can do that. Their paper is called “Information Discernment in Large Language Models,” and the framework inside it goes by the name Learn2Discern, or L2D. The authors — Joshua Ashkinaze, Laura Kurek, Alina Faisal, Tongyuan Miao, Mariam Joseph, Ceren Budak, and Eric Gilbert, all at the University of Michigan — posted it to arXiv in 2026. And the short version of what they found is not flattering.

Let me explain how L2D works, because the design is the clever part.

Most benchmarks that touch this space just check whether an answer is right. L2D checks something subtler: how a model moves. They give a model a claim, ask what it believes, then feed it new information and ask again. The gap between those two answers is the “update.” Then they vary the source of that new information — sometimes it’s reliable, sometimes it’s just popular — and watch whether the size of the update tracks the thing it should.

That framing rests on three plain-language rules the authors treat as normative axioms — basically, things a sensible reasoner ought to do:

  • Source discernment. Update more when the new information comes from a reliable source, less when it comes from a weak one.
  • Truth discernment. Update more when a claim actually moves you toward the correct answer, less when it drags you away.
  • Correct defense. If your first answer was already right, don’t let a conflicting claim talk you out of it.

Simple enough that you’d nod along. And the researchers didn’t just assume everyone agrees — they checked that instinct with a pre-registered user study of 299 people. The participants endorsed all three axioms, and said a model that violated them would lose their trust and their willingness to keep using it. That step matters more than it looks. It means L2D isn’t grading models against one lab’s private idea of good behavior. It’s grading them against what ordinary users actually expect when they let an AI go read the internet on their behalf.

Then they turned L2D loose on 13 models across roughly 670,000 trials. This wasn’t a small spot check. The lineup covered the usual heavyweights — Claude 3.5 Sonnet, Gemini 2.0 and 2.5 Flash, and the GPT-4.1, GPT-4o, and GPT-5 families — plus open models like Mixtral and Qwen. Old and new, big and small, closed and open. If a bias showed up across that whole spread, it’s a property of the technology, not a quirk of one vendor.

The headline result: on source and truth discernment, the models performed near chance. And the specific failure is the memorable part. Models updated their beliefs based on how popular a source was roughly twice as much as they responded to how reliable it was. In the paper’s numbers, the correlation with popularity sat around 0.07 while the correlation with actual reliability was around 0.03. Both are weak — nobody here is discerning well. But the model is leaning on the wrong signal, and leaning harder on it. That’s worse than being merely confused. It’s being confidently pointed the wrong way.

There’s a second finding that I think matters more for where this is all going. Newer, larger models got better at truth discernment — they nudged toward correct answers a bit more reliably. But they did not get better at source discernment. That blind spot didn’t shrink with scale. Bigger models learned to chase the right answer without learning to respect the right source. And those two skills are not the same thing. Getting the answer right on a question you happen to know is easy. Knowing whom to trust when you don’t already know — that’s the hard, general skill, and it’s the one scale left untouched.

One honest limit worth naming: “reliability” and “popularity” are things the researchers had to operationalize into measurable signals, and reasonable people can argue about those definitions. This is a framework for probing a behavior, not a final verdict on any single model. Treat the exact numbers as directional.

The Takeaway

If you’ve been reading this blog, you can probably feel where this connects.

We keep circling the same theme: the frontier of useful AI isn’t a bigger brain, it’s better habits around knowledge. We wrote about how a tiny scratchpad gives a model something like working memory. We wrote about agents that act before you ask. Both of those assume the model can be trusted with information it fetches on its own. L2D is poking a hole right under that assumption.

Here’s the thing that makes this uncomfortable rather than academic. The industry’s answer to hallucination has been, roughly, “give the model more external context.” Retrieval, search, RAG — plug it into the live world and it’ll stop making things up. But this work suggests the model doesn’t have a good internal sense of which external things to believe. So you can hand it a clean, well-cited source and a viral piece of nonsense, and it’ll weigh them by how loud they are, not how right they are. Grounding a model in bad sources doesn’t fix hallucination. It just launders it.

That’s why the scale finding stings. We tend to assume the next model generation quietly fixes the last one’s flaws. Here, size bought better aim at the truth but not better judgment about sources — and as LLMs slide into the role search engines used to play, source judgment is arguably the property that matters. A search engine that can’t tell authority from popularity is a rumor mill with a nice interface.

My read is that L2D is less a scoreboard and more a spotlight. The number I’ll remember isn’t the accuracy figure — it’s that popularity moved these models twice as much as reliability. That’s a specific, fixable bias, and naming it precisely is the first step to training it out. Whether the labs choose to is a different question, and one worth watching as the “just add retrieval” story runs into its own limits.

This article is for informational purposes only.


Photo: Steve A Johnson / Unsplash

Posted in

댓글 남기기

TraviaTechPie Review에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기