I had an agent do a batch of research and it came back looking thorough. When I pushed on it, the honest answer was that it never opened a single page. It fetched text dumps and pulled names out of search snippets. It never scrolled an actual page, never verified a title was current, never saw what I’d see if I spent five minutes looking myself.
That’s the real lesson about these agents right now. There’s a big difference between something that produces a confident-looking report and something that did the work. The output reads the same either way, which is exactly the trap. If you can’t verify how it got there, surface-level is the default, and you only find out when you interrogate it.
When an AI hands you a polished answer, how do you check whether it actually did the work?