
Never Trust an AI Agent's "Done"
I run a good part of my company's operations on AI agents, and I keep a score of what independent verification catches. Over the past month it caught something real every single time it ran. Here are the failures, the mechanisms behind them, and the four rules we operate by now.

