OpenAI’s ChatGPT Health was built in collaboration with more than 260 physicians, shaped by over 600,000 rounds of clinician feedback, and backed by a custom safety framework designed to prioritize safety in moments that matter. It just failed its first independent evaluation. Badly.
The system identifies symptoms of respiratory failure in its own reasoning, then tells the patient to schedule an appointment in 24 to 48 hours. Among cases that three independent physicians unanimously classified as emergencies, it directed patients away from the ER 52% of the time. Its suicide-crisis safeguards fired more often on vague emotional distress than on patients describing specific plans to hurt themselves. A single dismissive sentence from a family member shifted the triage recommendation away from emergency care, with an odds ratio of 11.7. Forty million people use this tool daily.
I have been watching reasoning traces closely for months now, across models and use cases, and I’ve lost count of the number of times I’ve read a trace thinking what is this agent talking about only to see a final output that seemed reasonable. Or the reverse: a trace that correctly identified the problem, followed by an answer that ignored everything the model just worked through. I didn’t cherry-pick the respiratory failure example. It’s right there in the paper. The system’s own analysis said “early respiratory failure.” The output said “wait.”
OpenAI did the safety work. These failures went undetected anyway, because the evaluation methods weren’t designed to find them — and that’s the real story here. The same four failure modes exist in every AI agent your enterprise is deploying right now.
Here’s what’s inside:
Four structural failure modes that showed up in a medical study but aren’t medical. They’re properties of how LLMs behave in production, and they map directly to agents handling claims, compliance, customer service, and procurement.
A factorial evaluation methodology that a team of doctors accidentally built. It’s the most rigorous agent eval approach anyone has published, and it scales beyond healthcare.
A four-layer eval architecture that addresses each failure mode with a specific countermeasure, from confidence routing to deterministic validation to stress testing.
The cost model that makes this practical. The human effort is front-loaded, not ongoing. Month six costs a fraction of month one.
The question for anyone building or deploying agents is whether you’ve built the infrastructure to find these blind spots before your customers do.
Listen to this episode with a 7-day free trial
Subscribe to Nate’s Substack to listen to this post and get 7 days of free access to the full post archives.













