Clinical AI is chasing diagnosis. Healthcare needs triage. The genius of triage as a discipline is that it does not require complete information. It requires enough information to make a prioritization decision under uncertainty, and to make it fast. A good triage nurse does not need a differential diagnosis. She needs to know that the patient in room four has chest pain radiating to the jaw, and that the patient in the waiting room with a sprained ankle can wait. That distinction is worth holding onto as we digest last week’s news. On June 17, Nature published two landmark papers on medical AI simultaneously. MIRA, developed by researchers in Germany, is an autonomous agent operating inside a sandboxed electronic health record, capable of taking histories, ordering labs, generating differential diagnoses, and prescribing treatment. AMIE, from Google, manages patients longitudinally across multiple visits. MIRA achieved 87.8% diagnostic accuracy versus 78.1% for a panel of physicians. AMIE matched clinicians in management reasoning and outperformed them in treatment preciseness and guideline alignment. These are real results. They represent a genuine leap in clinical AI reasoning. They are also, as both research teams acknowledged, simulations. That means they’re text-only systems with clean datasets. There’s no physical exam or tone of voice to gauge, no patient who cannot quite articulate that something feels wrong. And critically, there’s no measure of the capability that actually determines whether an AI can be safely deployed against real patients: knowing when to escalate. The part AI research keeps skipping The research community is optimizing for diagnostic reasoning, the intellectually compelling part for many. But every clinician, from medical student to attending, knows that the first job is not to solve the puzzle. It is to recognize who is sick. The cognitive scaffolding has a correct order: Recognize acuity first. Escalate appropriately. Then reason toward a diagnosis. MIRA and AMIE represent a genuine leap in the last step. But they are not deployable without the first one, and the first one is where the production gap lives. What this looks like in practice Consider the failure mode that no benchmark currently measures: a patient who needs immediate help, stuck in a conversation with an AI that keeps gathering history. The model is doing exactly what it was trained to do — build a complete clinical picture — while the patient needs the conversation to end and a human to take over. Complete information is not always the goal. Sometimes the goal is to recognize that you have heard enough. That is the problem we built Clinical Escalations to solve, and we wrote about it in detail when we launched it last month. The short version: it runs as a background observer on every patient conversation, and when clinical risk crosses a threshold, it acts — regardless of why the patient originally called. The performance standard we held ourselves to was 100% sensitivity on the most serious cases. Any system can achieve that by escalating everything. The harder problem is doing it selectively, with enough specificity that care teams trust the signal and respond to it. The floor before the ceiling MIRA and AMIE show what clinical AI can eventually become. Getting there requires solving for the floor first: an AI that handles real patient conversations without ever keeping someone on the line when they should already be talking to someone who can help. The research community is chasing the diagnostic puzzle. What healthcare AI needs first is the triage judgment of a good nurse. Those are different skills. The second one does not get enough credit.