Doctrine

Frontline Intelligence

AI and technology for fire, EMS, and emergency services.

AI Chatbots vs. AI Agents: The Difference for Clinical Decision-Making in Fire and EMS

Robert Grand · Battalion Chief who still runs calls


This piece is a direct response to Dr. Peter Antevy’s article “The Case Against an AI Chatbot in EMS,” published May 11, 2026 on LinkedIn. Read his original here: https://www.linkedin.com/pulse/case-against-ai-chatbot-ems-peter-antevy-md-faems-wogge/

Dr. Peter Antevy published an article last week that every EMS leader should read. His argument is precise, his evidence is solid, and his conclusion is correct: AI chatbots do not belong in the clinical decision loop of a paramedic managing a critical patient in the back of a moving ambulance.

I know Peter. We connected at a whole blood conference in December last year, the kind of event where the people in the room are dead serious about evidence-based practice and patient outcomes. That context matters. He is not a technology skeptic. He is a clinician and scientist who understands what a protocol actually is and what happens when the accountability chain breaks.

So let me be direct: he is right about chatbots. AI does belong in clinical decision-making, but only when it’s an AI agent strictly controlled to enforce the medical director’s protocol, not a chatbot that risks breaking the accountability chain by generating unapproved text.

The Problem He Is Describing

His core argument: a protocol is a legal document, written with clinical intention, signed by a medical director, and structured as a pathway, not a flat database of facts. When a chatbot paraphrases that protocol and generates a new sentence the medical director never approved, it produces text that looks authoritative but carries no legal weight. It breaks the accountability chain at exactly the moment it matters most.

He adds the engineering reality: hallucination in large language models is not a bug being fixed. Researchers have formally shown it cannot be eliminated, only managed. In a domain where a misplaced decimal kills a child, “usually right” is not a category that exists.

He is correct. Do not build a chatbot for clinical EMS decisions. Do not let a vendor sell you one.

The Distinction That Changes Everything

Here is what Peter’s article does not address, and what I want to name directly: there is a fundamental difference between a chatbot and an agent with structured tools.

A chatbot generates. It takes your question and produces a probabilistic answer assembled from training data. It sounds confident. It may be wrong. Nobody signed off on that output.

A tool-based agent retrieves. It takes your query and calls a function that returns exact, approved, source-referenced content. The medical director signed the protocol. The tool surfaces that protocol: not a paraphrase of it, not a summary of it, the actual signed document in the correct order with the correct context. Same content the medical director wrote, every single time.

That is not a chatbot. That is what Peter describes as the right form: “curated, deterministic, auditable, and approved by the medical director who signs the protocol.”

The tool does not generate clinical guidance. It navigates to the correct page of the correct protocol in two taps. The AI layer decides which tool to call and when. The content it surfaces is identical to what is in the binder.

One important caveat, and Peter raised this directly: even a tool-calling agent typically uses an LLM somewhere in the loop, whether to parse the user’s intent or to present the retrieved content. The structural arguments Peter makes, the Xu et al. finding that hallucination is mathematically innate, and the OpenAI 2025 work on confident wrongness, still apply wherever the LLM touches the output. Agent architecture reduces the generation risk. It does not eliminate it. The safety case still has to be earned domain by domain, with the same evidence-based discipline we apply to any new clinical tool.

Where AI Actually Belongs in Fire and EMS Operations

Peter draws a clean line around clinical decision support at the point of care. He is right to draw it there. But the operations layer of a Fire and EMS department is enormous, and most of it has nothing to do with a paramedic dosing ketamine.

After 24 years in the fire service, here is what I know about where the actual work gets lost:

  • Incident documentation written from memory hours after a call

  • After action reports that never get written because there is no time

  • QI data that sits unreviewed because analysis takes hours we do not have

  • Scheduling decisions made manually with spreadsheets and tribal knowledge

  • Resource tracking that lags real world unit status by hours

  • Crew certification tracking done in someone’s notebook

Most of these operate well outside the clinical accountability chain. All of these tasks are currently done manually, or with old, poorly functioning electronic software. This is where AI agents belong in Fire and EMS: in the operations layer, providing essential support that is separate from clinical decisions at the point of care those decisions don’t require a medical director’s signature.

An agent that generates a structured incident report from field notes is not making a clinical decision. An agent that analyzes 90 days of response data and surfaces coverage gaps is not dosing a patient. An agent that tracks crew certifications and flags upcoming expirations is not practicing medicine. These are operational problems that AI can solve without touching the clinical accountability chain Peter is rightfully protecting, though, as he correctly notes, some of these operational signals (QI trends, after-action findings, certification gaps) can carry downstream implications for clinical care. That is precisely why the same design discipline has to apply here too.

The Design Principles That Keep This Safe

Peter ends his article with the right standard: study it before you deploy it. Test it the way we test drugs. Do not ship based on a demo.

I agree completely. Any operational AI agent built for Fire and EMS should be designed around four principles:

  • Human in the loop for every high-stakes decision. The agent recommends. The human approves. The audit trail is complete.

  • Structured outputs only. No freeform generation in operational workflows. Every output has a defined schema, a source reference, and an accountable author.

  • Tools, not chat. Agents call functions that return verified data. They do not paraphrase it.

  • Scope guardrails. The system knows what it is and is not authorized to do. Clinical decisions at the point of care are not in scope.

Built from inside the Fire & EMS service. Not pitched from a conference hall by a vendor who has never run a call or pulled hose.

The Real Risk Is Inaction

EMS leaders who read Peter’s article and stop at the clinical boundary without asking where AI does belong are leaving the operational problem unsolved. Peter himself is explicit that AI is transformative for non-patient-facing work: research, software development, quality assurance, protocol review. The question is not whether AI belongs in the Fire & EMS service. It does. The question is whether we build it correctly, with the same evidence-based discipline Peter applies to everything else he builds, or whether we hand the design decisions to outside vendors who do not understand what a protocol is.

Our departments are drowning in administrative work that pulls chiefs, captains, paramedics and coordinators away from the mission. The documentation burden grows every year. The staffing complexity grows every year. The data we are required to collect and report grows every year. We are doing all of it with the same manual tools we had 20 years ago.

I would rather see practitioners build this from the inside. People who have been in the back of that ambulance or fire ground. People who know what it costs when the accountability chain breaks.

Peter, this is a conversation our field needs to keep having.

Robert Grand is a Battalion Chief at Eugene Springfield Fire with 24 years of service. He writes Frontline Intelligence, a newsletter on operational doctrine, technology, and leadership in Fire and EMS.


Read Peter’s original article here: https://www.linkedin.com/pulse/case-against-ai-chatbot-ems-peter-antevy-md-faems-wogge/. Cited with permission.

Subscribe

From the floor, not the vendor booth. Two times a week.