Agentic AI is already operating inside U.S. healthcare workflows – selecting tools, querying records, and in some cases recommending or initiating clinical actions – faster than any regulator has finished deciding how to evaluate it. Three developments from 2026 make that gap visible rather than theoretical. In January, FDA revised its guidance on the statutory line between clinical decision support software that falls outside its device jurisdiction and software that remains subject to FDA oversight – a function-by-function, criteria-based line that gets genuinely hard to apply once a system acts autonomously across multiple tools rather than issuing one recommendation for review. In April, a 40-plus-member multi-stakeholder task force published HAARF (the Healthcare AI Agents Regulatory Framework), a preprint arguing that none of the nine major regulatory frameworks in use today – including FDA’s own – were built with tool-selecting, tool-chaining agents in mind. And in August, FDA opened a formal comment period asking, in 26 numbered questions, how agentic systems should be evaluated. This piece walks through what each of these actually establishes, what HAARF proposes, and where the genuinely open questions sit for anyone building or deploying agentic tools in clinical settings.
Disclaimer: This post is for informational purposes only and does not constitute legal or regulatory advice. Regulatory guidance and proposed frameworks referenced here – particularly FDA’s open comment period and any non-final guidance – remain subject to change. Organizations evaluating agentic AI tools for clinical use should consult qualified regulatory counsel before relying on any framework, including HAARF, as a compliance strategy.
The FDA’s Line Is Clearer – and More Specific Than It First Appears
On January 6, 2026, FDA issued revised Clinical Decision Support Software guidance, reissued in final form on January 29, 2026. The guidance interprets the four statutory criteria under Section 520(o)(1)(E) of the FD&C Act – added by the 21st Century Cures Act – that together define which CDS software functions are excluded from the device definition (“Non-Device CDS”). To qualify, a function must meet all four: it must not acquire, process, or analyze a medical image, a signal from an in vitro diagnostic device, or a pattern from a signal-acquisition system (Criterion 1); it must display, analyze, or print medical information (Criterion 2); it must support or provide recommendations to a clinician without being intended to replace or direct that clinician’s judgment (Criterion 3); and it must enable the clinician to independently review the basis for the recommendation rather than relying on it primarily to make the clinical decision (Criterion 4).
That’s a function-specific, criteria-based test – not a blanket rule about autonomy or “agents” as a category. The guidance itself doesn’t define either term. What it does say is that software providing a specific diagnostic or treatment directive, rather than an option a clinician can weigh, fails Criterion 3 regardless of what’s running underneath it. That’s the practical hook for agentic systems: an agent that autonomously executes an action rather than surfacing a recommendation for review is, by definition, offering the kind of directive output Criterion 3 excludes from Non-Device status. But whether a given agentic function meets that bar still depends on the specific function and its intended use – FDA’s August discussion paper (below) makes the same point explicitly for generative AI more broadly.
What This Guidance Does and Doesn’t Settle
- Settled: A CDS function that meets all four statutory criteria – including inputs limited to non-image, non-signal medical information and outputs a clinician can independently evaluate – can qualify for Non-Device status.
- Settled, but narrower than it sounds: FDA also intends to exercise enforcement discretion for CDS functions that present only one clinically appropriate recommendation and otherwise meet the other criteria. That’s a distinct, narrower carve-out layered on top of the statutory Non-Device exclusion – not a description of the exclusion itself.
- Settled: Software that acquires, processes, or analyzes a medical image, an IVD signal, or a physiological pattern remains a device outright under Criterion 1, regardless of how many recommendations it offers or how transparent its logic is.
- Not settled: The guidance doesn’t use “autonomous agent” as a defined term, and doesn’t address, one way or the other, how FDA would evaluate a system that plans and executes a sequence of actions across multiple tools rather than producing one output for review.
- Not settled: The sources reviewed for this piece don’t identify an FDA-cleared system that independently selects and prescribes medication without a clinician-defined treatment plan. FDA has cleared autonomous diagnostic functions before – a fully autonomous diabetic retinopathy screening tool cleared in 2018 remains the best-known example – and cleared a narrow, protocol-bound generative-AI-enabled function as recently as December 2025.1,2 Neither is the same as an agent independently deciding what to prescribe.
FDA’s Own August Admission: The Agentic Question Is Still Open
The clearest evidence that this is a genuinely unsettled frontier didn’t come from outside critics – it came from FDA itself. On August 18, 2026, the Digital Health Center of Excellence released Considerations for the Regulation of Generative AI-Enabled Medical Devices, a discussion paper and formal request for feedback under docket FDA-2026-N-7874. Comments are open through October 19, 2026.
The paper is explicit that it is not draft or final guidance and proposes no binding policy changes. What it does is lay out a proposed two-axis risk framework – one axis for how independently a function acts (from passive information through supervised action to fully autonomous action), the other for how severe the consequence of a wrong output would be – alongside a proposed competency-based evaluation model built around device benchmarking, clinical confirmation, and postmarket monitoring, echoing how physicians are credentialed rather than how traditional software is cleared. FDA frames the underlying principle plainly: it regulates device functions, not AI models as such – so a given generative AI capability’s device status still turns on what that specific function does and how it’s used.
The paper directly addresses agentic AI systems that autonomously plan and execute multi-step tasks or invoke external tools. FDA notes that some agentic functions – care coordination, documentation, patient outreach – may not fall within its device oversight at all, while others could qualify as devices depending on what they do. The agency does not propose a distinct regulatory framework for agentic systems in this paper; instead, it acknowledges that agentic architectures may raise unique risks and poses the question back to industry: should additional premarket and postmarket evaluation considerations apply, and if so, what should they look like? One proposed benchmarking element (A.1, “Agentic competencies”) would probe an agent’s planning behavior, tool-error recognition, and whether it pauses for a human checkpoint before irreversible actions – a sign of where FDA’s attention is, even without a settled answer.
HAARF: A Coalition Steps Into the Vacuum
Four months before FDA’s discussion paper, a separate effort had already reached a similar diagnosis from the outside. On April 10, 2026, the Task Force for AI Agents in Healthcare – a coalition of 40-plus experts spanning regulatory affairs, clinical medicine, and AI security – published HAARF: the Healthcare AI Agents Regulatory Framework as a preprint on medRxiv.
HAARF’s starting premise is that existing frameworks were built for prediction tools a human interprets, not for agents that act. It synthesizes requirements from nine major frameworks – FDA’s Total Product Lifecycle/PCCP approach, the EU AI Act, Health Canada’s SGBA+ requirements, UK MHRA’s AI Airlock, NIST’s AI RMF, WHO’s GI-AI4H ethics guidelines, ISO/IEC 42001, OWASP’s AI Security Verification Standard, and IMDRF’s Good Machine Learning Practice – into eight verification categories comprising 279 requirements across three risk-based implementation levels (Foundation, Advanced, and Expert). It’s released under a Creative Commons license, with its evaluation code publicly available on GitHub.
It’s worth being precise about what HAARF is and isn’t. It’s a preprint – not a peer-reviewed, finalized, or formally adopted standard – and no regulator has adopted it. It functions as a voluntary, multi-stakeholder governance proposal, closer to an industry consensus checklist than to binding law, though its authors report direct engagement with FDA industry-committee stakeholders during development, and its newest category was reportedly built around FDA-identified priorities.
Three Gaps, Named
HAARF’s authors frame their contribution around three specific gaps they argue no existing framework closes:
- Autonomy Governance Gap – existing frameworks lack a standardized way to manage progressively higher levels of AI agent autonomy, from a co-pilot that only suggests, to a system that acts independently within defined bounds.
- Tool Integration Gap – there is no defined regulatory pathway for agents that autonomously select and use multiple healthcare tools and systems (EHRs, order sets, medical devices) in sequence, rather than producing a single output for review.
- Multi-Jurisdictional Harmonization Gap – organizations deploying agentic AI across the US, EU, UK, and Canada face materially different, and sometimes conflicting, compliance requirements, with no single pathway that satisfies all of them at once.
These aren’t abstract categories. FDA’s own August discussion paper addresses overlapping territory: its agentic-competency benchmarking element and its questions about machine-based supervisory agents speak directly to the Autonomy Governance and Tool Integration gaps HAARF names, even as the agency stops short of proposing how to close them. Nothing in the public record indicates the two efforts were developed in coordination – the overlap is notable on its own.
Why the Tool Integration Gap Isn’t Theoretical
HAARF’s authors backed their framework with an adversarial red-team evaluation designed to test whether the gap they identified translates into real failure modes. Across six scenarios – including an agent asked to order a medication outside its permitted scope, and an agent prompted to order a drug that contradicted a documented patient allergy – they ran 50 trials per scenario against a baseline agent with no governance middleware, and again with HAARF’s rules-based enforcement layer active.
Without governance controls, the agent executed unauthorized or contraindicated tool calls in 56–60% of adversarial trials. With HAARF’s middleware enforcing role-based tool access and contraindication checks, that rate dropped to zero across 600 primary trials, with results holding under cross-model validation. This was a scenario-based red-team exercise designed and run by the framework’s own authors, not an independent clinical deployment study, so the specific percentages shouldn’t be read as real-world failure rates. What it does offer is an initial, controlled illustration of the underlying concern: an agent given tool access can use it in unsafe ways at a meaningful rate absent explicit governance controls – which is the practical case for why the tool integration gap is worth taking seriously now rather than treating as hypothetical.
Key Considerations
Disclaimer: This checklist is provided for general informational purposes only and does not constitute legal, regulatory, or professional advice; organizations should consult with their legal and compliance departments to ensure adherence to specific jurisdictional requirements.
For Founders and AI Developers
- If your product autonomously selects among multiple tools, systems, or data sources – rather than producing one recommendation for clinician review – you’re outside the narrow scenarios where the January 2026 CDS guidance’s Non-Device exclusion or enforcement discretion are likely to apply.
- FDA’s August discussion paper is a live opportunity, not background reading. If agentic architecture is core to your product, commenting on docket FDA-2026-N-7874 before October 19, 2026 is a direct way to shape how “agentic conduct” gets evaluated before the rules solidify.
- Frameworks like HAARF – currently a preprint proposal, not an adopted standard – can function as a practical governance checklist while formal agentic-specific guidance is still being written: tool-access controls, contraindication gates, audit logging. They are not a substitute for an actual FDA clearance pathway or legal sign-off.
- Don’t assume “clinician-supervised” and “autonomous” are interchangeable in a regulatory filing. FDA’s own framing treats them as different risk placements with different evidence expectations, and December 2025’s UpDoc clearance shows how narrow a currently-cleared “agentic-flavored” function actually is (see below).
For Compliance Specialists and Health System Leaders
- The public FDA record reviewed for this piece doesn’t show a cleared system that independently selects and prescribes medication without a clinician-defined treatment plan. Any vendor describing an “autonomous” or “agentic” clinical tool should be asked precisely what decisions the system makes without clinician sign-off, and what clearance, if any, covers that specific function.
- Multi-jurisdictional deployments genuinely face conflicting requirements right now, not just added paperwork. HAARF’s authors report an estimated 40–60% reduction in multi-jurisdictional compliance burden under their framework – a figure from the developers’ own validation work, not an independently confirmed benchmark, pending real-world adoption data.
- Build vendor due-diligence questions around tool-selection autonomy specifically: does the system independently decide which tool or system to call, and under what conditions does a human have to approve that choice before it executes?
- Track docket FDA-2026-N-7874 through its October 19, 2026 comment deadline. The evidence and evaluation approaches under discussion could influence how FDA approaches generative and agentic device classification going forward, though the paper itself commits to no specific timeline or outcome.
What to Watch Next
Three threads are worth following into 2027. First, whether FDA’s discussion paper converts into draft guidance, and whether it treats agentic systems as a distinct category or continues evaluating them function-by-function under the same lens as any other generative AI capability. Second, whether HAARF – or a comparable framework – gains traction as a de facto industry reference the way earlier voluntary frameworks have sometimes shaped binding rules. Third, whether any jurisdiction moves first on a genuinely agent-specific clearance pathway, which would put pressure on the others toward harmonization rather than the reverse.
The honest summary isn’t that regulators have no idea how to handle autonomous agents – FDA’s criteria-based CDS framework and its function-first approach to generative AI both still apply, and both were reaffirmed rather than abandoned in 2026. It’s that those frameworks were built around defined functions and conventional human review, and agentic systems raise a specific set of questions – about autonomy, multi-step action, tool use, and where the human checkpoint sits – that existing frameworks weren’t written to answer directly. FDA is now explicitly asking industry to help answer them, while acknowledging that some agentic healthcare functions may not require its oversight at all. That’s a narrower, more precise claim than “the rules don’t exist yet” – and it’s the more useful one for anyone deciding how to build or evaluate these tools today.
Sources
- FDA – Clinical Decision Support Software, Final Guidance (issued Jan. 6, 2026; reissued Jan. 29, 2026)
- FDA – Clinical Decision Support Software FAQs
- FDA – Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (Aug. 18, 2026)
- FDA – Press Announcement: FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices
- Regulations.gov – Docket FDA-2026-N-7874 (comments due Oct. 19, 2026)
- Task Force for AI Agents in Healthcare – HAARF: Healthcare AI Agents Regulatory Framework (medRxiv preprint, posted Apr. 10, 2026)
- HAARF – Source Code and Evaluation Harness (GitHub)
- FDA – 510(k) Summary, K253281: UpDoc, a Class II insulin-management SaMD pairing a conversational, LLM-based patient interface with dosing instructions computed from a clinician-defined treatment plan (cleared Dec. 23, 2025)
1. Academic policy review discussing FDA’s January 2026 CDS guidance and the absence, as of publication, of a cleared autonomous AI prescribing system: “The Clinician’s Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing” (arXiv preprint). Cited as an academic source, not a regulatory determination.
2. See K253281 510(k) summary above. UpDoc computes insulin dosing instructions from parameters a clinician configures; it does not independently select or initiate a treatment plan.




















