Agentic AI Regulatory Gap

by | Sep 19, 2026 | healthcare AI performance | 0 comments

Agentic AI is already operating inside U.S. healthcare workflows – selecting tools, querying records, and in some cases recommending or initiating clinical actions – faster than any regulator has finished deciding how to evaluate it. Three developments from 2026 make that gap visible rather than theoretical. In January, FDA revised its guidance on the statutory line between clinical decision support software that falls outside its device jurisdiction and software that remains subject to FDA oversight – a function-by-function, criteria-based line that gets genuinely hard to apply once a system acts autonomously across multiple tools rather than issuing one recommendation for review. In April, a 40-plus-member multi-stakeholder task force published HAARF (the Healthcare AI Agents Regulatory Framework), a preprint arguing that none of the nine major regulatory frameworks in use today – including FDA’s own – were built with tool-selecting, tool-chaining agents in mind. And in August, FDA opened a formal comment period asking, in 26 numbered questions, how agentic systems should be evaluated. This piece walks through what each of these actually establishes, what HAARF proposes, and where the genuinely open questions sit for anyone building or deploying agentic tools in clinical settings.

 

Disclaimer: This post is for informational purposes only and does not constitute legal or regulatory advice. Regulatory guidance and proposed frameworks referenced here – particularly FDA’s open comment period and any non-final guidance – remain subject to change. Organizations evaluating agentic AI tools for clinical use should consult qualified regulatory counsel before relying on any framework, including HAARF, as a compliance strategy.

 
 

The FDA’s Line Is Clearer – and More Specific Than It First Appears

On January 6, 2026, FDA issued revised Clinical Decision Support Software guidance, reissued in final form on January 29, 2026. The guidance interprets the four statutory criteria under Section 520(o)(1)(E) of the FD&C Act – added by the 21st Century Cures Act – that together define which CDS software functions are excluded from the device definition (“Non-Device CDS”). To qualify, a function must meet all four: it must not acquire, process, or analyze a medical image, a signal from an in vitro diagnostic device, or a pattern from a signal-acquisition system (Criterion 1); it must display, analyze, or print medical information (Criterion 2); it must support or provide recommendations to a clinician without being intended to replace or direct that clinician’s judgment (Criterion 3); and it must enable the clinician to independently review the basis for the recommendation rather than relying on it primarily to make the clinical decision (Criterion 4).

That’s a function-specific, criteria-based test – not a blanket rule about autonomy or “agents” as a category. The guidance itself doesn’t define either term. What it does say is that software providing a specific diagnostic or treatment directive, rather than an option a clinician can weigh, fails Criterion 3 regardless of what’s running underneath it. That’s the practical hook for agentic systems: an agent that autonomously executes an action rather than surfacing a recommendation for review is, by definition, offering the kind of directive output Criterion 3 excludes from Non-Device status. But whether a given agentic function meets that bar still depends on the specific function and its intended use – FDA’s August discussion paper (below) makes the same point explicitly for generative AI more broadly.

 

What This Guidance Does and Doesn’t Settle

  • Settled: A CDS function that meets all four statutory criteria – including inputs limited to non-image, non-signal medical information and outputs a clinician can independently evaluate – can qualify for Non-Device status.
  • Settled, but narrower than it sounds: FDA also intends to exercise enforcement discretion for CDS functions that present only one clinically appropriate recommendation and otherwise meet the other criteria. That’s a distinct, narrower carve-out layered on top of the statutory Non-Device exclusion – not a description of the exclusion itself.
  • Settled: Software that acquires, processes, or analyzes a medical image, an IVD signal, or a physiological pattern remains a device outright under Criterion 1, regardless of how many recommendations it offers or how transparent its logic is.
  • Not settled: The guidance doesn’t use “autonomous agent” as a defined term, and doesn’t address, one way or the other, how FDA would evaluate a system that plans and executes a sequence of actions across multiple tools rather than producing one output for review.
  • Not settled: The sources reviewed for this piece don’t identify an FDA-cleared system that independently selects and prescribes medication without a clinician-defined treatment plan. FDA has cleared autonomous diagnostic functions before – a fully autonomous diabetic retinopathy screening tool cleared in 2018 remains the best-known example – and cleared a narrow, protocol-bound generative-AI-enabled function as recently as December 2025.1,2 Neither is the same as an agent independently deciding what to prescribe.

 
 

FDA’s Own August Admission: The Agentic Question Is Still Open

The clearest evidence that this is a genuinely unsettled frontier didn’t come from outside critics – it came from FDA itself. On August 18, 2026, the Digital Health Center of Excellence released Considerations for the Regulation of Generative AI-Enabled Medical Devices, a discussion paper and formal request for feedback under docket FDA-2026-N-7874. Comments are open through October 19, 2026.

The paper is explicit that it is not draft or final guidance and proposes no binding policy changes. What it does is lay out a proposed two-axis risk framework – one axis for how independently a function acts (from passive information through supervised action to fully autonomous action), the other for how severe the consequence of a wrong output would be – alongside a proposed competency-based evaluation model built around device benchmarking, clinical confirmation, and postmarket monitoring, echoing how physicians are credentialed rather than how traditional software is cleared. FDA frames the underlying principle plainly: it regulates device functions, not AI models as such – so a given generative AI capability’s device status still turns on what that specific function does and how it’s used.

The paper directly addresses agentic AI systems that autonomously plan and execute multi-step tasks or invoke external tools. FDA notes that some agentic functions – care coordination, documentation, patient outreach – may not fall within its device oversight at all, while others could qualify as devices depending on what they do. The agency does not propose a distinct regulatory framework for agentic systems in this paper; instead, it acknowledges that agentic architectures may raise unique risks and poses the question back to industry: should additional premarket and postmarket evaluation considerations apply, and if so, what should they look like? One proposed benchmarking element (A.1, “Agentic competencies”) would probe an agent’s planning behavior, tool-error recognition, and whether it pauses for a human checkpoint before irreversible actions – a sign of where FDA’s attention is, even without a settled answer.

 
 

HAARF: A Coalition Steps Into the Vacuum

Four months before FDA’s discussion paper, a separate effort had already reached a similar diagnosis from the outside. On April 10, 2026, the Task Force for AI Agents in Healthcare – a coalition of 40-plus experts spanning regulatory affairs, clinical medicine, and AI security – published HAARF: the Healthcare AI Agents Regulatory Framework as a preprint on medRxiv.

HAARF’s starting premise is that existing frameworks were built for prediction tools a human interprets, not for agents that act. It synthesizes requirements from nine major frameworks – FDA’s Total Product Lifecycle/PCCP approach, the EU AI Act, Health Canada’s SGBA+ requirements, UK MHRA’s AI Airlock, NIST’s AI RMF, WHO’s GI-AI4H ethics guidelines, ISO/IEC 42001, OWASP’s AI Security Verification Standard, and IMDRF’s Good Machine Learning Practice – into eight verification categories comprising 279 requirements across three risk-based implementation levels (Foundation, Advanced, and Expert). It’s released under a Creative Commons license, with its evaluation code publicly available on GitHub.

It’s worth being precise about what HAARF is and isn’t. It’s a preprint – not a peer-reviewed, finalized, or formally adopted standard – and no regulator has adopted it. It functions as a voluntary, multi-stakeholder governance proposal, closer to an industry consensus checklist than to binding law, though its authors report direct engagement with FDA industry-committee stakeholders during development, and its newest category was reportedly built around FDA-identified priorities.

 

Three Gaps, Named

HAARF’s authors frame their contribution around three specific gaps they argue no existing framework closes:

  • Autonomy Governance Gap – existing frameworks lack a standardized way to manage progressively higher levels of AI agent autonomy, from a co-pilot that only suggests, to a system that acts independently within defined bounds.
  • Tool Integration Gap – there is no defined regulatory pathway for agents that autonomously select and use multiple healthcare tools and systems (EHRs, order sets, medical devices) in sequence, rather than producing a single output for review.
  • Multi-Jurisdictional Harmonization Gap – organizations deploying agentic AI across the US, EU, UK, and Canada face materially different, and sometimes conflicting, compliance requirements, with no single pathway that satisfies all of them at once.

These aren’t abstract categories. FDA’s own August discussion paper addresses overlapping territory: its agentic-competency benchmarking element and its questions about machine-based supervisory agents speak directly to the Autonomy Governance and Tool Integration gaps HAARF names, even as the agency stops short of proposing how to close them. Nothing in the public record indicates the two efforts were developed in coordination – the overlap is notable on its own.

 
 

Why the Tool Integration Gap Isn’t Theoretical

HAARF’s authors backed their framework with an adversarial red-team evaluation designed to test whether the gap they identified translates into real failure modes. Across six scenarios – including an agent asked to order a medication outside its permitted scope, and an agent prompted to order a drug that contradicted a documented patient allergy – they ran 50 trials per scenario against a baseline agent with no governance middleware, and again with HAARF’s rules-based enforcement layer active.

Without governance controls, the agent executed unauthorized or contraindicated tool calls in 56–60% of adversarial trials. With HAARF’s middleware enforcing role-based tool access and contraindication checks, that rate dropped to zero across 600 primary trials, with results holding under cross-model validation. This was a scenario-based red-team exercise designed and run by the framework’s own authors, not an independent clinical deployment study, so the specific percentages shouldn’t be read as real-world failure rates. What it does offer is an initial, controlled illustration of the underlying concern: an agent given tool access can use it in unsafe ways at a meaningful rate absent explicit governance controls – which is the practical case for why the tool integration gap is worth taking seriously now rather than treating as hypothetical.

 
 

Key Considerations

 

Disclaimer: This checklist is provided for general informational purposes only and does not constitute legal, regulatory, or professional advice; organizations should consult with their legal and compliance departments to ensure adherence to specific jurisdictional requirements.

 

For Founders and AI Developers

  • If your product autonomously selects among multiple tools, systems, or data sources – rather than producing one recommendation for clinician review – you’re outside the narrow scenarios where the January 2026 CDS guidance’s Non-Device exclusion or enforcement discretion are likely to apply.
  • FDA’s August discussion paper is a live opportunity, not background reading. If agentic architecture is core to your product, commenting on docket FDA-2026-N-7874 before October 19, 2026 is a direct way to shape how “agentic conduct” gets evaluated before the rules solidify.
  • Frameworks like HAARF – currently a preprint proposal, not an adopted standard – can function as a practical governance checklist while formal agentic-specific guidance is still being written: tool-access controls, contraindication gates, audit logging. They are not a substitute for an actual FDA clearance pathway or legal sign-off.
  • Don’t assume “clinician-supervised” and “autonomous” are interchangeable in a regulatory filing. FDA’s own framing treats them as different risk placements with different evidence expectations, and December 2025’s UpDoc clearance shows how narrow a currently-cleared “agentic-flavored” function actually is (see below).

 

For Compliance Specialists and Health System Leaders

  • The public FDA record reviewed for this piece doesn’t show a cleared system that independently selects and prescribes medication without a clinician-defined treatment plan. Any vendor describing an “autonomous” or “agentic” clinical tool should be asked precisely what decisions the system makes without clinician sign-off, and what clearance, if any, covers that specific function.
  • Multi-jurisdictional deployments genuinely face conflicting requirements right now, not just added paperwork. HAARF’s authors report an estimated 40–60% reduction in multi-jurisdictional compliance burden under their framework – a figure from the developers’ own validation work, not an independently confirmed benchmark, pending real-world adoption data.
  • Build vendor due-diligence questions around tool-selection autonomy specifically: does the system independently decide which tool or system to call, and under what conditions does a human have to approve that choice before it executes?
  • Track docket FDA-2026-N-7874 through its October 19, 2026 comment deadline. The evidence and evaluation approaches under discussion could influence how FDA approaches generative and agentic device classification going forward, though the paper itself commits to no specific timeline or outcome.

 
 

What to Watch Next

Three threads are worth following into 2027. First, whether FDA’s discussion paper converts into draft guidance, and whether it treats agentic systems as a distinct category or continues evaluating them function-by-function under the same lens as any other generative AI capability. Second, whether HAARF – or a comparable framework – gains traction as a de facto industry reference the way earlier voluntary frameworks have sometimes shaped binding rules. Third, whether any jurisdiction moves first on a genuinely agent-specific clearance pathway, which would put pressure on the others toward harmonization rather than the reverse.

The honest summary isn’t that regulators have no idea how to handle autonomous agents – FDA’s criteria-based CDS framework and its function-first approach to generative AI both still apply, and both were reaffirmed rather than abandoned in 2026. It’s that those frameworks were built around defined functions and conventional human review, and agentic systems raise a specific set of questions – about autonomy, multi-step action, tool use, and where the human checkpoint sits – that existing frameworks weren’t written to answer directly. FDA is now explicitly asking industry to help answer them, while acknowledging that some agentic healthcare functions may not require its oversight at all. That’s a narrower, more precise claim than “the rules don’t exist yet” – and it’s the more useful one for anyone deciding how to build or evaluate these tools today.

 
 

Sources

1. Academic policy review discussing FDA’s January 2026 CDS guidance and the absence, as of publication, of a cleared autonomous AI prescribing system: “The Clinician’s Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing” (arXiv preprint). Cited as an academic source, not a regulatory determination.
2. See K253281 510(k) summary above. UpDoc computes insulin dosing instructions from parameters a clinician configures; it does not independently select or initiate a treatment plan.

Written by Grigorii Kochetov

Cybersecurity Researcher at AI Healthcare Compliance

Read more

Monthly News and Updates (August 2026)

Monthly News and Updates (August 2026)

August 2026 was a quieter month for Canadian AI regulatory activity, but the United States and the European Union both produced developments with direct implications for AI in healthcare. In the US, the FDA opened a public comment process on how it might regulate...

read more
Monthly News and Updates (July 2026)

Monthly News and Updates (July 2026)

In July 2026, AI-in-healthcare regulatory activity centered on clarifying how existing frameworks apply to fast-moving AI tools rather than creating new ones outright. Health Canada issued guidance on AI-generated content in device submissions, HHS committed...

read more
Monthly News and Updates (June 2026)

Monthly News and Updates (June 2026)

During June, 2026, AI-in-healthcare governance advanced on multiple fronts across our tracked jurisdictions. Canada named health and life sciences the first priority sector of its new national AI strategy and committed $300 million in combined health-data...

read more
Monthly News and Updates (May 2026)

Monthly News and Updates (May 2026)

May 2026 was a high-density month for AI healthcare regulation, with significant activity across Canada, the UK, and the EU. Canada published new federal guidance addressing the governance of agentic AI systems, including considerations that may be relevant to federal...

read more
Monthly News and Updates (April 2026)

Monthly News and Updates (April 2026)

April 2026 saw meaningful regulatory movement across several major jurisdictions. While activity in Canada and the United States was relatively modest compared to previous months - each producing one notable update - the EU and UK were particularly active, with a...

read more
Monthly News and Updates (March 2026)

Monthly News and Updates (March 2026)

March 2026 did not introduce any major or immediately actionable regulatory changes for AI in healthcare across Canada, the European Union, or the United Kingdom. Activity in these regions remained relatively stable, with no significant new guidance, enforcement...

read more
Practical impacts of using AI in Healthcare

Practical impacts of using AI in Healthcare

Artificial Intelligence (AI) is transforming healthcare systems globally - enhancing diagnostics, improving patient outcomes, optimizing workflows, and reducing costs. However, its adoption also brings challenges around data integrity, equity, and ethical use. Below...

read more
Monthly News and Updates (February 2026)

Monthly News and Updates (February 2026)

During February 2026, governments and regulators across Canada, the United States, and Europe advanced regulatory and governance measures directly affecting AI in healthcare. Key themes included quality system harmonisation, acceleration pathways for digital health...

read more
Monthly News and Updates (January 2026)

Monthly News and Updates (January 2026)

Editorial Update: Moving to a Monthly Schedule   To ensure we provide the most robust and actionable compliance intelligence for the healthcare AI sector, we are transitioning from weekly to monthly updates. This allows us to focus on high-impact regulatory...

read more
Weekly News and Updates (Jan 12-16, 2026)

Weekly News and Updates (Jan 12-16, 2026)

This week (January 12–16, 2026) marked a pivotal shift in AI healthcare regulation globally, characterized by the formalization of oversight and international harmonization. Key highlights include the joint FDA-EMA guiding principles for AI in drug development,...

read more
Weekly News and Updates (Jan 1-9, 2026)

Weekly News and Updates (Jan 1-9, 2026)

Between 1st and 9th January 2026, the first full week of the year marks a significant shift from theoretical frameworks to operational infrastructure in AI healthcare governance. Key developments include the UK’s closing of its “AI Growth Lab” consultation, the FDA’s...

read more
Weekly News and Updates (Dec 12 – 19, 2025)

Weekly News and Updates (Dec 12 – 19, 2025)

Between 12–19 December 2025, the regulatory landscape for AI in healthcare shifted decisively toward national-level consolidation and operational security: the U.S. White House issued a landmark Executive Order to centralize AI policy and preempt state-level...

read more
Weekly News and Updates (Nov 22 – 28, 2025)

Weekly News and Updates (Nov 22 – 28, 2025)

Between 22–28 November 2025, global regulators accelerated the shift from high-level principles to mandatory operational controls, particularly in Canada, which launched its first public AI Register detailing hundreds of government AI systems. The EU continued...

read more
Weekly News and Updates (Nov 8 – 21, 2025)

Weekly News and Updates (Nov 8 – 21, 2025)

Between 8-21 November 2025 regulators and international bodies emphasised moving from principles to practice: the EU launched COMPASS-AI to operationalise safe clinical AI; the UK (MHRA) published AI Airlock pilot outputs and announced AI drug-safety projects; the FDA...

read more
Prohibited AI Systems Under the EU AI Act

Prohibited AI Systems Under the EU AI Act

The European Union’s Artificial Intelligence Act (EU AI Act) establishes the world’s first comprehensive legal framework for governing artificial intelligence. It divides AI systems into four categories based on their potential impact on safety and fundamental rights...

read more