Traditional software validation frameworks are proving ill-equipped for the fluid, unpredictable nature of generative artificial intelligence in healthcare. As medical device developers navigate open-ended prompts, hallucination risks, and third-party foundation models, the FDA’s Center for Devices and Radiological Health (CDRH) has released a landmark discussion paper outlining a dedicated regulatory approach for GenAI-enabled devices. While this paper does not represent formal binding guidance, it signals an immediate mental shift within review divisions evaluating novel digital health submissions. Smart development teams must look beyond current static Software as a Medical Device (SaMD) rules and map their clinical strategies directly to these emerging expectations. By understanding this evolving philosophy today, regulatory leaders can de-risk their current pipelines and actively shape the future of medical AI policy.
While CDRH has cleared hundreds of traditional AI/ML software functions, GenAI presents an entirely different challenge. Traditional SaMD relies on deterministic or narrow algorithms with bounded inputs, predictable output structures, and explicit change control pathways. In contrast, GenAI relies on open-ended inputs, produces variable outputs, and generates emergent, multi-turn behaviors. Furthermore, GenAI devices carry severe risks of hallucination or confabulation, performance drift over long exchanges, and deep dependencies on third-party foundation models whose training data and inner mechanics are often opaque. Traditional unit testing and bounded verification methodologies simply cannot evaluate every conceivable prompt or output trajectory. Recognizing this shift, the FDA’s discussion paper outlines early considerations for a regulatory structure tailored directly to GenAI.
The Two-Axis Risk Assessment Framework: To address the fluid nature of GenAI, the FDA proposes evaluating device risk along two primary axes:
- Device Activity (Horizontal Axis): Evaluates how independently the device operates. This ranges from Informational Non-Directive (providing raw context or risk scores) to Informational Action-Directing (recommending specific clinical decisions), Action-Taking with HCP Supervision, and Fully Autonomous Action-Taking.
- Consequences of Error (Vertical Axis): Evaluates the severity of potential harm if an incorrect output is acted upon. For instance, an incorrect over-the-counter recommendation sits low on the risk curve, whereas an erroneous dosage alteration for acute care presents high severity.
The agency also introduces critical nuances within this framework. For example, “action-directing” text cannot be masked merely by placing disclaimers like “consult your physician” alongside a specific recommendation. Furthermore, patient-facing tools generally carry higher risk than HCP-facing tools because patients lack the clinical domain knowledge required to critically assess and catch a hallucinated output.
Because exhaustive input/output testing is impractical for GenAI, CDRH proposes shifting to a Competency-Based Evaluation model, conceptually borrowing from how human clinicians are credentialed through licensing exams and supervised clinical residency. Rather than evaluating an isolated algorithm, CDRH proposes assessing the fully configured, user-facing system through two main stages:
- Device Benchmarking: Scalable, non-clinical evaluation using standardized or specialized datasets across core competencies. These cover Safety (safety-critical recognition, scope maintenance/refusal, and uncertainty calibration), Clinical Proficiency (guideline fidelity, multi-turn information gathering, and numerical accuracy), Generalizability (robustness against adversarial prompt injection and demographic subgroup consistency), and Agentic AI Capabilities (tool usage and multi-step execution).
- Clinical Confirmation: Real-world performance validation tailored to device risk. Rather than mandating full prospective clinical trials for every device, sponsors could utilize shadow deployments, retrospective evaluation against expert panels, or standardized patient actor interactions to confirm clinical utility.
The paper also introduces two major policy suggestions:
- Foundation Model Master Files (MAFs): Device sponsors often build applications on top of third-party models (e.g., OpenAI, Anthropic, Google). CDRH proposes leveraging voluntary Master Files where model developers submit confidential system cards, architecture details, and guardrail performance directly to the FDA. Device developers could then reference these MAFs in their submissions without requiring the model vendor to expose proprietary intellectual property to the public.
- Accepting Premarket Uncertainty via Continuous Postmarket Monitoring: CDRH suggests it may be open to accepting greater premarket evidentiary uncertainty if backed by robust postmarket oversight. This includes periodic re-benchmarking, sample-based clinician reviews of live logs, performance degradation monitoring, and updating Predetermined Change Control Plans (PCCPs) to manage passive model updates.
This discussion paper represents a rare strategic window for medtech executives to participate in formal policy creation before draft guidance solidifies into law. Because FDA review teams are already applying these core concepts to evaluate ongoing GenAI submissions, alignment with this paper provides immediate utility for active programs. By pre-emptively structuring premarket dossiers around competency benchmarking and robust postmarket monitoring, developers can preempt extensive additional information queries during pre-submission and review cycles. Furthermore, engaging with the FDA’s public docket before the comment window closes empowers sponsors to help establish practical standards for foundation model integration and change control. Embracing this competency-based mindset today ensures your organization remains a proactive leader in the rapidly shifting landscape of generative AI regulation.