From Algorithms to Evidence: FDA Considerations for AI in Clinical Development
Artificial intelligence is moving from an emerging technology to a practical component of modern clinical development. Pharmaceutical and biotechnology companies are exploring machine learning for patient identification, trial enrichment, treatment-response prediction, safety assessment, endpoint analysis, and other development activities. Medical device manufacturers and clinical research organizations are also evaluating AI-enabled tools across clinical investigations. As these applications expand, FDA’s focus is increasingly centered on whether AI-generated information is reliable, appropriately validated, and suitable for its intended regulatory purpose. For sponsors, this means that AI in clinical trials should not be evaluated solely on technical accuracy. A model can demonstrate impressive predictive performance while still creating regulatory concerns if its intended use is unclear, its training data are poorly characterized, or its performance does not translate to the population in which it will be used. FDA’s January 2025 draft guidance proposes a risk-based credibility assessment framework for AI models used to generate information supporting regulatory decisions involving the safety, effectiveness, or quality of drugs and biological products.
One of the most important concepts in this framework is the context of use. Sponsors should clearly establish what question the model is intended to answer, how its output will be used, and what regulatory decision could be influenced by that output. An AI model used for exploratory subgroup analysis presents a different level of regulatory concern from a model used to determine eligibility for a pivotal study. Defining the context of use early allows sponsors to establish an appropriate evidence and validation strategy rather than attempting to justify the technology after it has already been incorporated into a clinical program.
Patient stratification illustrates why this distinction matters. AI can analyze multiple baseline characteristics simultaneously and identify patterns that may be difficult to detect through conventional approaches. Sponsors may use these capabilities to identify potentially responsive populations, support enrichment strategies, or understand differences in clinical outcomes. However, algorithmic selection can also introduce unintended bias or reduce the representativeness of a study population. Sponsors therefore need to demonstrate that the variables used by the model are scientifically justified and that the resulting population remains relevant to the intended indication. The credibility of predictive modeling in clinical research also depends heavily on data quality. Training and validation datasets should be sufficiently representative of the intended population and clinical environment. Sponsors should evaluate missing information, inconsistent measurements, differences between sites, endpoint definitions, data leakage, overfitting, and potential demographic or clinical biases. Performance should not be assessed only against an overall dataset; relevant patient subgroups may require separate evaluation to determine whether the model performs consistently across the population.
This makes AI model validation more than a statistical exercise. The validation strategy should be connected to the model’s intended clinical and regulatory application. Depending on the use case, relevant assessments may include discrimination, calibration, sensitivity, specificity, predictive values, robustness, and subgroup performance. Sponsors should also establish meaningful thresholds for acceptable performance before relying on model outputs in important trial decisions. The level of evidence should be proportionate to the consequences of an incorrect prediction. Another consideration is how AI interacts with the broader clinical trial design. If an algorithm influences enrollment, stratification, treatment assignment, monitoring, or endpoint assessment, sponsors should evaluate how that use affects the protocol, statistical analysis plan, trial population, and interpretation of results. AI should complement—not automatically replace—established statistical principles. Any model-dependent decision should be appropriately documented and integrated into the overall scientific rationale for the study.
FDA’s January 2026 Guiding Principles of Good AI Practice in Drug Development reinforce this lifecycle perspective. Developed with the European Medicines Agency, the principles emphasize human-centric design, a risk-based approach, clear context of use, multidisciplinary expertise, data governance, model development practices, risk-based performance assessment, lifecycle management, and clear information. These principles provide an important direction for organizations developing FDA AI compliance strategies. Documentation is therefore becoming a critical part of AI regulatory strategy. Sponsors should maintain records describing data sources, model architecture or methodology, development procedures, assumptions, intended population, validation activities, performance limitations, and version history. If a model is updated during development, the sponsor should assess whether the change could affect participant selection, trial conduct, safety conclusions, statistical analyses, or interpretation of previously generated evidence. A defensible documentation trail can make it easier to demonstrate how AI-related decisions were controlled throughout the development lifecycle.
The expectations are also relevant to AI-enabled medical devices. FDA’s January 2025 draft guidance for AI-enabled device software functions addresses lifecycle management, documentation, design, development, and risk considerations, while the Agency’s August 2025 final guidance provides recommendations for predetermined change control plans for AI-enabled device software functions. These approaches recognize that AI-enabled products may evolve over time and that planned modifications should be appropriately controlled to maintain reasonable assurance of safety and effectiveness. For clinical research organizations, the growing use of AI creates additional responsibilities around vendor qualification, data integrity, validation, documentation, and oversight. Sponsors should understand whether a third-party AI platform is being used merely as an operational tool or whether its outputs could become part of evidence submitted to FDA. The distinction can materially affect the level of oversight required.
The primary compliance risks are not limited to algorithmic errors. They can include unclear intended use, inadequate data governance, unrepresentative training populations, insufficient validation, undocumented model changes, poor transparency, inappropriate patient exclusion, and failure to evaluate performance across relevant subgroups. These issues can weaken confidence in AI-generated evidence and potentially complicate FDA review. Early regulatory engagement can help sponsors address these questions before an AI model becomes deeply embedded in a clinical development program. When an AI application is expected to materially influence trial design, patient selection, analysis, or evidence supporting regulatory decision-making, sponsors should be prepared to explain the model’s purpose, credibility assessment, limitations, validation strategy, and risk controls.
Frequently Asked Questions
FDA expects sponsors to establish that AI-generated information is credible and appropriate for its specific context of use, with evidence and controls proportionate to the associated risk.
Context of use defines the specific question the model addresses and how its output will influence a development or regulatory decision. It provides the foundation for determining appropriate credibility and validation activities.
Sponsors should evaluate scientific justification, data representativeness, subgroup performance, bias, generalizability, and the potential effect of AI-driven selection on trial validity and population relevance.
Model changes should be documented, assessed for their potential impact, and managed through an appropriate lifecycle process. For AI-enabled devices, FDA has specifically addressed predetermined change control plans for planned modifications.
Early engagement can be valuable when AI has a meaningful influence on clinical development or evidence intended to support regulatory decision-making, particularly when the approach is novel or has significant consequences for participants or study interpretation.