AI in Clinical Trials: FDA Strategies for Patient Selection and Predictive Modeling

Artificial Intelligence (AI) and machine learning (ML) are increasingly transforming clinical research by supporting patient identification, treatment-response prediction, trial optimization, and data-driven decision-making. As these technologies become more integrated into clinical development, sponsors must ensure that AI-supported approaches are scientifically justified, appropriately validated, and consistent with FDA expectations. The use of AI in Clinical Trial Design can provide important opportunities to improve trial efficiency, but it also introduces considerations related to data quality, model credibility, bias, transparency, and patient protection.

One of the most important applications is patient stratification in clinical trials. AI and ML models can evaluate demographic, clinical, laboratory, imaging, genomic, and other baseline characteristics to identify differences in prognosis or potential treatment response. These approaches may support enrichment strategies by helping sponsors identify populations that are more likely to benefit from an investigational intervention or that have specific characteristics relevant to the study objective. FDA has described the use of machine learning for identifying suitable patient populations as an emerging application within drug development.

For sponsors, the first step is to establish a clearly defined purpose for the model. An AI model should have a specific intended use, target population, and clinical or operational application. Sponsors should document why the model is appropriate for the proposed trial and explain how its output will influence participant selection, stratification, or other trial decisions. A clearly defined context of use is particularly important because the credibility requirements for an AI model depend on the decision it is intended to support.

Predictive modeling in clinical trials can also help sponsors identify prognostic factors and estimate the likelihood of specific clinical outcomes. Predictive models may contribute to enrichment strategies, reduce variability, support trial population selection, or improve the efficiency of clinical development. However, strong predictive performance alone does not establish clinical validity. Sponsors must demonstrate that the model is fit for its intended purpose and that its use does not introduce unacceptable risks or compromise the interpretability of trial results.

Data quality is fundamental to successful AI implementation. Clinical datasets used to develop and validate predictive models should be accurate, relevant, sufficiently representative, and appropriately controlled. Sponsors should evaluate missing data, inconsistencies, differences between data sources, and potential biases within development datasets. Data provenance and traceability should also be maintained so that the origin and processing of information can be demonstrated during regulatory review or inspection.

FDA’s emerging approach to AI emphasizes a risk-based assessment of model credibility. In its draft guidance on using AI to support regulatory decision-making for drugs and biological products, FDA proposes evaluating the credibility of an AI model according to its specific context of use (COU). The guidance is currently a draft and does not establish legally enforceable requirements, but it provides an important framework for sponsors considering AI-supported regulatory applications.

Model validation should therefore be proportionate to the potential impact of the AI-supported decision. Sponsors should establish appropriate training, testing, and validation procedures and evaluate model performance using datasets representative of the intended population. Performance should also be assessed across relevant demographic and clinical subgroups to identify potential disparities. Where appropriate, sensitivity analyses should be conducted to understand how changes in assumptions or inputs could affect model performance.

Patient protection and representativeness must remain central to AI-supported patient stratification in clinical trials. An algorithm that unintentionally excludes certain demographic or clinical groups could affect enrollment and reduce the generalizability of study findings. Sponsors should therefore assess whether AI-based eligibility or enrichment approaches could create unintended selection bias. FDA has emphasized the importance of appropriate eligibility criteria and enrollment practices to improve participation and ensure that clinical trial populations can adequately represent the patients who may ultimately use the product.

AI governance should also be integrated into the clinical quality management system. Sponsors should clearly define responsibilities for model development, validation, approval, monitoring, and change management. Clinical Operations, Biostatistics, Data Management, Regulatory Affairs, Quality Assurance, and Medical Affairs teams should collaborate to ensure that AI-supported processes remain controlled and documented. Any material modification to a model should undergo appropriate assessment to determine whether additional validation or regulatory evaluation is necessary.

Documentation is equally important. Protocols, Statistical Analysis Plans, model-development documentation, validation reports, and relevant regulatory submissions should clearly describe how AI or ML is being used. Sponsors should maintain records of datasets, assumptions, performance assessments, limitations, and model changes. This documentation supports transparency, reproducibility, inspection readiness, and regulatory confidence.

As AI technologies continue to evolve, their role in clinical research is likely to expand. However, technological sophistication does not replace the need for sound clinical science, statistical rigor, data integrity, and participant protection. The successful use of AI in Clinical Trial Design depends on integrating innovative technologies within established quality and regulatory frameworks.

Ultimately, AI-supported patient stratification in clinical trials and predictive modeling in clinical trials can help sponsors develop more targeted and efficient studies when appropriately designed and validated. By defining a clear context of use, maintaining reliable data, evaluating model credibility, monitoring potential bias, and implementing robust governance, sponsors can leverage AI while supporting scientifically credible evidence and meeting evolving FDA expectations.

Register now  to learn how AI and predictive modeling can support patient stratification, optimize clinical trial design, and align with evolving FDA expectations.

Frequently Asked Questions

Sponsors should define the model's context of use (COU) and establish a risk-proportionate credibility assessment. This should address the model's intended purpose, development data, performance characteristics, validation strategy, applicable limitations, and the potential impact of model error on trial participants or regulatory conclusions.

Sponsors should demonstrate that AI-based eligibility or enrollment approaches are scientifically justified and do not introduce inappropriate selection bias. The model should be evaluated for performance across relevant patient subgroups, and sponsors should assess whether its use could affect trial representativeness, generalizability, or the reliability of the resulting clinical evidence.

Sponsors should establish appropriate controls for data provenance, accuracy, completeness, relevance, traceability, and representativeness. Training and validation datasets should be sufficiently characterized, and sponsors should evaluate missing data, measurement variability, potential bias, and differences between development and intended-use populations.

AI model modifications should be governed through a documented change-control process. Sponsors should determine whether changes could affect model performance, patient selection, trial conduct, or the interpretation of clinical results. Depending on the nature and significance of the modification, additional verification, validation, documentation, or regulatory engagement may be warranted.

Sponsors should maintain documentation describing the model's intended use, development methodology, datasets, assumptions, performance evaluation, validation results, limitations, version history, change controls, and governance procedures. Where AI influences trial design or patient selection, the relevant methodology should also be appropriately reflected in the protocol and statistical documentation to support transparency and regulatory review.