In Silico Evidence in Clinical Trials: Understanding FDA’s Regulatory Expectations
The use of artificial intelligence, computational modeling, and advanced data-generation techniques is reshaping clinical research. Among these developments, synthetic data is emerging as a potentially valuable tool for clinical development, particularly when traditional data sources are limited or difficult to obtain.For sponsors, however, the key question is not whether synthetic data is innovative, but whether it can generate credible, fit-for-purpose evidence that supports a regulatory decision.FDA has shown increasing interest in innovative approaches to evidence generation, including model-informed drug development, artificial intelligence, real-world evidence, and new approach methodologies. This evolving landscape creates opportunities for sponsors, but it also raises important questions about validation, reliability, transparency, and regulatory acceptance.
Synthetic data is artificially generated data designed to reproduce relevant characteristics and relationships found in real-world or clinical datasets. It can be generated using statistical techniques, computational models, simulations, or artificial intelligence.The potential applications are broad. Synthetic data may support algorithm development, analytical testing, trial simulations, modeling, and other research activities. In certain situations, it may also provide supplementary evidence when access to adequate patient-level data is challenging.
However, synthetic data should not automatically be considered equivalent to data collected directly from clinical trial participants. Its value depends on the quality of the underlying information, the methodology used to generate the data, and the evidence demonstrating that the resulting dataset is appropriate for its intended purpose.
FDA’s Focus on Scientific Credibility
FDA’s approach to innovative evidence generation is increasingly focused on whether a methodology is scientifically credible and appropriate for the regulatory question being addressed.In June 2026, FDA finalized ICH M15, which provides general principles for planning, evaluating, documenting, and communicating evidence generated through model-informed drug development. The framework reflects the importance of understanding how models are developed, evaluated, and used in regulatory decision-making.FDA has also been developing recommendations for the use of new approach methodologies and AI-related technologies. These initiatives indicate that innovative computational approaches can have a role in modern drug development, provided that their scientific reliability can be demonstrated.
For sponsors exploring synthetic data in clinical trials, this means the methodology should be considered as part of an evidence-generation strategy rather than simply as a technology solution.
When Can Synthetic Data Be Useful?
The usefulness of synthetic data depends largely on its intended application.For example, synthetic datasets may be useful for developing or testing computational algorithms, evaluating analytical methods, conducting simulations, or exploring different clinical trial scenarios. They may also help researchers investigate specific questions when appropriate source data are available for model development.The regulatory expectations become more significant when synthetic data are intended to support conclusions about safety or effectiveness. Sponsors should clearly define what question the synthetic data are intended to answer and determine whether the methodology is capable of generating sufficiently reliable evidence for that purpose.
Data Quality and Representativeness Matter
One of the biggest challenges with synthetic data is ensuring that it accurately reflects the characteristics that matter clinically.A dataset may statistically resemble its source data while failing to reproduce important clinical relationships or population characteristics. If the source data contain bias or limitations, those issues may also affect the generated dataset.Sponsors should therefore evaluate the representativeness of the data, relevant patient characteristics, important clinical variables, missing information, potential bias, and uncertainty. The goal should not simply be to create data that look realistic. The objective is to demonstrate that the data are scientifically appropriate for the intended use.
Validation Is Central to Regulatory Acceptance
Validation is one of the most important considerations when developing in silico evidence for FDA review.Sponsors should establish appropriate criteria for assessing whether the methodology performs as expected. Validation may involve comparing synthetic outputs with relevant real-world observations, evaluating model performance, assessing reproducibility, and examining whether important relationships are preserved.FDA’s work involving AI-enabled technologies similarly emphasizes the importance of understanding how synthetic data are generated and demonstrating that they are fit for purpose. A sophisticated artificial intelligence system is not automatically a scientifically valid evidence-generation tool. Its performance must be evaluated in the context in which it will be used.
Synthetic Data Is Not a Universal Replacement for Clinical Evidence
There is an important distinction between supporting clinical development and replacing clinical evidence.Synthetic data may provide valuable supplementary information, but sponsors should not assume that it can eliminate the need for appropriate clinical studies.The strength of evidence required depends on the regulatory question. A methodology that is appropriate for trial simulation or algorithm development may not be sufficient to support a primary conclusion about a product’s safety or effectiveness. This fit-for-purpose principle should guide the design and interpretation of synthetic evidence.
Considerations for Rare and Difficult-to-Study Populations
Synthetic and computational approaches may be particularly interesting when clinical research involves small or difficult-to-recruit populations.Rare diseases, for example, may involve limited numbers of eligible patients, making conventional evidence generation challenging. Innovative approaches could potentially provide additional information or support more efficient trial planning.However, the use of synthetic data in these situations still requires careful scientific justification. A small patient population does not automatically justify replacing observed clinical evidence with generated data.
What Should Sponsors Do Before Using Synthetic Data?
Sponsors considering FDA synthetic data guidance should begin by defining the regulatory purpose of the proposed approach.The development plan should address the source of the underlying data, methodology, assumptions, validation strategy, intended use, limitations, and relationship to other evidence.
Early regulatory engagement may also be valuable when the approach is novel or could significantly influence the clinical development program. Discussing the methodology before generating extensive evidence can help sponsors identify potential regulatory concerns and establish an appropriate development strategy.
The Future of In Silico Evidence
The role of computational evidence in drug development is likely to expand as technology and regulatory science continue to evolve.FDA’s work across model-informed drug development, artificial intelligence, new approach methodologies, and other innovative approaches suggests that the agency is increasingly interested in evidence-generation methods that can provide meaningful, scientifically reliable information. The future is unlikely to be about choosing between traditional clinical evidence and synthetic data. Instead, successful development programs may increasingly integrate multiple evidence sources while maintaining appropriate standards for scientific validity.
Conclusion
Synthetic data has the potential to become an important component of modern clinical development. It can support research, modeling, simulation, and potentially broader evidence-generation strategies when appropriately designed. However, regulatory acceptance will depend on scientific credibility rather than technological sophistication. Sponsors must demonstrate that their methodology is validated, transparent, reliable, and fit for the intended regulatory purpose.Understanding FDA’s evolving position is therefore essential for organizations considering synthetic or in silico approaches in clinical development.
Frequently Asked Questions
FDA's assessment will depend on the intended use of the synthetic data, the quality and provenance of the source data, the methodology used to generate it, validation results, representativeness, potential bias, and the overall reliability of the evidence.
Not automatically. Synthetic data may supplement clinical evidence or support specific development activities, but its ability to replace observed clinical data depends on the regulatory question, scientific justification, and demonstrated fitness for purpose.
Sponsors should establish predefined validation criteria appropriate to the intended use. Validation may include comparison with relevant observed data, assessment of preserved clinical relationships, evaluation of model performance, reproducibility, bias, and uncertainty.
Sponsors should document the source and characteristics of the underlying data, data-generation methodology, assumptions, model development, validation procedures, performance results, limitations, and intended regulatory use. This documentation should allow regulators to understand and assess the credibility of the evidence.
For novel approaches that could materially influence clinical development or regulatory decision-making, early interaction with FDA can be valuable. Sponsors can use appropriate regulatory engagement to discuss the proposed methodology, intended use, validation strategy, and evidence needed to support its application.