How Large Language Models Uncover Hidden Changes in Clinical Trials 

In the world of medical research, prespecified primary outcomes are supposed to be the immovable target that ensures scientific honesty. Yet, historical monitoring of thousands of clinical trials has long been bogged down by tedious manual curation. Enter the power of artificial intelligence, bringing unprecedented scale to the oversight of clinical evidence. A groundbreaking new study reveals that nearly one-third of clinical trials quietly alter their primary outcomes after participant enrollment begins.  

When researchers design a clinical trial, the bedrock of scientific integrity relies on prespecifying primary outcomes before a single participant walks through the door. This golden rule of clinical research prevents selective reporting and ensures that investigators cannot reshape their hypotheses after seeing the data. However, tracking whether trials strictly adhere to their original blueprints across tens of thousands of public records has traditionally been an impossible task for human reviewers alone. A recent landmark study published in JAMA Network Open changes the game by leveraging advanced large language models (LLMs), specifically OpenAI’s GPT-5 mini, to successfully analyze over 15,000 interventional clinical trials registered on ClinicalTrials.gov between 2015 and 2024. The results provide a breathtaking look at how often clinical trial goals shift after the starting gun has sounded.  

The findings reveal that an estimated 30.8% of clinical trials experienced at least one meaningful primary outcome change after enrollment began. When breaking down the specific types of alterations, time frame modifications reigned supreme at 17.8%, followed by the addition of new primary outcomes at 12.5%, and the complete removal of prespecified outcomes at 8.5%. While a ~31% modification rate sounds alarming, the study uncovered a major silver lining: a sharp, encouraging downward trend over time. Back in 2016, nearly half (48.9%) of trials showed primary outcome changes, but by 2024, that number plummeted to just 8.0%. Statistical models confirmed this steady improvement, showing a decreasing odds ratio of outcome changes with each passing year.  

Not all clinical trials are created equal, as the multivariable analysis exposed distinct predictors associated with a higher likelihood of outcome modification. Trials backed by industry funding had significantly higher odds of changing their primary outcomes compared to non-industry-funded trials. Furthermore, compared to drug trials, device trials and biologic trials were substantially more likely to shift their targets, while certain medical fields like neoplasms faced the highest odds of alteration.  

Evaluating these conclusions reveals important strengths and limitations in how we interpret this data. On the positive side, deploying an LLM to parse over 15,000 records solves a massive logistical bottleneck, proving scalable oversight innovation. Documenting the steep temporal decline and incorporating robust statistical adjustments for classification errors further strengthens the empirical value of the findings. However, a closer look highlights notable cons, including potential model accuracy blind spots due to a limited training sample, the exclusion of nuanced metric wording shifts, and the conflation of malicious p-hacking with legitimate, scientifically justified mid-trial adaptations. Additionally, relying on a commercial, rapidly evolving proprietary AI introduces reproducibility challenges, while observational models cannot definitively prove malicious intent behind industry or oncology modifications. Ultimately, this methodology successfully demonstrates a practical framework for regulators, journal editors, and funders to monitor clinical evidence at scale.  

Ultimately, this study proves that while shifting primary outcomes remains a notable hurdle in clinical research, the overall trajectory of scientific transparency is heading in the right direction. The dramatic decline in modifications from 2016 to 2024 signals a maturing research ecosystem that increasingly values strict pre-registration adherence. Furthermore, the integration of large language models opens up exciting new frontiers for automated, large-scale oversight by regulators and independent watchdogs. By shining a brighter, automated light on trial registries, we can better protect patients, physicians, and policymakers from biased reporting. Moving forward, embracing these AI-driven monitoring tools will be essential to keeping clinical research honest, accountable, and robust. 

Author

FDA Purán Newsletter Signup

Subscribe to FDA Purán Newsletter for 
Refreshing Outlook on Regulatory Topics