An AI model can be right and still not be trusted. That’s what Bristol Myers Squibb (BMS) found six months into rolling out AI-powered adverse event processing: while F1 performance scores were tracking upward, the audit trail found reviewers still reopening cases the system had already cleared. Sometimes more than one reviewer had checked the same file, even though not specified as part of quality control.
Once accuracy stops being the open question, the harder problem is behavioral, not technical. All of this risks compromising the return on investment, particularly when it comes to freeing up specialist time, one of the main goals of AI use.
Accuracy Is No Longer The Problem
Stated confidence in AI adoption across pharma often belies what’s actually in place. BMS is an exception in a pharmacovigilance context. Its data extraction system, which pulls structured findings from an incoming adverse event report before a reviewer sees it, has been in production long enough, and across a sufficient case volume, for the audit trail to mean something.
The company processes roughly 400,000 individual case safety reports a year. Automated data extraction has taken care of roughly 60% of that volume from June 2025, and the company plans to extend this proportion as confidence builds – as long as it can find a way to distance itself from the 100% quality-control regime that has long surrounded its case processing.
Although some checks can be put down to legitimate second-guessing, persistent checks just for the sake of it point to a level of confidence that doesn’t yet tally with the technology’s demonstrated performance. It may simply be that behavior and process haven’t fully caught up with what the technology can now do.
Part of the lag is technical. Reviewers trained on rules-based software expect identical inputs to produce identical outputs every time; AI doesn’t work that way, so a system that’s right 95% of the time can look shaky, even when 95% comfortably clears what operations require. Adding more rules to chase consistency usually backfires, too, in that each additional constraint narrows what the model can handle without a human having to step in.
Handling of AI-generated case narratives illustrates the issue. Here, many of the edits reviewers make reflect house writing preferences rather than clinical or factual error. If team members log every edit as a correction, though, it can look as though the AI underperformed.
Making Trust Measurable
Applying blanket quality control to an AI-driven process doesn’t make sense if organizations want what the technology actually promises in a PV setting. The opportunity isn’t just processing more cases faster; it’s also about freeing up the valuable time of experienced specialists.
Trust, though, can’t be measured by asking reviewers how confident they feel in AI’s output. It doesn’t matter how much scientists say they trust the system, if they insist on rechecking its output.
The idea of a formal trust coefficient is to draw together several operational signals, including model performance, reviewer behavior, validation and governance maturity, task risk, and an organization’s overall readiness to rely on AI outputs - into one structured view. The goal of this wouldn’t be a fixed formula, so much as a consistent way of deciding how much independent human verification a validated system actually needs, so oversight can become lighter as evidence builds - without governance growing looser.
PV already has service-level metrics for uptime and turnaround time, but nothing equivalent for managing the shift from blanket AI-workflow checking to evidence-based oversight. A trust coefficient would belong to a different category from existing operational KPIs altogether: not tracking performance or efficiency but measuring how much reliance a validated system has actually earned. That’s where the real work in AI operationalization happens over the next two years, not in raw accuracy.
More Autonomy for AI, Same Accountability for Humans
A trust coefficient will need to account for agentic AI too. Where earlier AI-enabled workflow tools supported individual tasks, agentic systems offer to complete entire sequences of preparatory work - pulling data, checking fields, and assembling evidence - before a case even reaches a human reviewer. This redirects specialists’ attention toward interpretation, judgment calls, and the genuinely uncertain cases that need their expertise.
The distinction, here, is between reviewers checking every step of a system’s work and reviewers stepping in only where risk or clinical significance calls for it. None of that changes who answers for the outcome. Agentic AI can shift where the groundwork happens, but responsibility for patient safety stays with qualified PV professionals.
Redesign First, Automate Second
BMS’s experience carries a broader lesson: that successful AI deployment starts with redesigning the process, not automating one that already exists. Before introducing AI in case processing, the company took pains to streamline its workflow, which had previously featured multiple hand-offs and interfaces. By modernizing its core safety platform, it ensured that any AI-enabled automation would build on a simplified process rather than being layered onto one designed for manual work.
BMS’s work hasn’t finished: its rolling three-year program aims to remove further hand-offs and manual touchpoints, and its AI capability will keep evolving. Rather than fixating on choosing the right technology at a single point in time, it has prioritized continuous improvement as the technology advances.
Regulators Are Championing Change Too
An overlap with regulatory ambition is already evident. International expert groups have been busy publishing guidance around the idea that AI governance should be proportionate, risk-based, and built into a system across its lifecycle, rather than bolted on later.
In January 2026, the European Medicines Agency and the US Food and Drug Administration jointly issued 10 guiding principles for AI use across the medicines lifecycle, covering evidence generation from early research through manufacturing to post-marketing safety monitoring, and calling for validation proportional to intended use, solid data governance, and ongoing performance monitoring.1 The Council for International Organizations of Medical Sciences reached similar conclusions in its Working Group XIV report on AI in pharmacovigilance, published in December 2025: rather than legislating for specific technologies, it favors governance frameworks flexible enough to keep pace with how AI capabilities evolve.2
The case for doing something proactive is only getting stronger as medicines grow more complex and adverse event volumes climb. VigiBase, the World Health Organization’s global database of individual case safety reports, was recorded to hold more than 40 million reports by the end of 2024, with roughly 70% filed within the previous decade.3 But processing that volume is now less of a barrier than understanding when a validated system warrants routine checking - and when it doesn’t.
As things stand, there is no agreed method yet for calculating a trust coefficient, which is understandable given how new this territory is. But scale strengthens the case for building evidence-based reliance and reducing disproportionate checks, calibrated to the particular use case and its identified risk — with improved confidence following as a natural byproduct, rather than being the thing organizations aim at directly. An organization processing several hundred thousand safety reports a year wouldn’t need a dramatic shift in oversight to free up meaningful human capacity. A modest, well-governed adjustment would probably suffice to redirect expert time to the cases that actually need it.
This article builds on themes discussed in a recent life sciences industry podcast, which can be accessed in full at https://www.arisglobal.com/podcast/.
References
- European Medicines Agency and U.S. Food and Drug Administration, ‘EMA and FDA set common principles for AI in medicine development’, 14 January 2026. Available at: https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
- Council for International Organizations of Medical Sciences (CIOMS), Artificial Intelligence in Pharmacovigilance’, CIOMS Working Group XIV report, Geneva, December 2025. Available at: https://cioms.ch/working_groups/working-group-xiv-artificial-intelligence-in-pharmacovigilance/
- Brand JS, Gauffin O, Sartori D, Fusaroli M, Sköld H, Bergvall T, Sandberg L, Wallberg M, Hjelmström P, Norén GN. ‘VigiBase: Resource Profile Update with a Summary of Global Patterns and Trends in Adverse Event Reports for Medicines and Vaccines’. Drug Safety. 2026;49(6):613–629. DOI: 10.1007/s40264-025-01642-6
About the Authors
Jason Bryant is General Manager of AI Platforms at ArisGlobal, based in London. A data science actuary by training, he has built his career across fintech and health-tech, specializing in AI-driven, human-centric product innovation, and previously led an AstraZeneca digital incubator before turning to agentic AI in drug safety.
Samuel Wallis is Head of Case Processing at Bristol Myers Squibb, based in London, responsible for pharmacovigilance operations across one of the industry’s highest ICSR volumes. He has led major transformation programs spanning global safety operations, including the migration of BMS’s global safety database to the cloud and the deployment of AI-enabled case processing.