AI implementation in pharma: a thought-provoking series
Why most AI pilots in pharma stall, what an agent needs to work reliably, what to ask before investing, and where the FDA and EMA are heading.
Article6 min read9 Sep 2026
Almost every pharmaceutical company is evaluating AI, or already implementing it. This series approaches that challenge from four angles: why most pilots never reach production, what AI agents need in order to work reliably with a company's information, what to ask yourself before investing, and where the leading regulators are focusing.
01 · From pilot to production
95% of AI implementation projects never make it past the pilot phase.
The MIT study1 behind that figure concluded that the main obstacle to scaling is not how capable AI agents are, but their integration with each company's workflows and the context they need in order to operate. Expecting an AI to work reliably on scattered information (spread across documents and disconnected systems, without structure or context) is an act of faith, given how the technology works. Whether a pilot scales depends on the structure it runs on.
02 · What AI needs to work
The AI you implement will run on the data you already have. Could a new hire work with that data without asking anyone a question?
A new hire arrives with zero context about the company, learns through onboarding, and accumulates knowledge by asking questions and getting answers. An AI agent starts from the same place, but it cannot ask, and it cannot build up context on its own. Its only way of learning about your company is the information it has access to.
In a 2026 benchmark2, Claude Opus 4.7, Claude Sonnet 4.6 and GPT-5.4 answered business questions over the same database and got between 45% and 50% right; given a context layer defining what the data meant, all three answered close to 68% correctly. With context, the three models were statistically indistinguishable.
In pharma the effect is even more pronounced: on Bayer's internal pharmacovigilance data, GPT-4 went from 8% to 78%3 accuracy once it was given explicit business context.
In short: the reliability of an AI is defined not by the model, but by the context it is given.
03 · What to ask before you invest
3 out of 4 pharma companies invest in AI without having defined what success looks like, or how to measure it.
A McKinsey study4 found that about 75% of pharma and medtech leaders say their organization lacks a roadmap with clearly defined success measures tied to business priorities, and that only 5% have realized consistent financial value from generative AI.
Four questions that apply to any process:
- How long does it take? Developing a product, releasing a batch, implementing an already-approved change.
- How much human capital does it consume? Documentation reviews, quality investigations, verifications across systems.
- How many errors or delays does it generate? Regulator observations, failed registrations, rejected variations.
- What does each error or delay cost? A launch that slips, an approved change left unimplemented, materials destroyed.
The structure of your information is key both to measuring success and to an implementation that pays off.
04 · Where regulators are heading
In 2026, the FDA and EMA published ten principles of good AI practice. Six of those ten are about your data.
In January 2026, the FDA5 and the EMA6 jointly published ten principles of good AI practice across the entire medicines lifecycle. Six of the ten address the data AI operates on:
Figure 1
| Principle | What it says about data |
|---|---|
| 3. Adherence to standards | AI technologies must adhere to relevant legal, technical, scientific and regulatory standards, including Good Practices (GxP), whose backbone is data integrity. |
| 6. Data governance and documentation | Data source provenance, processing steps and analytical decisions are documented in a detailed, traceable and verifiable manner, in line with GxP requirements. |
| 7. Model design and development practices | Development leverages data that is fit-for-use. |
| 8. Risk-based performance assessment | Assessments evaluate the complete system, including human-AI interactions, using fit-for-use data and metrics appropriate to the context of use. |
| 9. Life cycle management | Scheduled monitoring and periodic re-evaluation to ensure adequate performance, for example to address data drift. |
| 10. Clear, essential information | Plain-language information on the AI's context of use, performance, limitations and underlying data. |
Source: FDA; EMA. Guiding Principles of Good AI Practice in Drug Development, January 2026.
The other four principles address governance and human-centricity.
To align with all ten, you need a system.
Krevyos is the operating system for AI in life sciences: a single source of truth built on structured data, with governed workflows and documented decisions, for every product in every market.
For your team, it means no longer spending hours on searching, reconciling and verifying scattered information, and spending them instead on defining strategy, exercising judgment, and everything else AI cannot contribute. For the AI, it means taking on the operational workload reliably, because it works on a controlled, auditable context.
By
- Julieta Villafañe
Founder & CEO
References
- MIT NANDA. The GenAI Divide: State of AI in Business 2025 (2025).
- Rumiantsau, M.; Fokeev, I. Semantic Layers for Reliable LLM-Powered Data Analytics: A Paired Benchmark of Accuracy and Hallucination Across Three Frontier Models. arXiv:2604.25149 (2026).
- Painter, J. L.; Chalamalasetti, V. R.; Kassekert, R.; Bate, A. Automating Pharmacovigilance Evidence Generation: Using Large Language Models to Produce Context-Aware SQL. arXiv:2406.10690 (2024).
- McKinsey & Company. Scaling gen AI in the life sciences industry.
- FDA; EMA. Guiding Principles of Good AI Practice in Drug Development (January 2026).
- EMA. EMA and FDA set common principles for AI in medicine development (14 January 2026).