Data and AI

AI Under GMP

The rules for machine learning in regulated manufacturing are being written right now, in public, and most of the people building the models have not read them. What the draft Annex 22, the rewritten Annex 11 and the FDA’s credibility framework actually say, and what to build before any of them is final.

Field guide ~13 min read September 2026

The rules for machine learning in regulated pharmaceutical manufacturing are being written right now, in public, and most of the people building the models have not read them. None of the instruments below is final. All of them are close enough to design against, and they agree with each other about more than they disagree.

There is a version of this subject that consists of the word governance repeated until the slide ends. This is not that. What follows is what the documents actually say, what status each one is in as of September 2026, where they contradict each other, and what you can do now that will still be right whichever way the open questions land.

One disclosure before the substance: I have a commercial interest here, in that this is part of what my practice does. That does not change what the drafts say, and every claim below is cited so you can check it against the source rather than against me.

Where the rules are, as of September 2026

Five separate instruments touch a machine learning model used in or around a medicinal product. They come from three different authorities, they are on three different clocks, and they are not coordinated with one another.

The instruments, and what state each one is in
InstrumentAuthorityStatusWhat it governs
EU GMP Annex 22, Artificial IntelligenceEuropean Commission, with PIC/SDraft, published 7 July 2025, consultation closed 7 October 2025, not yet in forceAI and ML models used in the manufacture of active substances and medicinal products
EU GMP Annex 11, Computerised Systems, revisedEuropean Commission, with PIC/SDraft, same consultation window, not yet in forceThe computerized system the model runs inside
EU GMP Chapter 4, Documentation, revisedEuropean CommissionDraft, same consultation windowRecords, including electronic ones
FDA draft guidance, Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological ProductsFDA, CDER and CBERDraft, January 2025, comments closed 7 April 2025, not finalizedAI that produces information submitted in support of a regulatory decision
EU AI Act, Regulation (EU) 2024/1689, as amendedEuropean Union, horizontal lawIn force, applying in stagesAI systems as products, independent of whether they touch a medicine

Two further documents are not law and are worth more than most law. GAMP 5 Second Edition (ISPE, 2022) added Appendix D11 on artificial intelligence and machine learning, and ISPE followed it in July 2025 with a standalone GAMP Guide: Artificial Intelligence. Inspectors do not enforce GAMP. They do read it, and so does the consultant who will audit you.

On the word draft

Draft is not the same as speculative. A published consultation draft is the authority telling you, on the record, how it currently thinks. Text moves between draft and final. The underlying position rarely moves far. Design against the draft and budget for the delta.

What Annex 22 actually says

Annex 22 is short and it is specific. It applies where an AI or ML model is used in a GMP-critical application, meaning one with a direct impact on product quality, patient safety or data integrity. Inside that boundary it asks for things that will be familiar to anyone who has ever qualified a piece of equipment, and one thing that will not.

The familiar parts: the intended use has to be written down, explicitly, including the boundary of the decision the model is permitted to touch. Personnel have to understand that intended use and the risks that come with it. Test data has to be independent of training data, and the testing has to justify that the model is responding to features that are actually relevant rather than to an artifact of how the data was collected. Everything is risk-based and everything is documented. None of this is new thinking. It is qualification, applied to a statistical object instead of a chromatograph.

The unfamiliar part is the line the draft draws between two kinds of model.

The static model line, which is the whole argument

Annex 22 as drafted distinguishes static models, which are locked after training and do not change during operation, from adaptive models, which continue to learn from data they encounter in use. In critical GMP applications the draft permits static models and excludes adaptive ones. Generative models and large language models fall outside the critical boundary on the same reasoning.

This has been read in a good deal of commentary as regulators being frightened of new technology. That reading is lazy. Consider what an adaptive model in production actually is, in the vocabulary of the quality system that already governs your site: it is a process that changes itself without raising a change control record. Validation establishes a state. Change control protects that state, and requalification re-establishes it when something moves. A model that retrains on Tuesday against data it saw on Monday has moved, silently, with no assessment, no approval, and no record of what the previous version would have decided. The whole apparatus of GMP is built to make that impossible.

A model that learns in production is a process that changes itself without a change control record.The argument Annex 22 is making, in GMP's own language

Framed that way, the static-only position is not conservatism. It is consistency. The interesting question is not whether regulators are right to be careful, it is whether the static and adaptive categories are the right two boxes. A model retrained quarterly, under change control, with a documented comparison against the previous version and a requalification, is nominally static and behaves like a managed process. A model that is technically frozen but whose input distribution has drifted a long way from its training set is nominally compliant and quietly wrong. The category that actually matters is whether the model's behavior is under control, which is a monitoring question, not an architecture question.

Whether the drafted line survives consultation is being argued now, and the generative exclusion in particular has drawn sustained comment. I would not bet either way. I would note that it does not much matter, for reasons in the last section.

The FDA is asking a different question

The FDA's draft does not carve the world into static and adaptive. It asks how much you need to prove, and it makes that a function of consequence. The mechanism is a risk-based credibility assessment framework organized around a context of use: you state the question the model will answer, you state the precise role its output will play in the decision, you assess the risk that arises from the model being wrong in that role, you plan the evidence that risk demands, you execute the plan, and you document what happened.

If that shape feels familiar, it should. It is analytical method validation. You do not validate a method in the abstract, you validate it for a purpose, and the evidence you owe scales with what happens if the answer is wrong. An assay supporting release carries a heavier burden than one supporting a development decision. A model is the same object under a different name, and the FDA has essentially said so.

The practical consequence

The unit you qualify is not the model. It is the model, plus the data it learned from, plus a written statement of the decision it is allowed to touch. Change any one of the three and you are looking at a different object with a different evidence burden. This is why data lineage turns out to matter more than model architecture, and why the team that cannot reconstruct its training set has a problem no amount of accuracy will fix.

The FDA also published Guiding Principles of Good AI Practice in Drug Development in January 2026. The January 2025 credibility draft has not been finalized as of this writing.

The AI Act is a separate machine

It is a common and expensive mistake to assume the EU AI Act is the pharmaceutical AI rule. It is not. It is horizontal product law that applies to AI systems by risk category regardless of sector, and it runs on its own clock, which has recently moved.

General-purpose AI model obligations began applying on 2 August 2025. The Act's general application, including the Article 50 transparency duties, lands on 2 August 2026. The high-risk obligations, which were originally due in August 2026, were deferred by Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026: standalone high-risk systems under Annex III move to 2 December 2027, and AI embedded in products already regulated under Annex I moves to 2 August 2028.

Most manufacturing analytics will not be high-risk under the Act. Some things adjacent to your program will be, and a delay is not a repeal. The operational point is that the AI Act and Annex 22 impose overlapping but non-identical documentation on the same model, and you want one set of artifacts that satisfies both rather than two programs that discover each other late.

What to do now, that will survive either answer

Everything below is doable today, requires no vendor, and is correct under every version of the drafts I have read.

  1. Write the intended use down, first. One paragraph per model: the question it answers, the decision it informs, the decision it is not allowed to make, and who owns the outcome. Almost nobody has this. It takes an hour and it is the artifact every one of these frameworks starts from.
  2. Make the training set reconstructible. Not documented, reconstructible: you can regenerate the exact rows the model learned from, on demand, a year later. If you cannot, you cannot answer the first question an inspector or a reviewer will ask.
  3. Version the model and the data together, and keep the old ones. A model version without its data version is not a version. You will need to be able to say what the previous model would have decided.
  4. Classify your models by consequence, not by cleverness. Which of them, if silently wrong for six months, would change a disposition, a specification, a submission or a patient outcome? That short list is your GxP-critical set. Everything else is a productivity tool and should not be carrying the same overhead.
  5. Monitor the input distribution, not only the accuracy. A frozen model whose world has moved is the failure mode the static rule does not catch. Set the drift threshold that triggers review before you need it.
  6. Define the requalification trigger in advance. Retraining schedule, drift threshold, process change, instrument change, supplier change. Write it into the quality system you already have rather than inventing a parallel one.
  7. Keep the human decision human where the model is generative. See below.

Where this leaves large language models

If Annex 22's exclusion survives, generative models are out of critical GMP applications. That is a narrower statement than it sounds, and it is worth being precise about, because the panicked version of it is wrong.

Out of critical applications does not mean out of the building. It means the model may not be the thing that decides. The design that survives the exclusion is the one where the model retrieves and presents, and a qualified human decides: it finds the three prior deviations that look like this one and quotes the relevant paragraph with a link to the source record, and an investigator who can open that record reaches the conclusion. The model never claims the decision, so the decision never inherits the model's credibility problem, and the reviewer can check the evidence rather than trusting a summary.

This is not a workaround. It is a better architecture on its own merits, because it fails visibly. A summary that is subtly wrong is indistinguishable from one that is right. A quoted paragraph with a link either supports the claim or does not, and the person reading it can tell in seconds.

The categories will move. Something will be final by the end of the year, the text will differ from the drafts in ways nobody can currently predict, and the second-order guidance will take another two years to settle. None of that changes what you should be building, because the artifacts that survive are the boring ones: a written intended use, a reconstructible data set, a version history, a consequence classification, and a monitoring plan. Anyone selling you a compliant AI platform in September 2026 is selling you a guess about text that does not exist yet. The work above is not a guess, and you will need it whichever way the guess lands.

References

  1. European Commission, EudraLex Volume 4, Annex 22: Artificial Intelligence, consultation draft, 7 July 2025. Consultation closed 7 October 2025. Reasons for changes, health.ec.europa.eu.
  2. European Commission, EudraLex Volume 4, Annex 11: Computerised Systems, consultation draft, 7 July 2025. Reasons for changes, health.ec.europa.eu.
  3. PIC/S, joint stakeholder consultation on the revision of Chapter 4, Annex 11 and Annex 22. picscheme.org.
  4. ECA Academy, EU GMP Annex 22 (Draft 2025): Artificial Intelligence. gmp-compliance.org.
  5. FDA, Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025. fda.gov.
  6. FDA, FDA Proposes Framework to Advance Credibility of AI Models Used for Drug and Biological Product Submissions, press announcement. fda.gov.
  7. FDA CDER, Artificial Intelligence for Drug Development, programme page. fda.gov.
  8. Regulation (EU) 2024/1689 (Artificial Intelligence Act), as amended by Regulation (EU) 2026/1744, Official Journal, 24 July 2026, deferring the high-risk application dates to 2 December 2027 (Annex III) and 2 August 2028 (Annex I).
  9. ISPE, GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems, Second Edition, 2022, Appendix D11, Artificial Intelligence and Machine Learning. ispe.org.
  10. ISPE, GAMP Guide: Artificial Intelligence, July 2025.
  11. 21 CFR Part 11, Electronic Records; Electronic Signatures.
  12. Rizkin, B. The Analytical Procedure Lifecycle. The method-validation argument this piece leans on, in its original setting.

Status of every draft above verified September 2026. Drafts change. Check the source before you rely on it.