Blog Post

Can AI Automate a SAP?

Can AI Automate a SAP?

Beneath the Surface of the Statistical Analysis Plan

There is an emerging commercial narrative that AI can automate the drafting of a Statistical Analysis Plan (SAP) directly from a clinical protocol. Extract the endpoints, populate a template, link to table shells, and the plan is ready. For anyone who has lived through the hours behind the process of drafting a SAP, the appeal is immediate.

The promise of a shortcut, however, reveals as much about the one proposing it as it does about the task itself. When a company promises to automate the SAP, the question isn’t really about the technology. It’s about whether they understand the nature of the work at all.

The SAP carries a study from database lock to regulatory submission. The reality of SAP development in medical device and IVD research is typically deeper, more iterative, and more technically demanding than any extraction-and-template workflow can capture. This is not a document that falls out of the protocol once the right fields are tagged. It is a crafted instrument that is negotiated, designed, and stress-tested. The gap between that reality and the automation promise is worth examining closely, especially for sponsors and early-career statisticians who may be tempted by the prospect of a faster path.

The SAP Doesn’t Start with the Protocol Alone

A protocol states the primary endpoint: “change from baseline in six-minute walk distance at 90 days.” But before a statistician can write the corresponding section of the SAP, they need the annotated eCRF. What exactly is collected? Is the distance recorded as a continuous value in metres, or is it captured as a categorical range? If it’s the latter, the analysis cannot proceed with the ANCOVA that was provisionally planned. The derivation of “change from baseline” also demands a precise baseline definition, such as “the last non-missing assessment on or before the procedure date”, and a visit windowing rule that selects the observation closest to Day 90, with ties broken by the later date. None of these operational details live in the protocol.

Then there is the Data Management Plan, which specifies (for example) how missing and implausible values are queried, and that adverse events are coded using MedDRA version 29.x. That version number must be cited in the SAP, and the rules for treatment-emergent attribution must align with the cleaning conventions. If CDISC standards are in use, the SAP must name the exact ADaM variable that will hold the derived change-from-baseline, perhaps CHG in ADEFF at PARAMCD = "SIXMWD" and AVISIT = "Day 90", and map it to the SDTM source field. For an IVD study, the reference standard algorithm determines which subjects are true positives; a plan written without it may misclassify indeterminate results and inflate sensitivity.

A tool that sees only the protocol cannot audit these connections. It cannot flag that the eCRF captures an ordinal pain scale when the SAP specifies a linear mixed model. It cannot notice that the imputation model must include the treatment arm to avoid bias toward the null. It is building on a partial foundation. The result is a document that may read smoothly but is silently disconnected from the data the study will later receive, a flaw that surfaces only when the programmer starts coding.

The Integration Tax of a Thousand Micro-Decisions

To be fair, artificial intelligence is incredibly capable of resolving an individual statistical quandary. If you ask an AI to define the Boolean logic for a treatment-emergent adverse event, or to calculate the design effect for a clustered sample size, it can often provide a highly competent sounding, localised response. AI vendors love to demonstrate these isolated feats of reasoning.

It pays, however, to notice what those demonstrations conceal. Each question arrives already framed: someone has decided what counts as treatment-emergent, chosen the clustering unit, fixed the ICC. The AI is resolving a problem a human has already isolated. Writing a SAP is the opposite operation. Instead of answering questions one at a time, it involves a) deriving the specific bespoke questions that must be answered in this novel edge case (which represents the entire clinical investigation from an evidentiary and regulatory perspective); b) holding every relevant statistical resolution in view at once; and c) weighing them into a single, fully defensible conclusion that satisfies all relevant points of consideration.

A SAP is not a single, easily-reconcilable problem. It is a composite of count-less minor considerations and micro-decisions, each of which must align with the others. Crucially, they do not align in a neat, linear sequence – they tend to conflict. A choice made to satisfy a clinical nuance might violate an operational constraint in the eCRF. A regulatory precedent might demand a conservative sensitivity analysis that causes the primary statistical model to fail to converge. The work is not executing a chain of dependent steps; it is holding a range of competing variables in tension simultaneously, weighing trade-offs across clinical, operational, and statistical domains until a single, defensible compromise is reached.

There are some linear, cascading decisions which AI would currently struggle with in a regulated environment but could more-or-less handle with sufficient training and study-specific input documents. For example: A baseline variable definition dictates the corresponding change-from-baseline variable derivation; this derivation in turn dictates the missing data handling; the missing data handling dictates the imputation model; the imputation model dictates the table shells- and so forth. Even so, this represents one of a potential myriad of decision chains that comprise a SAP.

When you account for the distributed nature of study sites, the variances in EDC builds, and the iterative estimand negotiations, the engineering required to orchestrate a system of agents capable of maintaining logical consistency across all competing input is far more complicated and energy intensive than a statistician simply drafting the document manually.

The commercial pitch is that an AI generated draft, checked and revised by a human, nets out cheaper than a human-written one. This is a dubious assumption once the practicalities are considered. Verifying a SAP is not mere proofreading; it requires the formal, structured QC demanded by a highly regulated, high-commercial-stakes environment. That QC means independently re-deriving the document’s logic. A statistician reviewing their own draft carries the mental model that produced it: they know why the baseline definition took its form and what it constrains downstream. In independent QC, a statistician reviewing someone else’s draft must reconstruct that model from the text, decision by decision. Between two humans in the same field this is tractable, because the reviewer can generally assume the author reasoned their way to each decision, and an error tends to show up as a wrong step in that reasoning. With AI generated text the reviewer has less to lean on, and is closer to verifying each decision in its own right than to following a line of argument. Fluency makes this harder, not easier: confident, well-formed prose is exactly where fabricated detail and quietly incompatible decisions are hardest to spot. The drafting hours supposedly saved by AI are spent, with interest, on verification and subsequent re-drafting. The residual temptation to lighten that verification because a draft reads well is exactly the failure mode a GxP environment exists to prevent.

The Plan is Forged in Conversation, Not just Extraction

A protocol might say: “Subjects in whom the device cannot be deployed will be replaced.” In a cross-functional kickoff, the statistician asks whether a deployment failure due to anatomical unsuitability is a screening issue, while a failure due to a fractured catheter is a device deficiency. The two scenarios lead to different analysis set assignments: one exclusion from mITT with careful documentation, the other an inclusion in the Safety Set and a composite strategy for the efficacy endpoint. The SAP must capture this logic as explicit Boolean conditions referencing eCRF fields. No protocol text spells it out.

These discussions are where estimands are built under ICH E9(R1). For a peripheral atherectomy device, the intercurrent event of “bailout stenting” might be handled with a composite strategy: the subject is assigned a worst-case residual stenosis value, rather than being excluded. The five estimand attributes: (the treatment condition, the population, the variable, the handling of intercurrent events, and the population-level summary) must be stated for each endpoint. For a key secondary endpoint evaluated at six months, the population might be the mITT set, the variable the patient-reported VAS pain score, the intercurrent event of device explant handled by a hypothetical strategy, and the summary measure the difference in least-squares means from a mixed model with Kenward–Roger degrees of freedom.

Such choices do not emerge from a template. They tend to emerge from the statistician playing devil’s advocate: “If the wound hasn’t healed by Day 30, and the patient stops attending, is that missing at random, or missing not at random?” That question determines whether the primary analysis uses multiple imputation under MAR, or whether a tipping-point sensitivity analysis must be pre-specified with a grid of delta shifts. The SAP is the record of that reasoning, but the reasoning itself happens in whiteboard sessions, not in document extraction.

Every Study Carries Its Own Fingerprint

In MedTech, while clinical trials often repeat a relative pattern, they do so with a good dose of individual nuance and some novelty thrown in. One study might involve clustered lesions within patients, requiring a generalised linear mixed model with a random patient intercept and an intra-class correlation coefficient (ICC) justified from literature. The sample size calculation must inflate the naïve estimate by the design effect (1 + (m-1)ρ), where m is the average cluster size and ρ is the ICC – an adjustment an automated power calculation might overlook. Another might be a multi-reader multi-case study whose variance must generalise across readers and cases alike, not a simple DeLong AUC comparison.

For a Bayesian adaptive design borrowing strength from a previous device generation, the SAP must specify the power prior, the discounting parameter, and the simulation plan that demonstrates operating characteristics. The prior itself must be defensible to a Notified Body reviewing under MDR Annex XIV, an expectation that is context-specific and not codified in any public training corpus. A third trial might involve a safety composite endpoint like MACE, where the Boolean logic must adjudicate cardiovascular death, target-vessel myocardial infarction, and clinically driven revascularisation using data from an independent Clinical Events Committee, not site-reported terms.

Each of these requires technical specifications that no off-the-shelf template can supply. The SAP must name the SAS procedure (PROC MIXED with ddfm=KR), the R package (lme4 with lmerTest, using ddf = "Kenward-Roger"), the confidence interval method for sensitivity (Clopper-Pearson because the numerator is small), and the pre-specified convergence fallback: if the unstructured covariance matrix fails, simplify to compound symmetry and document the change. An automation that fills in a generic “mixed model” or “appropriate nonparametric test” is writing a placeholder, not a plan. And in a document that must survive regulatory scrutiny, a placeholder is a liability.

The SAP Is Read Under Pressure, Not at Leisure

A SAP is an operational manual for a programmer who must translate each sentence into production code, and for the statistician who will later interpret the outputs and draft the clinical evaluation report. In MedTech, statistical programming is rarely straightforward. It often involves codifying highly sophisticated models and simulations that cannot be derived from the protocol alone, nor readily pieced together from a combination of input documents unless the SAP has explicitly synthesised them alongside the conversational decisions and considerations. Programmers and statisticians use the SAP as a workflow document. If they did not draft it themselves, it must be painstakingly clear about the methodological nuances and exactly how and why they play out practically. Both are working to deadlines, holding multiple rules in mind. A document that forces them to piece together the population definition from Section 3, the covariates from Section 4, and the missing data rule from Section 5 is a document that invites error and inefficiency at the QC stage, never mind the programming and analysis stages. Calibration is the craft: restate everything and the document bloats until nothing stands out; cross-reference everything and the reader assembles each analysis from fragments. Deciding what to repeat, where, and in how much detail is a delicate set of judgements in itself.

A proponent of automation might counter this by suggesting that if the SAP is this precisely specified, the statistical programming itself could simply be automated next. If the exact specifications are in the SAP, and the analysis is exactly specified in the code, why not let an AI write the R or SAS scripts? The lived experience of statistical programming tells a different story. Certainly, AI has a role to serve as an efficiency booster in any coding context. This is, however, miles away from any kind of automation. The statistical programming process for a clinical study usually ends up unearthing as much nuance as the SAP draft and study design specification themselves. Writing the code forces a granular confrontation with the actual data- revealing undocumented edge cases, unexpected structural quirks, or logical paradoxes that remained invisible at the planning level. The code is not a mechanical translation of the SAP; it is the final, practical stress-test of the statistical logic. Translating a complex MACE composite or a Bayesian adaptive simulation into executable code requires a continuous stream of micro-decisions that loop right back to the original study design conversations. Automating the programming doesn’t eliminate the nuance; it just moves it one step further down the chain, where an AI is potentially even less equipped to resolve it.

Why Capturing the Conversations Doesn’t Close the Gap

A natural response to the points above is to ask whether the missing conversations could simply be fed into the automation as well. After all, tools now exist (such as Aimanuensis Clinical Trials) that document and summarise clinical trial meetings – capturing every word of the kickoff discussions, the estimand negotiations, and the whiteboard debates. The protocol, eCRF, DMP, SDTM specs, plus the meeting transcripts: surely that all bridges the gap?

In practice, it doesn’t. One reason lies in the non-linear nature of the conversations themselves. Clinical meetings are not linear briefings; they are iterative, exploratory, and unavoidably messy. The other lies in what the transcript would demand of its reader: not the resolution of any single question, but the same multi-front integration the meeting itself performed – weighing each stated preference against the eCRF structure, the estimand strategy, the regulatory precedent, and every decision already provisionally made. Dedicated software can extract decisions and action items from individual meetings. That is different to the kind of synthesis required for SAP development.

A clinical meeting about a device study loops. A statistician raises a concern about an endpoint definition; the clinician pushes back with a real-world example; a regulatory colleague mentions a Notified Body precedent; the statistician revises the proposed handling and sketches a sensitivity analysis. Ten minutes later, the discussion circles back after a new point about the eCRF structure casts the earlier decision in a different light. The final agreed approach emerges not from any single documented decision, but from the synthesis of the entire exchange.

Handing an AI a list of “decisions” doesn’t solve the integration problem. Decisions made on Monday are often invalidated by a competing practical consideration discovered on Wednesday. The SAP is the final, synthesised resolution of a constantly shifting web of clinical, operational, and regulatory constraints. An AI can be fed the individual pieces from a repository of decisions, but it cannot independently or reliably weigh competing considerations, resolve ambiguities, or fill in the gaps that the discussion itself never fully closed.

Many of the most important resolutions happen between meetings, in the statistician’s own work. The quiet afternoon spent checking whether the proposed mixed model will converge with 40 subjects and an average cluster size of 1.2 – these insights are not captured in any meeting transcript until they are discussed after-the-fact. They are the result of solitary technical reflection, and they feed back into the next discussion, altering its course.

Room for Tools, Space for Craft

When a company offers to turn a protocol into a SAP at the push of a button, the deeper question is not about the algorithm. It’s about what they think the SAP is. The assumption baked into such tools is that the SAP is a derivative document, obtained by mapping protocol sections to plan sections to some standard statistical boilerplate.

Ultimately, a SAP is the product of well-considered decisions made analytically by the biostatistician, balancing competing practical considerations. Once all the input documents and talking points have been digested, the proposed methodological approaches must be simulated and validated to confirm they hold up under the study’s specific constraints. An AI can draft a statistical method, but it cannot balance those competing considerations, nor can it independently design, run, and interpret the simulations required to prove the method is empirically sound. The final SAP is not just a written plan; it is an empirically tested engine that powers the study’s evidence gathering.

For sponsors, the risk of trusting automation is a SAP that passes superficial review but contains methodological gaps, gaps a Notified Body may identify, or that surface only during the CER phase when rework is expensive. For early-career statisticians, the risk is absorbing a model of the work that mistakes document assembly for statistical design.

None of this argues against automation. Intelligently extracting decisions and action items from clinical trial meeting transcripts, such as in the case of Aimanuensis, brings real efficiency. Tools with the ability to cross-check endpoint consistency between protocol and SAP, flag missing sections, or generate first-draft table shells aligned with the SAP’s population definitions would lighten the administrative layer of the work and be very much welcomed.

In medical device and IVD research, where nearly every study deviates from the standard template, human judgement is not a bottleneck to be engineered away. It is the very thing that gives the SAP its authority. While removing human judgement might be a good idea for self-driving cars – writing a SAP is not a task that should be completed on auto-pilot with no recollection of the journey that just took place. Those suggesting otherwise are selling a shortcut to a destination they haven’t fully mapped.

https://anatomisebiostats.co/wp-content/uploads/2026/07/MedTech_SAP_Handbook.pdf

Related Posts