Blog Post

Clinical Evaluation of Medical Devices: Equivalence Margins, Operator Learning Curves, and Evidence That Keeps Moving

Series: Advanced Biostatistics for MedTech: Bridging Clinical Evaluation and Engineering

The first clinical investigation plan from a device manufacturer tends to send a pharmaceutical statistician looking for things that are not there. There may be no placebo arm, no dose-escalation cohort, no blinded packaging, and for a large share of devices the clinical evidence package contains no new interventional study at all. What sits in their place is a predicate device, a literature appraisal, a plan for managing the operator learning curve, and a set of endpoints that concern the device itself as much as the patient.

The instinct is to read this as a deficit. This diagnosis would be wrong. A statistician from biosimilars arrives fluent in equivalence testing, an oncology statistician has modelled death as a competing risk for years, and rare-disease programmes have made external controls an accepted, though still justified, option. Most of the machinery is shared. What differs is structural: where the margin comes from, what role the operator plays, who can be blinded, what the endpoint is actually measuring, and the awkward fact that the evaluated object keeps changing after approval.

Why a Device Trial Has No Placebo Arm: Substantial Equivalence and determining the Equivalence Margin

Pharmaceutical development leans on equivalence and non-inferiority testing far more than the placebo-dominated reputation suggests. Bioequivalence for generics rests on codified acceptance limits of 80 to 125% for log-transformed pharmacokinetic parameters. Biosimilarity adds a stepwise totality-of-evidence approach in which statistical equivalence testing is applied directly to analytical assay data. Non-inferiority designs with M1 and M2 margins, and the constancy assumption that underpins them, are standard whenever the aim is to demonstrate retention of a reference effect, including settings where a smaller effect is consciously traded for a better safety profile.

Device equivalence uses the same machinery. The Two One-Sided Tests procedure of Schuirmann (1987) governs both.

Under MDR Article 61, equivalence to a predicate device is evaluated under strict conditions, alignment of technical, biological, and clinical characteristics as set out in Annex XIV together with access to the underlying data, but the statistical limits must be constructed case by case from engineering tolerances, analytical performance goals, and the historical performance of the predicate. The claim is also about a system rather than one or two parameters: the margin negotiation spans engineering and clinical judgement at once. Documenting the rationale, and showing sensitivity of conclusions to the margin choice, is a substantial part of the work.

The FDA 510(k) pathway demands its own demonstration of substantial equivalence, with evidentiary conventions that differ from the EU regime, so a sponsor operating across both jurisdictions defends the same product against two sets of expectations.

For more information on equivalence and non-inferiority methods in a medtech context see here.

The Operator Learning Curve: Why Early Cases Make a Good Medical Device Look Bad

Pharmaceutical development is not totally devoid of technique effects. CAR-T administration depends on centre experience, with treatment volume associated with outcomes and accreditation requirements encoding that dependence. Surgical trials within oncology carry surgeon learning curves. Site effects in multi-centre drug trials are real and routinely modelled. The distinction is degree and constitutiveness: in a drug trial the technique surrounds the intervention, while for an implanted or surgical device the technique is part of what the product’s performance means.

A medical implant is placed by a clinician whose procedure is often as new as the device itself, and early cases carry worse outcomes for reasons unrelated to the product. The learning curve confounds the device effect in a way that can work in either direction, and clinical study designs absorb it explicitly. A training phase or run-in excludes the first cases per operator. Eligibility may require minimum procedure volumes. Where exclusion is wasteful, the curve is modelled directly, with case sequence or cumulative site experience entering the analysis as a covariate. Operator and centre enter as random effects, and clustering by site is often stronger than a drug statistician would expect from a comparable number of centres. The specific failure mode is attribution: crediting late operator mastery to a device improvement that never happened, or charging early inexperience to the device.

Blinding: What to Do When the Surgeon Knows What They Implanted

Blinded adjudication is shared territory across both pharma and medtech studies. Cardiovascular outcome trials in pharma depend on blinded clinical events committees, open-label periods are managed rather than forbidden, and single-arm studies with external controls are accepted in rare disease and oncology, with platform trials sharing control arms across programmes.

Devices shift the default. The surgical implanter cannot be blinded, and the patient often knows what was placed. Sham-controlled device trials exist, renal denervation being the canonical example, but they remain rare and carry procedural risk that an inert tablet does not. The consequence is tiered blinding wherever possible. Imaging assessors and event adjudicators are blinded even when operators and patients are not, and endpoints are tilted towards objective adjudicated outcomes. The predicate-based comparison then operates as the architectural norm rather than a special pathway. The external-control arm analytic burden that pharma reserves for settings without a randomisation option can often be the default of device evaluation.

Device Endpoints That Track the Product: Competing Risks When the Patient Dies First

Competing risks are part of any pharma study. Oncology statisticians spend careers modelling patient death as a competing risk for disease progression. In drug studies recurrent cardiovascular or other clinical events are tracked to give a better understanding of disease progression compared to time-to-first-event alone.

Device study endpoints frequently track the product as well as the patient. Composite device-success measures, target-lesion revascularisation, mechanical fracture, lead insulation failure, migration, and battery depletion sit alongside clinical outcomes. Time-to-device-failure competes with death in a peculiar way: a patient who dies of other causes can never experience the device failure that would otherwise have occurred, and naive Kaplan-Meier estimation overstates the failure risk accordingly, so cumulative incidence functions replace one-minus-survival plots. Re-interventions, revisions, and recalibrations call for recurrent-event models rather than first-event-only analyses.

The feature with no pharmaceutical parallel is the feedback loop. A device-failure finding travels back into design control and the risk management file, altering the product’s engineering. An efficacy finding in a drug trial can change the label or the prescribing guidance, but never the molecule; a device-failure finding can change the product itself.

Clinical Evidence for IVDs: When Agreement Against a Reference Standard Replaces the Treatment Effect

In pre-clinical pharma, biosimilar analytical similarity applies equivalence testing to assay data directly, and bioanalytical method validation is an established discipline. The difference is which side of the instrument the product sits on. In a biosimilar comparison the molecule is the measured object and the assay is trusted infrastructure. In an in vitro diagnostic the device is the measuring instrument, and the entire evidential question is whether its readout agrees with a reference procedure across patient samples.

Deming regression and its non-parametric counterpart, Passing-Bablok regression, are used in IVD studies in place of OLS regression and other popular pharma methods. This is largely because ordinary least squares assumes the reference standard is error-free and that biases the slope toward zero, which would usually be incorrect in this context. The correlation coefficient, a familiar statistic from drug trials, is nearly useless here, since it inflates with the width of the concentration range rather than with true agreement. Error grids add a layer with no direct pharmaceutical counterpart. These map measurement error of IVDs onto clinical risk so that agreement claims can be read as patient-safety claims.

For more information on regression methods in a medical device context see here.

Bayesian Adaptive Designs in Device Trials: Pooling Product Performance

In early-phase oncology, Bayesian methods solve dose-finding. Continual reassessment learns a dose-toxicity curve as cohorts arrive, and the posterior drives escalation decisions. The prior concerns a parameter of the pharmacology, or where the safe dose sits. Platform trials apply the same approach to control sharing, letting arms borrow a common control arm so each treatment is compared against a strengthened estimate. Rare-disease programmes use it for external evidence borrowing, pulling natural history or registry data into trials too small to stand alone. Bayesian adaptive designs and interim analyses are common to both industries as well, with application differing slightly.

In a device clinical investigation, a typical Bayesian task is pooling performance data across studies of the same product: a predicate generation, an earlier configuration, or concurrent multi-site validation of identical hardware. The prior concerns the device’s own performance distribution, and the posterior is used to justify that the current iteration meets its performance claims with fewer subjects than a standalone trial would need. This is why the FDA’s 2010 guidance on Bayesian statistics in medical device clinical trials is built around pooling across studies and why hierarchical models with between-study variability estimated, not assumed away, are the signature device application.

The exchangeability caveat is shared across all four settings: a prior from data that differs in design, population, or conduct will mislead, and reviewers probe it. But the object of the exchangeability argument differs. In dose-finding the question is whether the prior toxicity model matches this drug. In device pooling it is whether the predicate’s performance distribution matches the new iteration. Since the sponsor has the predicate’s actual data, that question is answered with engineering records rather than literature alone.


For more information on Bayesian methods in a medical device context see here.

The Technical File as a Single Argument: Engineering, Analytical, and Clinical Data Under One Set of Claims

Pharmaceutical submissions do integrate heterogeneous evidence, and CMC and clinical data travel in one filing. What CDISC provides is a mandated, standardised architecture for the clinical core, and what the pharma dossier maintains is a firewall between manufacturing evidence and clinical evidence, each reviewed by its own specialists.

A device technical file makes no such separation. Engineering data such as tensile strength and fatigue life, analytical data such as biosensor accuracy and assay precision, clinical data such as Major Adverse Cardiovascular Events and patient-reported outcomes, and real-world data such as post-market telemetry and complaint logs are bound together by the risk management file, and claims in the clinical evaluation must align with verification results and risk entries across documents that use different distributions and vocabularies. Reviewers check for that internal consistency directly. Modern connected devices intensify the demand: a continuous glucose monitor streams raw sensor telemetry, algorithmic confidence scores, and patient-facing clinical outputs simultaneously, and the analytical precision of the sensor dictates the clinical decision it supports.

When the Algorithm Is the Product: Diagnostics, Prognostics, Monitoring, and AI Inside the Device

Both the pharma and medtech industries use machine learning heavily in R&D. In pharma it accelerates molecule discovery and development; in medtech it does equivalent service in design and engineering. The difference arrives at clinical evaluation. Not least because in medtech, the algorithm can be the product. Diagnostic, prognostic, and monitoring devices have clinical output that is derived from an algorithm which must be clinically validated.

When the model sits inside a medical device, it is the product, and its clinical evaluation splits into genuinely different evidential questions depending on what the algorithm is for. A diagnostic algorithm is evaluated on whether its calls match a reference standard, which is the agreement statistics of the previous section applied to an algorithmic readout, with sensitivity and specificity carrying the claim. A prognostic algorithm is evaluated on calibration and discrimination, whether the predicted risks correspond to observed event rates and rank patients correctly, a framing familiar from risk scores anywhere in medicine but here attached to a product claim that must be re-established after any update. A monitoring algorithm is evaluated on whether the stream’s outputs stay trustworthy over wear time and use, drift and calibration over months rather than a single encounter, which makes the longitudinal data itself part of the claim.

Across all three, version control becomes a first-class statistical variable. Analyses are stratified by hardware revision and software version, with pooling rules specified before data are seen, and training and validation datasets must be strictly independent, since a model evaluated on its own training data reports an optimism that does not survive contact with new patients. The IMDRF framework for Software as a Medical Device organises the work into three pillars, valid clinical association between output and clinical condition, analytical validation of the processing, and clinical validation that the output achieves its intended purpose, with the last often resting on retrospective studies of blinded datasets with known outcomes. Post-market, the monitoring burden follows the product: drift, distribution shift as populations and practice change, and update governance through predetermined change control plans, which function as a statistical contract specifying which modifications are pre-approved, how their performance impact will be assessed, and what update algorithm governs the model over time.

The practical consequence for trial design is version discipline. Because the evaluated object is software that can be updated, the protocol must specify which software version contributes to the primary analysis, stratify or pool across versions only under pre-specified rules, and keep training and validation datasets strictly separate. It is the pre-specification instinct a pharma statistician already has, applied to the version of the product rather than to the analysis plan alone.

For more information on types of medtech product algorithms and how to validate them see here. And for more technical look at ML methods in this context see here.

To understand the difference between AI and traditional algorithms in medtech see here.

References

  • European Parliament and Council. (2017). Regulation (EU) 2017/745 on medical devices (MDR).
  • European Commission. (2016). MEDDEV 2.7/1 revision 4, Guidelines on medical devices, Clinical evaluation.
  • Schuirmann, D. J. (1987). A comparison of the two one-sided tests procedure and the power approach for assessing the equivalence of average bioavailability. Journal of Pharmacokinetics and Biopharmaceutics, 15(6), 657-680.
  • International Medical Device Regulators Forum. (2017). Software as a Medical Device (SaMD): Clinical Evaluation. IMDRF/SaMD WG/N41FINAL:2017.
  • U.S. Food and Drug Administration. (2010). Guidance for the Use of Bayesian Statistics in Medical Device Clinical Trials.
  • Clinical and Laboratory Standards Institute. (2018). EP09c, Measurement procedure comparison and bias estimation using patient samples.

Related Posts