Introduction
Every new surgical device arrives with a hidden variable. The engineers who built it know exactly how it behaves. The surgeon holding it for the first time does not, and may not do for a while. Between the first case and the fiftieth, the device does not change; the operator does. That change is the learning curve, and it is one of the most consistent sources of bias and avoidable harm in device clinical investigations.
The learning curve is well described in the literature, yet it is still routinely under-planned in clinical investigation plans (CIPs). This article sets out what the evidence says, where the trade-offs lie for sponsors and study designers, and what can be done about it.
What a learning curve actually is
Cook, Ramsay and Fayers characterised a learning curve using three features: (1) the initial level of performance (2) the rate at which performance improves and (3) the plateau at which it stabilises [1]. Those three parameters are what a study designer needs to know, and they are rarely known in advance for a novel device.
The scale of the effect varies enormously. In the Southern Surgeons Club analysis of 8,839 laparoscopic cholecystectomies by 55 surgeons, 90% of bile duct injuries occurred within a surgeon’s first 30 cases; surgeon experience was the only significant predictor of an adverse outcome in multivariate analysis [2]. That is a short, steep curve with serious early harm.
At the other end, Vickers and colleagues analysed 4,702 laparoscopic radical prostatectomies across 29 surgeons. Five-year recurrence risk fell from 17% for a surgeon with 10 prior cases to 9% at 750 cases, and improvement accrued more slowly than for open surgery. Surgeons with prior open experience did worse than those who started laparoscopically, suggesting the skills did not transfer [3]. A separate margin analysis found the curve plateaued at roughly 200 to 250 cases [4]. That is a long, shallow curve where a study of a few hundred patients could sit entirely within the learning phase without anyone noticing.
These two precedents bracket the problem. A sponsor planning a device study needs to know which end of that spectrum they are on, and most have no idea.
Why it matters for interpretability

A learning curve affects three things at once.
Patient safety. Early cases carry higher risk. This is an ethical issue before it is a statistical one; ethics committees increasingly expect the CIP to address it directly.
Bias against the new device. In a comparative study the control is usually an established technique that surgeons already perform well. The investigational arm starts at the bottom of a curve. As Cook and colleagues noted, this can distort the true impact of the intervention and lead to false conclusions against the new procedure [1,5]. The ROLARR trial is the clearest recent example. The primary analysis found no significant difference between robotic and laparoscopic rectal cancer surgery. A pre-planned reanalysis adjusting for surgeon experience found an odds ratio for conversion of 0.40 in favour of robotics at the mean experience level, and the authors concluded that without adjustment the trial would have been confounded by learning effects [6].
Generalisability. The opposite problem. If a sponsor over-trains and proctors every case, the study demonstrates what the device can do in ideal conditions rather than what it will do in routine use. FDA’s pivotal study guidance is explicit that the training provided in the study should mirror what will be provided at market, and that if no training will be offered commercially, study personnel should not receive it either [7]. Ramsay and colleagues found that learning curves were rarely assessed formally in health technology assessment and that where they were, the statistical methods and reporting were weak [8]. A regulator or HTA body reading your report will want to know where on the curve your data came from. It’s also worth noting that this is the principal reason why data generated in clinical investigation is not a substitute for proper human factors or usability evaluation.
The trade-offs

There is no design that removes the learning curve. There are only choices about where to put it.
Restrict to experienced operators. Cleaner treatment effect, fewer early complications, faster enrolment per site. The cost is external validity and, for a novel device, the practical problem that nobody is experienced yet.
Include early cases and adjust. Better generalisability and honest data on real-world adoption. The cost is complexity in the statistical analysis plan, larger sample sizes to preserve power and higher early adverse event rates that you will need to explain to the ethics committee and in your clinical evaluation report.
Delay the study until the curve is climbed. Cook and colleagues note this risks loss of surgeon equipoise, since by the time surgeons are proficient they have usually formed a view [5]. It also delays market access.
Heavy sponsor support at every case. Lowers risk and variance. Risks creating an artificial performance level, plus the training burden you document in the study becomes the training burden you owe the market [7].
The IDEAL framework offers a way through. It describes surgical innovation in stages (Idea, Development, Exploration, Assessment, Long-term study) and places learning curve characterisation in the exploration stage, before a comparative trial [9,10]. In device terms this means using the pilot or feasibility study to measure the curve rather than discovering it in the pivotal (sample size permitting).
Mitigation for study designers

The following are practical and, in our experience, achievable for small and mid-sized sponsors.
- Estimate the curve before the pivotal. Use feasibility data, published data on analogous devices and structured elicitation from surgeons. Cook and colleagues have shown surgeons’ beliefs about learning can be captured with a simple questionnaire [11]. FDA expects this to be addressed in the exploratory stage [7].
- Define proficiency in the CIP. State the minimum experience for an investigator to enrol, whether measured in cases, proctored cases or simulator performance. ROLARR mandated a minimum experience level per technique as a surgeon inclusion criterion [6].
- Use roll-in cases where appropriate. Pre-specify that the first n cases per surgeon or site are treated as roll-in and excluded from the primary analysis but retained for safety reporting. State n and its justification.
- Record operator experience as a covariate. Capture each surgeon’s prior experience with the device and with comparator procedures at baseline and periodically. This is straightforward to do and is what made the ROLARR reanalysis possible [6].
- Pre-specify the analysis. Bayesian hierarchical models [5], CUSUM monitoring and simple experience-adjusted models are all established. Put the approach in the statistical analysis plan before unblinding, not after a null result.
- Structure the on-site support and document it. Proctoring, device specialists in theatre and site training logs reduce early risk. They also need to be recorded, because the training delivered in the study defines the training you will be expected to deliver commercially [7]. ISO 14155 requires investigators to be qualified and trained on the device and the CIP, and the fourth edition, published in March 2026, tightens expectations on risk management for study procedures that sit outside routine practice [12].
- Monitor learning during the study. Site-level tracking of operative time, conversions and device-related events against the expected curve gives the data monitoring committee something to act on.
- Report it. Ramsay’s minimum standard is the number and experience of operators and a full description of data collection [8]. Include a learning curve section in the clinical investigation report.
What this does to your data

If you do none of the above, three things follow.
- Early adverse events cluster and inflate the safety signal
- The comparative effect estimate is biased towards the null or against the device
- The evidence base cannot say what the device will do in the hands of the surgeons who will actually use it
A notified body, FDA reviewer or NICE committee can see all three.
If you plan for it, the learning curve becomes reportable evidence in its own right. You can state how many cases proficiency requires, what training achieves it and what outcomes look like on either side of the curve. That is commercially useful and it is what the market will ask for at adoption.
Practice makes perfect. Your study should be designed to show it.
References
- Cook JA, Ramsay CR, Fayers P. Statistical evaluation of learning curve effects in surgical trials. Clinical Trials. 2004;1(5):421-427. doi:10.1191/1740774504cn042oa
- Moore MJ, Bennett CL. The learning curve for laparoscopic cholecystectomy. The Southern Surgeons Club. American Journal of Surgery. 1995;170(1):55-59. doi:10.1016/S0002-9610(99)80252-9
- Vickers AJ, Savage CJ, Hruza M, et al. The surgical learning curve for laparoscopic radical prostatectomy: a retrospective cohort study. Lancet Oncology. 2009;10(5):475-480. doi:10.1016/S1470-2045(09)70079-8
- Secin FP, Savage C, Abbou C, et al. The learning curve for laparoscopic radical prostatectomy: an international multicenter study. Journal of Urology. 2010;184(6):2291-2296. doi:10.1016/j.juro.2010.08.003
- Papachristofi O, Jenkins D, Sharples LD. Assessment of learning curves in complex surgical interventions: a consecutive case-series study. Trials. 2016;17:266. doi:10.1186/s13063-016-1383-4
- Corrigan N, Marshall H, Croft J, Copeland J, Jayne D, Brown J. Exploring and adjusting for potential learning effects in ROLARR: a randomised controlled trial comparing robotic-assisted vs. standard laparoscopic surgery for rectal cancer resection. Trials. 2018;19:339. doi:10.1186/s13063-018-2726-0
- US Food and Drug Administration. Design Considerations for Pivotal Clinical Investigations for Medical Devices. Guidance for Industry, Clinical Investigators, Institutional Review Boards and FDA Staff. November 2013. https://www.fda.gov/media/87363/download
- Ramsay CR, Grant AM, Wallace SA, Garthwaite PH, Monk AF, Russell IT. Statistical assessment of the learning curves of health technologies. Health Technology Assessment. 2001;5(12):1-79. doi:10.3310/hta5120
- McCulloch P, Altman DG, Campbell WB, et al. No surgical innovation without evaluation: the IDEAL recommendations. Lancet. 2009;374(9695):1105-1112. doi:10.1016/S0140-6736(09)61116-8
- Sedrakyan A, Campbell B, Merino JG, Kuntz R, Hirst A, McCulloch P. IDEAL-D: a rational framework for evaluating and regulating the use of medical devices. BMJ. 2016;353:i2372.
- Cook JA, Ramsay CR, Carr AJ, Rees JL. A questionnaire elicitation of surgeons’ belief about learning within a surgical trial. PLoS ONE. 2012;7(11):e49178. doi:10.1371/journal.pone.0049178
- International Organization for Standardization. ISO 14155:2026 Clinical investigation of medical devices for human subjects. Good clinical practice. 4th edition. March 2026. https://www.iso.org/standard/83968.html
