The PMCF decision guide

The PMCF decision guide

Derk Arts, Castor CEO and Founder

Derk Arts

CEO and Founder

On this page

Executive summary. The compliance reality behind the easing

The proposed European medical device simplification removes a reporting wrapper while preserving the statutory obligation to generate clinical evidence. The European Commission proposal COM(2025) 1023 final would eliminate the standalone post-market clinical follow-up (PMCF) evaluation report under Annex XIV Section 7. It would also relax several fixed reporting cadences.[1][2] This proposal leaves the core mandate to proactively collect and evaluate clinical data completely intact (Annex XIV Part B Section 5). The PMCF plan requirements in Sections 6, 6.1, and 6.2 also remain in full force.[8] You must still determine what clinical uncertainty remains. You must select a proportionate proactive method and defend that choice to your notified body.

This distinction changes the practical PMCF question. Instead of asking which report to prepare, you must identify the remaining evidence gap. You need to determine which method can close it and how to generate the evidence efficiently without weakening traceability or scientific control.

Evidence convention. Regulatory statements rest on statutory texts and official guidance.[1][4][8] Performance findings originate from peer-reviewed informatics literature[10][17] or Castor operational benchmarks from live production workflows.[36] Financial results rely on modeled first-principles scenarios. Full provenance appears in Appendix G.

Keep reading, and get the designed PDF

The rest of this guide continues below: the method-selection matrix, the notified-body defense dossiers, the AI control domains, and the full financial model. Enter your work email to keep reading and to receive the designed PDF of the complete guide.

1. The report may disappear while the evidence obligation remains

The European Commission proposal COM(2025) 1023 final has generated understandable optimism across the medical device sector.[1][3] The proposal aims to remove duplicated documentation, relax fixed update cycles, and shift obligations toward a risk-based cadence. These changes reduce the administrative burden for stable legacy devices on the EU market.

You still bear the statutory responsibility to understand the clinical performance and safety of a device throughout its lifecycle. Annex XIV of Regulation (EU) 2017/745 (MDR) continues to mandate proactive clinical data collection.[8] The proposed simplification alters how you document PMCF findings. It leaves the underlying legal obligation to generate clinical evidence completely intact where residual uncertainty remains.

1.1 What is law today

Under current MDR law, PMCF operates as an ongoing and proactive lifecycle process.[8] You must continuously update your clinical evaluation report (CER), risk management file (RMF), and post-market surveillance (PMS) plans based on real-world clinical experience.

COM(2025) 1023 final currently sits as an ordinary legislative proposal under procedure 2025/0404(COD).[2] The European Parliament Committee on Public Health (SANT) is reviewing the proposal. The Rapporteur tabled Draft Report PE787.987 on 30 June 2026, and it awaits a committee decision.[19] Current MDR requirements remain fully applicable until formal trilogue adoption and subsequent entry into force.

1.2 What the proposal changes and what stands intact

Proposed administrative relief (COM(2025) 1023)Clinical obligations that stand intact (MDR)
Standalone PMCF report wrapper removed. Findings documented directly in the CER and technical file (Annex XIV Section 7 replaced).[1]Evidence gaps must still be identified. Annex XIV Part B Section 5 mandates proactive collection and evaluation of clinical data across the lifecycle. Sections 6, 6.1, and 6.2 specify what the PMCF plan must contain.[8]
Fixed update cadence relaxed. Annual CER updates for Class III and implants shift to clinically triggered updates (Article 61(11)).[1]PMCF methods must still be selected. Annex XIV Section 6.2(a) and 6.2(b) require general and specific proactive methods.[8]
PSUR delivery schedule simplified. Biennial updates for Class IIb and III after Year 1, Class IIa on request (Article 86).[1]Method appropriateness must still be justified. Annex XIV Section 6.2(c) mandates a documented defense to notified bodies.[8]
Non-applicability clarified. The proposal adds an explicit justification route for mature, low-risk devices.[1]Non-applicability must still be justified. Under current MDR the post-market surveillance plan must contain a PMCF plan, or a justification as to why PMCF is not applicable, documented in the clinical evaluation (Annex III Section 1.1).[8]

Regulatory basis (per Sanne Derks review). Annex XIV Part B Section 5 establishes the proactive-collection mandate. It states that “when conducting PMCF, the manufacturer shall proactively collect and evaluate clinical data from the use in or on humans of a device which bears the CE marking.” Sections 6, 6.1, and 6.2 govern what the PMCF plan must specify and how methods are chosen and justified.

1.3 What this means for manufacturers

As a practical consequence, you must continue PMCF while designing it more intelligently. You can move resources away from repetitive report formatting. You can direct them toward acquiring clinical data that directly answers unresolved CER and risk-management questions.

1.4 Worked example. The hypothetical drug-eluting coronary stent

Methodological note. To illustrate how the regulatory, methodological, technical, and economic principles in this guide apply in practice, we follow a single hypothetical worked example throughout. We examine a Class IIb/III drug-eluting coronary stent facing an Annex XIV Section 6.2(b) long-term follow-up evidence gap. All protocol parameters, data workflows, and cost models are illustrative. They derive from operational benchmarks across Castor’s PMCF and clinical registry portfolio.

A Class IIb/III drug-eluting coronary stent has strong short-term 12-month pre-market trial data, but long-term target lesion revascularization (TLR) and late stent thrombosis in high-risk diabetic subgroups remain uncharacterized at 5 years. Grounded in a Section 6.1(b) PMCF objective (identifying and analyzing previously unknown risks or side-effects in a subpopulation), the notified body issues a formal condition for CE-mark renewal. You must actively collect 5-year post-market clinical data on this diabetic cohort. The selected proactive method to meet that objective sits under Section 6.2(b) (see Section 2). Throughout this guide, we use this example to show how each regulatory requirement translates into method selection, AI-driven execution with human review, and budget optimization.

Objective and method (per Sanne Derks review). The notified body condition maps to a Section 6.1(b) PMCF objective. The retrospective chart-review method chosen to satisfy it is a Section 6.2(b) specific method. Objective and method are distinct clauses and should not be conflated.

2. Start with the evidence gap, then choose the method

You must base your PMCF method selection on a defined clinical uncertainty in the CER or risk-management file that requires resolution. You cannot select a method simply because a registry, survey platform, or extraction tool happens to be available.

A defensible PMCF plan begins with four questions.

  • Clinical uncertainty. What safety, performance, or benefit-risk question remains unresolved?
  • Population uncertainty. Which subgroup, clinical indication, use setting, or user group is underrepresented?
  • Temporal uncertainty. Is short-term evidence adequate while long-term survivorship or wear data are absent?
  • Decision threshold. What quantitative result would confirm, or challenge, the current benefit-risk profile?

2.1 The method-selection hierarchy (Annex XIV Section 6.2)

MDR Annex XIV Section 6.2 establishes two tiers of proactive clinical data collection, plus a justification duty.[8]

  • Section 6.2(a) general methods. Gathering routine clinical experience, feedback from healthcare professionals (eCOA and surveys), and screening scientific literature.[4] This approach remains proportionate when clinical uncertainty is low and the technology is mature.
  • Section 6.2(b) specific methods. Evaluation of suitable device registries, prospective clinical investigations, and retrospective chart reviews.[5] These methods are necessary when objective, patient-level endpoints are required to evaluate residual risks.
  • Section 6.2(c) documented justification. You must present a documented justification explaining why the chosen methods are adequate to fulfil the statutory objectives of Section 6.1.[8]

2.2 The evidence-gap-to-method pathway

Residual evidence gapBest available data sourceCandidate PMCF methodMandatory fitness checksMinimum defense deliverable
Mature device, zero residual design risks, known usability.High-volume surgical users, published literature.Literature screening plus HCP eCOA surveys (6.2(a)).[4]Response-rate validity (at or above 80%), representative user sampling, absence of complaints.Screening protocol, validated survey tool, statistical summary report.
High-quality national registry exists with UDI tracking.National clinical quality registers (for example NJR, BCIR).Registry extraction and linkage (6.2(b)).[5][27]Data completeness (above 90%), device-level UDI capture, longitudinal follow-up.Registry governance agreement, extraction protocol, data audit.
Patient-reported outcomes, recovery quality, pain scores.Direct patient follow-up cohorts.Digital ePRO / eCOA follow-up (6.2(a)).[28]Long-term retention (above 70%), documented Articles 6 and 9 basis, informed consent where required, validated scales.Validated ePRO audit trail, compliance logs, retention strategy.
Implantable device, long-term safety and survivorship gaps.Hospital electronic health records, surgical notes.Multi-center retrospective chart review (6.2(b)).[10][13]Source-document availability, certified-copy integrity, 100% source provenance.GCP-compliant protocol, eSource audit trail, statistical analysis plan.
Novel mechanism of action, unobservable in routine care.Prospectively enrolled clinical cohort.Prospective post-market clinical study (6.2(b)).[8]Ethical-approval feasibility, site recruitment capacity, protocol compliance.ISO 14155 GCP clinical investigation plan, monitoring dossier.[13]

ISO 14155 scope (per Sanne Derks review). Registries and any other non-interventional but prospective studies are within the scope of ISO 14155. A study is treated as prospective for the regulatory framework whenever it carries any prospective element, which affects the applicable compliance obligations.

2.3 Worked example. Defending the stent PMCF strategy

Applying this method-selection logic to the worked coronary stent example.

  • Evidence gap. 5-year cumulative TLR and late stent thrombosis in diabetic patients are uncharacterized in the pre-market dossier.
  • Rejected alternatives. Literature screening (6.2(a)) is rejected because published papers lack device-specific serial tracking in diabetic subsets. A new prospective randomized trial is disproportionate because the device is already widely implanted and objective source records exist across hospital cath labs.
  • Selected method. Multi-center retrospective chart-review cohort (N=200 patients across 10 centers) linked to a national cardiovascular registry (6.2(b)).[10][27]

Primary estimand and acceptance rule (illustrative). In this single-arm retrospective cohort of N=200 diabetic patients evaluated against a historical control benchmark (for example a historical 5-year TLR rate of 4.2% with a pre-specified non-inferiority margin of 2.5%), the primary objective is to demonstrate that the upper bound of the one-sided 95% confidence interval for 5-year TLR remains at or below 6.5%. Cumulative 5-year MACE at or below 7.5% serves as a descriptive secondary endpoint. Final sample size and power must be formally established in the statistical analysis plan. The N=200 figure serves as a standardized operational benchmark for study execution and budget modeling.

Model Section 6.2(c) defense language. “Pursuant to MDR Annex XIV Part B Section 6.2(c), specific proactive clinical data collection is executed via a multi-center retrospective cohort (N=200) linked to national cardiovascular registers under Section 6.2(b). General surveys under Section 6.2(a) are insufficient to detect low-frequency 5-year stent thrombosis. The retrospective design captures longitudinal hard endpoints, while AI-driven source extraction with 100% human review mitigates manual transcription error and preserves full source provenance.”

2.4 The execution bottleneck in registry and custom-cohort studies

Selecting a defensible method solves the regulatory design problem, but executing registry extraction and custom PMCF cohorts (site uploads) creates a severe operational bottleneck. When sponsors extract longitudinal outcomes from hospital EHRs, cath-lab records, pathology reports, and paper source worksheets, clinical research coordinators spend hundreds of hours retyping data into eCRFs. The next section introduces Castor Catalyst. This AI-driven extraction and site-upload engine with human review removes this bottleneck while maintaining 100% human-in-the-loop auditability.

3. Castor Catalyst. Regulated AI-driven extraction with human review and site uploads for PMCF

3.1 The source-data bottleneck

For data-intensive PMCF studies, the main operational failure mode occurs during manual chart abstraction. Clinical research coordinators and abstraction teams hit several failure points when reviewing unstructured inpatient and outpatient records.

  • High transcription discordance. Systematic reviews show manual chart abstraction error rates averaging 6.6%, ranging up to 14.8% across unstructured hospital records.[10] Variable definitions drift across sites, and complex endpoints suffer from subjective interpretation.[12]
  • Source data verification (SDV) has shifted from 100% to risk-based. Historically, when 100% on-site SDV was standard, it consumed a large share of trial budgets while rarely changing primary-endpoint conclusions. Risk-based monitoring has since reduced SDV, and ISO 14155 requires the monitoring plan to justify the amount of SDV, allowing reduced or, where risks are mitigated by documented training, meetings, or extensive guidance, no on-site SDV.
  • Severe query latency. In traditional site-based models, queries are generated weeks or months after data entry. Resolving a single disputed variable costs an average of $180 per resolved query in cross-functional labor.

SDV reframing (per Sanne Derks review). The prior 2012 figure (“15 to 30% of budgets, alters conclusions in under 0.1% of cases”) reflected a 100%-SDV era and has been removed. Risk-based monitoring is now standard, and ISO 14155 requires the monitoring plan to justify the amount of SDV, up to and including no on-site monitoring where risks are adequately mitigated.

3.2 The Catalyst architecture. Site uploads and direct-to-patient

  • Site uploads (global). Clinical research coordinators drop raw source files, including unstructured PDF discharge summaries, scanned catheterization logs, operative notes, and paper PRO worksheets, into a secure portal. Catalyst parses the documents, extracts protocol-specific variables, and maps them to standard CDISC CDASH domains.
  • Direct-to-patient (US). Consented participants authorize automated medical-record retrieval via standard FHIR API gateways and national health information networks (TEFCA), pulling structured and unstructured histories into the study repository without site-staff intervention.

Castor Catalyst architecture

1. Ingestion

  • Hospital EHR PDFs and scans
  • Cath-lab operative notes
  • Paper PRO worksheets
  • Direct FHIR / TEFCA feeds

2. AI processing

  • Layout-aware vision OCR
  • Protocol-specific extraction
  • Calibrated confidence (0 to 1)
  • MedDRA and WHODrug coding

3. Clinical QC

  • 100% human-in-the-loop review
  • Side-by-side highlight UI
  • Interactive bounding boxes
  • 0.8% reviewer override rate

4. Compliance

  • Castor EDC and CDMS
  • 21 CFR Part 11 and EU Annex 11 audit trail
  • Remote SDV-ready trace
  • Direct export to CER and PSUR

Governance: deterministic execution (temperature 0), GAMP 5 Category 4/5, zero model training on client data. Standards: ISO 14155, MDR Annex XIV, FDA RWE guidance.

3.3 The controlled six-stage AI-driven extraction pipeline with human review

Catalyst operates through a deterministic, six-stage pipeline designed to satisfy ISO 14155 GCP and to align with the FDA/EMA good AI practice principles (applied by analogy).[13][31]

  1. Document ingestion and verification. Raw source documents are ingested, hashed (SHA-256), and validated for readability. Degraded scans (under 0.2%) route to a manual intake queue.
  2. Pre-processing and layout-aware OCR. Documents are parsed with multimodal vision-language OCR, preserving tabular structures, handwritten annotations, and temporal sequencing.
  3. Protocol-aware variable extraction. The engine extracts only variables prespecified in the study protocol (for example stent diameter, bifurcation involvement, post-procedure TIMI flow, 5-year TLR events), generating reasoning chains and calibrated confidence scores (0.00 to 1.00).
  4. Automated medical coding. Unstructured adverse events, comorbidities, and medications map to standard ontologies (MedDRA, WHODrug, ICD-10), and device deficiencies and problems map to IMDRF device-event terminology, with semantic proximity scores.
  5. 100% human-in-the-loop adjudication. Every extracted value is presented to a medically qualified reviewer in a side-by-side interface with an interactive source bounding box. The reviewer verifies, edits, or rejects the proposal before data lock.
  6. Commit to a 21 CFR Part 11 and EU Annex 11 aligned EDC. Approved values are written to the Castor EDC with an immutable audit log linking field value, user ID, timestamp, and source-document coordinates.

Governance. Deterministic execution (temperature 0), GAMP 5 Category 4/5, zero model training on client data. Standards include ISO 14155, MDR Annex XIV, and FDA RWE guidance.

3.4 Validated production benchmarks

Castor Catalyst performance spotlight  ·  verified production telemetry

Controlled AI with expert oversight, at enterprise speed

0.8%

Reviewer override rate (99.2% accepted as extracted)[36]

6 vs 39 min

Per-chart abstraction ($9.75 vs $48.75)[36]

99.8%

Clean document ingestion, 0.2% routed to manual queue[36]

In live production deployments of the Catalyst site-uploads workflow and internal benchmarking against real EMR records, Castor observed a 0.8% adjudication override rate. Medically trained reviewers accepted or confirmed 99.2% of AI-driven proposals, compared to a 6.6% manual chart-abstraction baseline.[10] The system delivered 100% human-in-the-loop completion and a 99.8% clean ingestion rate. We also documented a five- to sixfold abstraction acceleration, dropping time from 39 minutes to 6 minutes per chart. This lowered direct abstraction cost from $48.75 to $9.75 per chart, with zero model training on client data.[36]

Across the Castor platform, device teams run 893 MedTech studies to date (289 live) and 298 post-market studies including PMCF (98 live), the footprint behind the operational telemetry above.[36]

3.5 Five non-negotiable AI control domains

To defend AI-driven PMCF data with human review to notified bodies and health authorities, you must enforce five governance controls.[31][32]

  • Auditability and traceability. Every data point must trace directly to a highlighted bounding box in a certified source document, satisfying ISO 14155 source-data integrity (ALCOA+).[13]
  • Determinism and model versioning. Model weights, system prompts, temperature (T=0), and extraction schemas are version-locked and validated under GAMP 5 Category 4/5.
  • 100% human-in-the-loop verification. Autonomous database commits are prohibited for safety and efficacy endpoints. Human clinicians retain final sign-off.
  • Data privacy and contractual zero-retention. Client clinical data is processed in dedicated isolated environments and is contractually prohibited from training public foundation models.
  • Continuous calibration and quality audits. Inter-rater reliability (Cohen’s kappa at or above 0.90) and statistical error sampling are audited across sites throughout the study lifecycle.

3.6 Worked example. Applying Catalyst AI controls to the stent study

In the worked stent example, the AI-driven system extracts 60,000 target fields from multi-center cath-lab records for human review. Sites upload de-identified operative reports, angiographic discharge summaries, and diabetic clinic follow-up notes via Catalyst site uploads. Catalyst extracts lesion length, reference vessel diameter, stent deployment pressures, and follow-up ischemic events. It maps adverse cardiac events to MedDRA preferred terms. A centralized core-lab research nurse reviews every extracted variable side-by-side against the highlighted operative PDF bounding box. You export an ALCOA+, source-linked dossier demonstrating that 100% of 5-year TLR events are verified against certified hospital source documents, with an adjudicated override rate under 1.0%.[36]

4. The economic reality. Comparing PMCF operational models

With the Catalyst architecture and human-in-the-loop controls established, how does this translate into study budgets? This section presents a first-principles economic comparison across three operational models for a standardized post-market study.

4.1 The three-column first-principles model

We model a standard data-intensive PMCF study. It includes N=200 patients across 10 clinical centers, 5-year follow-up, and 300 variables per patient (60,000 total target fields).

$1,102,000

Model A. Traditional on-site ($5,510 / patient)

$789,000

Model B. Centralized manual ($3,945 / patient)

$551,000

Model C. Castor Catalyst AI + HITL ($2,755 / patient)

4.2 Line-item budget decomposition and net savings

Budget line itemA: traditionalB: centralized manualC: Catalyst AI + HITLPrimary driver of efficiency
Protocol and regulatory design$45,000$45,000$45,000Standardized regulatory dossier and SAP.
Site identification and contracting$100,000$60,000$60,000Centralized contracting, no EDC training at sites.
EDC build and system validation$55,000$55,000$45,000Pre-validated extraction schemas, CDISC CDASH.
Site management and CRA monitoring$180,000$120,000$96,000Remote SDV replaces routine CRA travel.
Direct chart abstraction labor$120,000$100,000$70,0006.67 hrs/patient manual reduced to a fixed $350/patient AI + HITL fee (a $50,000, 41.7% direct-labor saving).
Data management and query resolution$145,000$85,000$45,000AI-driven pre-validation with human review eliminates transcription syntax errors.
Medical writing, CER and technical file$65,000$50,000$40,000Automated structured evidence export into the CER.
Project management and governance$92,000$74,000$50,000Accelerated timeline reduces monthly CRO fees.
Site grant fees ($1,500/patient)$300,000$200,000$100,000Reduced site burden lowers per-patient compensation.
Total fully loaded study budget$1,102,000$789,000$551,00050.0% net saving vs traditional, 30.2% vs centralized manual.
Per-patient fully loaded cost$5,510$3,945$2,755Capital efficiency from combining centralization with AI-driven extraction and human review.

4.3 Direct chart-abstraction labor reconciliation

For the traditional and centralized manual build, 300 variables per patient across multi-year unstructured records take an average of 6.67 hours of coordinator labor per patient (about 1.33 minutes per variable). At a fully loaded $90/hour, manual abstraction costs $600 per patient, or $120,000 across 200 patients. For the Catalyst build, the AI-driven system ingests source files and extracts 300 variables in seconds for human review. A centralized research nurse then reviews the highlighted bounding boxes in about 45 minutes per patient. Including platform software and managed HITL review, the fully loaded cost is $350 per patient, or $70,000 across 200 patients. This creates a $50,000 (41.7%) direct-labor reduction.

4.4 Query cost reconciliation and friction bridge

At a 6.6% error rate,[10] manual abstraction produces about 3,960 discrepant data points. Internal validation rules catch about 2,000 syntax errors. The remaining 900 complex clinical queries require formal cross-functional resolution. Resolving 900 queries at $180 each costs $162,000 in gross operational friction. With a 0.8% override rate and AI-driven pre-commit validation with human review, complex queries drop to about 200. This reduces gross friction to $36,000, representing a 77.8% friction saving.

4.5 Sensitivity across sample sizes

Sensitivity modeling across cohort sizes shows consistent behavior. N=100 saves 49.3% ($345k vs $680k), and N=500 saves 51.8% ($1.18M vs $2.45M). Studies capturing more than 150 variables per patient from unstructured hospital records achieve the highest return.

4.6 Worked example. Stent study budget application

For the worked stent study (N=200 diabetic patients, 10 centers, 60,000 target fields), applying the AI-driven Catalyst system with human review reduces the fully loaded 5-year PMCF budget from $1,102,000 to $551,000. This is a $551,000 (50.0%) reduction. Direct abstraction labor falls from $120,000 to $70,000. Gross query friction drops from $162,000 to $36,000. This lets you maintain CE-mark compliance within commercial margin targets.

5. Design the study once. Use the evidence responsibly.

Clinical evidence generated for MDR PMCF should not stay trapped in a single regulatory silo. When structured correctly, a governed real-world dataset can serve multiple commercial and scientific objectives.

5.1 Three strategic evidence-reuse pathways

  • MDR ongoing compliance. Direct integration into the CER, PSUR, and RMF under MDR Articles 61 and 86.[8]
  • HTA and reimbursement. Submission of longitudinal real-world survivorship and complication data to national reimbursement bodies (for example France’s HAS, Germany’s G-BA, the UK’s NICE).[33]
  • Evidence-based claims under Article 7 MDR. Supporting narrowly defined clinical claims in strict compliance with the Article 7 non-misleading requirements.[8]

5.2 Six conditions for evidence reuse

To reuse real-world PMCF data legally and scientifically, sponsors establish six governance controls at protocol inception. These include prospective estimand prespecification, full source traceability, data completeness and quality thresholds (above 90% completeness, under 1.0% error rate[36]), transparent bias quantification, a defensible GDPR compliance architecture,[9][35] and regulatory consistency with published SSCP files.[8]

5.3 GDPR processing-operation and governance framework

Governance note. Illustrative allocation only. Controller and processor roles, the Articles 6 and 9 conditions, and applicable national-law requirements must be confirmed for each processing operation based on the actual purposes and means of processing. The primary PMCF conduct rests on the MDR legal obligation (Art 6(1)(c)) with Art 9(2)(i). Where processing is for secondary scientific research beyond that legal duty, the Article 6(1)(e)/(f) and 9(2)(j) research basis applies (row 3). Reference [9] (EDPB Opinion 3/2019) is a clinical-trials-regulation opinion applied here by analogy.

Processing operationControllerProcessorLawful basis (Art 6 / 9)Core safeguards
1. PMCF study conduct and abstractionDevice sponsorTechnology vendorArt 6(1)(c) (MDR duty), Art 9(2)(i) (public health / safety)Art 4(5) pseudonymization at site, unique study codes, strict RBAC.
2. Vigilance and safety reportingDevice sponsorClinical siteArt 6(1)(c) (MDR duty), Art 9(2)(i) (high safety standards)Secure E2B XML reporting, pseudonymized adverse-event logs.
3. Secondary HTA and reimbursement analysisDevice sponsorAcademic partnerArt 6(1)(e)/(f) (public / legitimate), Art 9(2)(j) (scientific research)Fully anonymized dataset, cell suppression for low-frequency categories.
4. Technical support and system hostingDevice sponsorTechnology vendor (sub-processors)Processing under the controller’s documented instructions and underlying lawful basis. Art 28 DPA controls apply.Zero PHI/PII visibility, encrypted logs, Chapter V SCCs and TIAs.[35]
5. AI model training and fine-tuningProhibitedProhibitedProhibited (strict zero retention)Client clinical data is segregated and never used for foundation-model training.

5.4 Worked example. Maximizing the stent evidence asset

For the worked stent example, the 5-year retrospective dataset resolves the CER safety question for the notified body. Because the protocol prospectively prespecifies the long-term outcomes to be extracted and preserves complete source provenance, you can reuse the governed evidence asset for a purpose-specific HTA analysis (for example with France’s HAS or the UK’s NICE[33]). You can evaluate whether it supports a narrowly defined external claim under Article 7 MDR. This potentially avoids a separate seven-figure evidence-generation study.

6. A practical 90-day path to a defensible PMCF model

Transitioning to a modern, AI-driven PMCF operating model with human review does not require an immediate overhaul of the entire portfolio. You should follow a phased 90-day roadmap.

  • First 30 days (portfolio triage and risk mapping). Audit active devices against CER residual-risk tables, categorize into 6.2(a) surveys versus 6.2(b) cohorts, and identify the top one or two devices facing upcoming notified-body renewal deadlines.
  • Days 31 to 60 (evidence and method design). Author the PMCF plan per MDCG 2020-7,[4] define estimands and the SAP, conduct AI vendor due diligence (Appendix D), and draft DPIAs and Article 28 DPAs.
  • Days 61 to 90 (controlled pilot). Launch a multi-center pilot (1 to 2 sites, 20 to 30 records), validate the Catalyst site-uploads workflow, and audit inter-rater reliability, confidence calibration, and bounding-box precision.
  • After 90 days (scale selectively). Expand the validated workflow across centers, standardize extraction schemas across product families, and integrate structured data into annual PSURs and CER dossiers.

6.1 Three executive leadership decisions

  • Method-selection governance. Mandate that CER residual-risk gaps drive PMCF method selection rather than historical inertia.
  • Standardized AI vendor criteria. Require GAMP 5 validation, deterministic extraction logs, and contractual zero-retention guarantees from every external software vendor.
  • Cross-functional asset strategy. Establish a unified clinical-evidence council bridging Regulatory, Clinical Affairs, HEOR, and Commercial to maximize the downstream value of every PMCF dataset.

7. Beyond static reports to living clinical evidence

COM(2025) 1023 final would signal a welcome maturation of the European medical device framework. By removing redundant report wrappers and shifting toward risk-based cadences, the proposal creates space for you to focus on substantive clinical science. The statutory duty to demonstrate long-term safety and performance under MDR Annex XIV Part B remains. If you rely on manual chart transcription and fragmented site monitoring, you will keep facing cost inflation and audit vulnerability. By pairing rigorous Annex XIV method selection with centralized operations and validated AI-driven extraction with 100% human-in-the-loop review, you can turn PMCF from a costly compliance burden into a permanent, high-trust clinical-evidence asset.

Frequently Asked Questions

No. COM(2025) 1023 would delete the standalone PMCF evaluation report under Annex XIV Section 7 and relax some reporting cadences, but the mandate to proactively collect and evaluate clinical data (Annex XIV Part B Section 5) and the PMCF plan requirements in Sections 6, 6.1, and 6.2 remain fully in force. Until the proposal is adopted and enters into force, current MDR requirements apply.

Annex XIV Part B Section 5. It requires the manufacturer to proactively collect and evaluate clinical data from the use of a CE-marked device across its lifetime. Sections 6, 6.1, and 6.2 govern what the PMCF plan must specify and how methods are chosen and justified.

Start from the residual evidence gap in the clinical evaluation, select a proportionate method under Annex XIV Section 6.2 (general methods under 6.2(a) or specific methods under 6.2(b)), and provide the documented justification required by Section 6.2(c) explaining why the chosen method meets the Section 6.1 objectives.

Yes, with human review. AI-driven source extraction is defensible when every extracted value is reviewed by a qualified person before it enters the record (100% human-in-the-loop), with full source traceability, deterministic version-locked execution, and ISO 14155 source-data integrity.

In a first-principles model of a 200-patient, 10-center, 5-year study, a traditional on-site build runs about $1,102,000, a centralized manual build about $789,000, and Castor Catalyst with human review about $551,000, a 50% reduction versus traditional. The savings come from the data work, direct abstraction labor and query friction, not the study design.

Appendix

Appendix A. Legislative status and SANT delta

Specific amendments proposed under COM(2025) 1023 final and European Parliament SANT Committee Draft Report PE787.987 (30 June 2026).[1][19]

MDR provisionCurrent MDR (2017/745)Commission proposal COM(2025) 1023SANT draft report (PE787.987)Operational impact
Annex XIV Part B Section 5 (proactive-collection mandate)Proactive collection mandatory. Non-applicability must be justified and documented in the clinical evaluation (Annex III Section 1.1).[8]Adds an explicit non-applicability justification route for mature, low-risk devices. The collection mandate is unchanged.[1]Requires documented clinical proof of zero residual risk per Annex III Section 1.1.[19]Notified bodies require rigorous clinical proof before granting non-applicability.
Annex XIV Part B Section 6 (PMCF plan and duties)PMCF plan, objectives, and method justification mandatory.[8]Unchanged. Planning and method selection remain mandatory.[1]Emphasizes risk-proportionate application of specific versus general methods.[19]Core plan obligations remain fully intact.
Annex XIV Part B Section 7 (PMCF report)Standalone annual PMCF report mandatory for Class IIa, IIb, III.[8]Deletes Section 7. Findings documented in the CER.[1]Endorses replacement, mandates clear cross-referencing in technical files.[19]Eliminates standalone report formatting.
Article 61(11) (CER cadence)Annual CER updates for Class III and implants.[8]Shifts to clinically triggered updates.[1]Clarifies implants reviewed at least biennially.[19]Reduces calendar reporting, requires continuous PMS signal detection.
Article 86 (PSUR)Annual PSUR for Class IIb/III, biennial for IIa.[8]Biennial for IIb/III after Year 1, IIa on request.[1]Clarifies electronic submission via EUDAMED.[19]Aligns PSUR delivery with CER updates.

Dossier 1. Class IIa reusable orthopaedic surgical reamer

Reusable rotary surgical reamer for total hip arthroplasty (Class IIa, Rule 6). Mature technology on the EU market for more than 12 years with more than 250,000 lifetime uses. Residual gap. Long-term cutting efficiency, dimensional wear, and sterilization resilience after more than 50 autoclave cycles. Selected method. Systematic literature screening plus a validated surgeon eCOA survey at high-volume centers (6.2(a)).[4] Acceptance. N=60 surgeons across 4 member states, at or above 85% completion, at or above 90% satisfaction, zero intraoperative fractures. Model 6.2(c) defense. “Pursuant to MDR Annex XIV Part B Section 6.2(c), PMCF is executed via high-volume surgeon eCOA survey and systematic literature surveillance under Section 6.2(a). Specific clinical investigations under Section 6.2(b) are unjustified as the reamer is a well-established Class IIa technology with well-characterized clinical risks.”

Dossier 2. Class IIb AI-based software for stroke triage (SaMD)

Standalone SaMD deep-learning algorithm (Class IIb, Rule 11) analyzing emergency head CT scans for large-vessel-occlusion ischemic stroke and hemorrhage. Residual gap. Real-world sensitivity and specificity across heterogeneous CT hardware and elderly patients with leukoaraiosis. Selected method. Multi-center retrospective imaging database extraction and diagnostic-accuracy evaluation (6.2(b)).[5][10] Acceptance. N=500 consecutive emergency stroke CTs across 5 centers with neuroradiologist consensus ground truth, lower bound of the 95% CI sensitivity above 90.0% (target at or above 93.0%), specificity above 88.0% (target at or above 91.0%). Model 6.2(c) defense. “Pursuant to MDR Annex XIV Part B Section 6.2(c), clinical follow-up is conducted via multi-center retrospective diagnostic-accuracy evaluation (N=500) under Section 6.2(b). General surveys under Section 6.2(a) cannot validate diagnostic accuracy.”

Device risk classResidual gapPrimary data sourceRecommended methodPMCF method clauseEst. budgetDuration
Class I / IIa (mature)Usability, routine feedbackClinicians, published literatureLiterature screening + HCP survey[4]Annex XIV 6.2(a)$15k to $40k2 to 4 months
Class IIa / IIb (orthopaedic)Long-term survival, UDI trackingNational registries (NJR, EPRD)[27]Registry extraction and linkage[5]Annex XIV 6.2(b)$80k to $220k6 to 12 months
Class IIa / IIb (digital / SaMD)Patient recovery, quality of lifeEnrolled patient app usersDigital ePRO / eCOA cohort[28]Annex XIV 6.2(a)$50k to $140k4 to 8 months
Class IIb / III (implantable)Survivorship, rare complications, subgroupsHospital EHRs, surgical records[10]Multi-center retrospective chart review[13]Annex XIV 6.2(b)$200k to $550k6 to 14 months
Class III / novel techUncharacterized biological responseDe novo enrolled clinical sitesProspective PMCF clinical study[8]Annex XIV 6.2(b)$850k to $2.5M+18 to 36 months
Control domainVerification requirementMandatory deliverableRegulatory benchmark
1. Validation and QMSGAMP 5 validation, ISO 13485 / ISO 9001 quality system.IQ/OQ/PQ dossier, system traceability matrix, release reports.GAMP 5 Cat 4/5, 21 CFR 11, MDR Annex IX.
2. TraceabilityDeterministic execution (T=0), visual source bounding-box audit trail.Model architecture spec, bounding-box verification demo.ISO 14155 source-data integrity, FDA RWE guidance.[13][34]
3. Human-in-the-loop100% human verification workflow, medical qualification logs.Reviewer UI workflow docs, adjudication role SOPs.FDA/EMA good AI practice, MDCG 2020-7.[4][31]

Operational basis. N=200 patients across 10 centers, 5-year follow-up, 300 variables/patient (60,000 total target fields across operative notes, cath logs, and outpatient visits). Traditional decentralized model. Site activation $10k/site ($100k), CRA monitoring and travel $18k/site ($180k), labor 6.67 hrs/patient at $90/hr ($120k), 900 complex queries at $180 ($162k gross friction), site grants $1,500/patient ($300k), total $1,102,000 ($5,510/patient). Centralized manual model. Centralized onboarding ($60k), remote monitoring ($120k), manual abstraction $100k, DM and queries $85k, site grants $1,000/patient ($200k), total $789,000 ($3,945/patient, saves 28.4%). Castor Catalyst AI + HITL model. Pre-validated schemas ($45k), remote SDV ($96k), Catalyst software plus managed HITL review fixed $350/patient ($70k, 41.7% direct-labor saving), post-HITL query friction $36k (77.8% saving), site grants $500/patient ($100k), total $551,000 ($2,755/patient, 50.0% net saving). Sensitivity. Across N=100 to N=500, total savings range 49.3% to 51.8% versus traditional on-site execution.

Governance note. Illustrative allocation only. Controller and processor roles, the Articles 6 and 9 conditions, and applicable national-law requirements must be confirmed for each processing operation. This matrix expands Section 5.3 with detailed privacy safeguards per operation.

Processing operationControllerProcessorLawful basis (Art 6 / 9)Privacy safeguards and controls
1. Primary study conduct and abstractionDevice sponsorTechnology vendorArt 6(1)(c) (MDR duty), Art 9(2)(i) (public health / safety)Pseudonymization at the clinical site, unique patient study codes, role-based access control.
2. Vigilance and safety reportingDevice sponsorClinical siteArt 6(1)(c) (MDR duty), Art 9(2)(i) (high safety standards)Secure E2B XML reporting channels, pseudonymized adverse-event logs.
3. Secondary HTA and reimbursement analysisDevice sponsorAcademic partnerArt 6(1)(e)/(f) (public / legitimate), Art 9(2)(j) (scientific research)Fully anonymized dataset generation, cell suppression for low-frequency categories.
4. Technical support and system hostingDevice sponsorTechnology vendor (sub-processors)Processing under the controller’s documented instructions and underlying lawful basis. Art 28 DPA controls apply.Zero access to clinical records, encrypted logs, Chapter V SCCs and TIAs.[35]
5. AI model training and fine-tuningProhibitedProhibitedProhibited (strict zero retention)Client clinical data is strictly segregated and never used for foundation-model training.

Evidence tiers. Tier 1 (statutory texts and regulatory authorities). Regulation (EU) 2017/745 (MDR),[8] COM(2025) 1023 final,[1] SANT Committee Report PE787.987,[19] MDCG guidance (2020-5, 2020-7, 2020-8, 2021-24, 2022-21, 2024-10), ISO 14155,[13] FDA guidances,[25][34] and joint FDA/EMA good AI practice.[31] Tier 2 (Castor production benchmarks[36]). Operational telemetry from live Catalyst site-uploads deployments. Definitions. 100% HITL review = every extracted field reviewed and approved by a qualified clinical professional before commit. 0.8% override rate = adjudicated human edits divided by total verified fields. 99.8% clean ingestion = source documents parsed without exception-queue escalation. Time and cost efficiency = 6 minutes vs 39 minutes per chart ($9.75 vs $48.75). Tier 3 (modeled financial scenarios). First-principles bottom-up cost model for an illustrative 200-patient, 10-center, 5-year study.

References

  1. European Commission. COM(2025) 1023 final, 16 Dec 2025; accompanied by SWD(2025) 1050, 1051, 1052.
  2. European Parliament. Procedure 2025/0404(COD): targeted simplification of certain post-market requirements. Legislative Observatory, Aug 2026.
  3. European Commission. Study on the availability of medical devices on the EU market (3rd EO Survey). DG SANTE, 2025/2026.
  4. Medical Device Coordination Group. MDCG 2020-7: guidance on PMCF plan template. European Commission, Sept 2020.
  5. International Medical Device Regulators Forum. IMDRF MDCE WG/N65FINAL:2021: post-market clinical follow-up studies. IMDRF, March 2021.
  6. Medical Device Coordination Group. MDCG 2020-8: guidance on PMCF evaluation report template. European Commission, Sept 2020.
  7. Medical Device Coordination Group. MDCG 2020-5: guidance on clinical evaluation, equivalence. European Commission, April 2020.
  8. Regulation (EU) 2017/745 on medical devices (MDR). OJ L 117, 5.5.2017, p. 1 to 175; consolidated text 2026.
  9. European Data Protection Board. EDPB Opinion 3/2019 on CTR and GDPR. Adopted 23 Jan 2019.
  10. Garza M, Del Fiol G, Tenenbaum J, et al. Comparing medical record abstraction error rates: a systematic review and meta-analysis. BMC Med Res Methodol. 2024;24:304. doi:10.1186/s12874-024-02424-x.
  11. Timbie JW, et al. Five-year clinical outcomes of ICD therapy: a post-market evaluation. Am Heart J. 2023;258:45 to 56.
  12. Bauer CA, et al. High discordance in manual abstraction of complex clinical variables from unstructured inpatient EHRs. J Clin Epidemiol. 2019;112:54 to 62.
  13. International Organization for Standardization. ISO 14155: GCP for medical device clinical investigations. Related: ISO 14971:2019, ISO/TR 20416:2020.
  14. Okafor B, et al. Health economic evaluation of national arthroplasty registers. Arch Orthop Trauma Surg. 2025;145:408.
  15. Mauch V, et al. Long-term operational sustainability and data completeness in a device registry: lessons from IROS. Trials. 2021;22:845.
  16. Poser M, et al. Improving reliability and accuracy of structured data extraction using consensus LLMs in multiple sclerosis. Front Artif Intell. 2026;9:1658575.
  17. Wornow M, et al. Evaluating clinical foundation models for phenotyping across longitudinal EHRs. NPJ Digit Med. 2024;7:58.
  18. European Parliament. Draft report on the proposal for a regulation amending MDR/IVDR. Committee on Public Health (SANT), Rapporteur draft PE787.987, 30 June 2026.
  19. European Commission. Presentation of results of the 3rd study on device availability. DG SANTE, 2025/2026.
  20. Team-NB. Team-NB survey 2025: sector overview on notified body capacities and transition rates. Team-NB, 2025.
  21. SNITEM. Observatoire du Reglement General sur les Dispositifs Medicaux: enquete de conjoncture 2024-2025. SNITEM, 2024.
  22. Medical Device Coordination Group. MDCG 2022-21: guidance on PSUR according to MDR. European Commission, Dec 2022.
  23. Medical Device Coordination Group. MDCG 2024-10: clinical evaluation of orphan medical devices. European Commission, June 2024.
  24. U.S. Food and Drug Administration. Electronic systems, electronic records, and electronic signatures in clinical investigations: questions and answers. Guidance for Industry, October 2024. Docket FDA-2017-D-1105.
  25. Garza M, et al. Advancing clinical trial efficiency and data accuracy through direct EHR-to-EDC integration. AMIA Joint Summits. 2026; PMID 42317828.
  26. Sedrakyan A, et al. Coordinated registry networks for real-world evidence generation. BMJ Surg Interv Health Technol. 2022;4(Suppl 1):e000123.
  27. U.S. Food and Drug Administration. Principles for selecting, developing, modifying, and adapting patient-reported outcome instruments for use in medical device evaluation. Guidance for Industry and FDA Staff, Jan 2022. Docket FDA-2020-D-1564.
  28. Medical Device Coordination Group. MDCG 2021-24: guidance on classification of medical devices. European Commission, Oct 2021.
  29. TransFAIR Study Group. TransFAIR study: EHR2EDC technology compared to manual eCRF collection. BMJ Health Care Inform. 2023;30:e100602.
  30. U.S. FDA / EMA. Guiding principles of good AI practice in drug development. Joint guiding principles (applied by analogy), Jan 2026.
  31. NIST. Artificial intelligence risk management framework: generative AI profile (NIST AI 600-1). U.S. Dept of Commerce, July 2024.
  32. NICE. NICE real-world evidence framework. Corporate guidance ECD9, 2026.
  33. FDA. Use of real-world evidence to support regulatory decision-making for medical devices. FDA-2023-D-4395, 18 Dec 2025.
  34. EDPB. Guidelines 07/2020 on the concepts of controller and processor in the GDPR. Adopted 07 July 2021.
  35. Castor. Operational telemetry and performance benchmarks from live production deployments of Catalyst AI-assisted clinical data abstraction and Castor PMCF platform studies, 2024 to 2026. Controlled internal evidence disclosure.

Reference note. Reference [14] (Tudur Smith et al., PLoS ONE 2012) was removed with the stale 100%-SDV statistic per Sanne Derks’ review. The numbering is otherwise held stable to avoid citation drift, renumber on a final pass if preferred.

Build changelog (not for publication, strip before publish). v3, 2026-09-15. v2 was worked up FRESH from Derk draft_v6 (full long-form incl. Appendices A to G), superseding condensed v1, with Sanne Derks’ four review comments applied (Section 5 mandate vs 6/6.1/6.2 plan; stent worked example grounded in a 6.1(b) objective with method held at 6.2(b); ISO 14155 scope note at 2.2; SDV reframed with the 2012 Tudur Smith stat and ref [14] removed). v3 applies the medical-device-reviewer pre-screen. ISO edition year DROPPED per Kevin (bare “ISO 14155”, the MDR-harmonized edition is EN ISO 14155:2020/A11:2024, not the 4th ed), F1 Annex III cite corrected 1.1(b) to 1.1, F4 dropped the ISO Section 7.3 subclause (source-data integrity / ALCOA+), F5 FDA/EMA good AI practice cast as principles applied by analogy, not requirements, F6 proposal-not-law tense fixed in the callout and conclusion, F8 “audit-ready” to ALCOA+, F10 IMDRF device-event coding added, F11 Appendix F matrix restored, F13 21 CFR Part 11 paired with EU Annex 11, Appendix C column relabelled “PMCF method clause”. F2 resolved in-house. The non-applicability justification is anchored to current Annex III Section 1.1 rather than asserted as added to Annex XIV Section 5. GDPR resolved in-house. Primary PMCF conduct on Art 6(1)(c) + 9(2)(i), secondary research on 6(1)(e)/(f) + 9(2)(j), EDPB 3/2019 noted as applied by analogy. Castor stats reconciled to the cleared claims register. Footprint set to the current cleared device figures (893 MedTech / 289 live, 298 post-market incl PMCF / 98 live). Core Catalyst stats cleared (0.8%/6.6% per Pollux 2026-09-14, $48.75 to $9.75, 6 vs 39 min, 99.8% ingestion ref [36]). Retired <0.7% removed. Gates waived by Kevin (Derk-authored, Sanne review only), qc_scan done, Gemini voice pass run, human Sanne signs off on the FINAL DESIGN FILE (Kevin, 2026-09-15).

Related Posts

To read the rest of this content, please provide a little info about yourself

EDC For Researchers, Designed By Researchers

Discover all the features offered by Castor EDC

Discover Now