
MPI strategy shifts based on population scale. Understanding the transitions prevents architectural surprises as your patient volume grows.
Under 500k patients
Deterministic matching on identifiers + name + DOB catches 92-95% of true matches. Simple embedded MPI in the FHIR server sufficient. Manual review queue is manageable.
500k to 5M patients
Probabilistic matching becomes necessary. False-positive rate on deterministic climbs. Dedicated MPI service (Verato, NextGate, or Aidbox MDM) worth the investment.
5M+ patients
Empirical weight calibration matters — off-the-shelf weights don't reflect regional distributions. HIM review team required for merge queue processing. Nightly duplicate detection, weekly reconciliation cycles.
Key operational metrics
| Metric | Under 500k | 500k-5M | 5M+ |
|---|---|---|---|
| Auto-merge accuracy | >95% | >97% | >98% |
| Manual review queue depth | <50 | 50-500 | 500-5000 |
| Weekly new duplicates | <20 | 20-200 | 200-2000 |
| Nightly detection runtime | <5 min | 5-30 min | 30 min-2h |
Data model implications
Patient.link supports merges with replaced-by and replaces types. Downstream FHIR resources that reference merged Patients must follow link chains.
Vendor selection
1. Verato — enterprise, automated weight calibration. 2. NextGate — enterprise, manual calibration. 3. Aidbox MDM — bundled with Aidbox stack. 4. MITRE FRIL — open source, minimal tooling.
Common MPI mistakes at scale
1. Skipping empirical weight calibration. 2. Auto-merge threshold too aggressive. 3. No death registry reconciliation. 4. Weekly duplicate detection insufficient. 5. Manual merge queue without HIM staffing.
MPI is a five-year operational commitment. Match strategy to current and projected scale.
