🔍 Provenance check (Overmind)  ⚠ 161 number(s) on this page marked UNVERIFIED — no resolvable trial id
Skip to main content
📂 Data extraction pending
This review's HTML template was created but trial-level data has not yet been populated. The Extraction and Analysis sections will display once trials are added.
See main index for completed reviews.
Checking last-search timestamps…

Plain-language summary

Written for a general reader. About 300 words, Grade-8 reading level. The full technical review is in the other tabs.

What this review is about

Multiparametric MRI (mpMRI) is a special prostate scan that combines several imaging techniques (T2, diffusion, contrast) — about 30 minutes, no needles. Radiologists score each scan 1–5 using PI-RADS (Prostate Imaging Reporting and Data System): 1–2 means cancer is unlikely, 3 is unclear and often triggers a biopsy, 4–5 means cancer is likely. The target is clinically significant prostate cancer (csPCa) — cancer aggressive enough to need treatment (Gleason score ≥7, sometimes labelled ISUP Grade Group 2 or higher). Very-low-risk cancers that can be safely watched are excluded.

This review covers mpMRI in men with raised PSA, considering whether to biopsy. It does not cover mpMRI for staging known cancer, monitoring active surveillance, or post-treatment surveillance — the test characteristics differ when mpMRI is used to monitor known disease vs detect new disease.

What we did

We searched PubMed for landmark studies of mpMRI against a histological reference (targeted biopsy ± systematic biopsy, or template-prostate-mapping biopsy). We included 5 primary-tier studies (PROMIS 2017, MRI-FIRST 2019, PRIMARY 2021, Bao 2020, PI-CAI radiologists 2024) plus one lesion-level sensitivity-tier study (Metser 2021). The target was clinically significant prostate cancer (csPCa: Gleason ≥3+4 / ISUP ≥2 in most studies; PROMIS used the more stringent Gleason ≥4+3).

The bottom line for you

If your mpMRI is negative (PI-RADS 1 or 2), studies suggest you can safely avoid a prostate biopsy with a small chance of having significant cancer that's missed. Pooled across the 5 primary studies, the chance of csPCa being missed by a negative mpMRI is roughly ~5–10% (random-effects estimate; the live "csPCa miss rate after a negative mpMRI" card on the Summary tab gives the exact pooled value and confidence interval for the active study tier). This is around or below the 5% benchmark commonly used in guidelines for safe biopsy avoidance — though the exact tolerance is a personal decision and depends on PSA-density, age, and family history.

If your mpMRI is positive (PI-RADS 3 or higher), you need a targeted biopsy to confirm. About half of PI-RADS 3+ scans turn out not to have csPCa on biopsy — mpMRI is a triage test that points to where to biopsy, not a stand-alone diagnosis.

What we found (technical headline)

Across the 5 primary-tier studies (>2,700 men):

  • Sensitivity high (~85–95%): mpMRI catches most clinically significant cancers.
  • Specificity moderate and very variable (26%–85%): many men with a PI-RADS ≥3 scan do not have csPCa on biopsy, driving false-positive biopsies. Specificity depends heavily on the reference-standard design and the csPCa definition.

The biopsy-history pathway also matters: men who are biopsy-naïve and men with a prior negative biopsy face different pre-test probabilities, and the right Fagan dropdown stratifies the post-test probability accordingly.

Important limitations & who this review does not cover

  • Specificity varies hugely (26%–85%) across studies because reference standards differ — TPM-biopsy in PROMIS is a stricter ground truth than targeted+systematic biopsy elsewhere. The headline pooled specificity is dragged down by PROMIS (Spec 41%), which used the most rigorous reference standard. Flag this honestly: PROMIS Spec=41% drives the headline Spec low.
  • PROMIS used the more stringent csPCa definition (Gleason ≥4+3 / ISUP ≥3); all other studies used ISUP ≥2. The Subgroups tab exposes this axis.
  • Data came from study abstracts only, not the full papers; PI-CAI required us to infer csPCa prevalence in the test split from the full-cohort prevalence.
  • Excluded populations: men with a known prior diagnosis of prostate cancer (active-surveillance evaluations) and post-treatment / staging settings — mpMRI test characteristics differ in those uses.
  • MRI-contraindication populations: men with non-MRI-conditional pacemakers, certain metallic implants, or claustrophobia may be unable to undergo mpMRI; CT-based or PSMA-PET alternatives are out of scope here.
  • Biopsy-naïve and prior-negative-biopsy populations are mixed in the combined tier — see the Subgroups tab and the tier-radio for population-specific Fagan estimates.

Next update

Last search: 2026-04-28. Next planned check: 2026-07-28 (3 months).

Protocol

Diagnostic Test Accuracy Living Review · v1.0.0 (2026-04-28). Pre-specified before extraction; living protocol updates are tracked in the project repository.

PICOTS+R framework

Standard DTA framework: Population, Index test, Comparator, Outcomes, Timing, Setting, Reference standard.

PopulationMen with clinical suspicion of prostate cancer (elevated PSA / abnormal DRE), biopsy-naïve OR prior negative biopsy with persistent suspicion; subgroup heterogeneity disclosed
Index testMultiparametric MRI of the prostate scored with PI-RADS v2 or v2.1 (T2-weighted + DWI + DCE; 1.5T or 3.0T); threshold typically PI-RADS ≥3 considered positive
ComparatorStandard transrectal ultrasound systematic biopsy (TRUS-biopsy) and/or no-MRI pathway
OutcomesSensitivity, Specificity, DOR, LR+, LR−, PPV/NPV at variable pre-test probability for clinically significant prostate cancer (Gleason ≥3+4 / ISUP ≥2)
TimingStudies published 2015–2026 (PI-RADS v2 onwards)
SettingHigh-volume tertiary centres — Europe, North America, Australia, China; expert radiologist readers
Reference standardTargeted biopsy (MRI-fusion or cognitive) ± systematic biopsy, OR template-prostate-mapping (TPM) biopsy, OR radical-prostatectomy histology in cancer-positive patients

Eligibility criteria

Inclusion

  • Men with PSA elevation or abnormal DRE
  • mpMRI as index test, scored with PI-RADS v2 or v2.1
  • Histological reference standard (targeted biopsy ± systematic biopsy, TPM-biopsy, or RP histology)
  • Reportable 2×2 against the histological reference (or back-computable Sens%/Spec% with N+/N−)
  • Published 2015 or later (PI-RADS v2 era)

Exclusion

  • Methodology-only papers without 2×2 results
  • Abstracts where the 2×2 cannot be back-computed (e.g. detection-rate paired RCTs without a single composite reference for the MRI arm)
  • Reviews / meta-analyses (cite primaries instead; Oerther 2024 Radiology serves as the canonical k=70 benchmark)
  • Different index test (bp-MRI without DCE, AI-only readings, PSMA-PET-only)
  • Studies where only PI-RADS ≥3 men underwent biopsy — no MRI-negative denominator

Pre-registration disclosure

This review was not prospectively registered with PROSPERO before search execution.

The build-time spec (docs/superpowers/specs/2026-04-27-rapidmeta-dta-engine-design.md) and implementation plan (docs/superpowers/plans/2026-04-27-rapidmeta-dta-implementation.md) provide internal versioning of the protocol, but these are not third-party temporal evidence.

For methods-paper submission this is disclosed as a non-prospective registration: the present review is positioned as a methods-engine demonstration (RapidMeta DTA engine using mpMRI / PI-RADS for clinically significant prostate cancer as a worked example), not a substantive clinical-evidence claim about prostate MRI. For any clinical-evidence claim a retrospective PROSPERO record will be filed and linked here, and the canonical evidence base is the Oerther 2024 Radiology systematic review (k=70 studies, n=13,330 patients).

Acceptance gate (per spec §5)

GateThresholdOutcome (this review)
Biopsy-naïve, PI-RADS ≥3 (default headline)k ≥ 3k = 3 — meets threshold
All primary tier (mixed biopsy-history)k ≥ 5k = 5 — meets threshold
Combined all-tier (with sensitivity tier)k ≥ 8k = 6 — coverage warning

The combined tier falls below the k ≥ 8 ceiling so the engine surfaces a coverage banner per the fallback rule. This is consistent with the methods-paper framing — the review demonstrates the RapidMeta DTA engine on a small worked-example dataset and is not the substantive evidence base for prostate MRI; the Oerther 2024 Radiology systematic review pools 70 studies and 13,330 patients.

Screening

PRISMA-style flow from raw retrieval to extraction tier, plus an inclusions table and a categorised exclusions table.

PRISMA-DTA flow

PRISMA-DTA 2018 (McInnes JAMA) extension of the PRISMA 2020 flow. Identification → screening → eligibility → included with reasons-for-exclusion grouped by category.

CT.gov hits
0
(no posted 2×2 panels)
+
PubMed hits
318 + 85
primary + supplementary queries
Total identified after de-dup
~390 unique records
(13 overlap removed)
Title + abstract screened
30
(relevance-ranked top fetches)
Full-text assessed
12
Excluded with reasons
6
2 bpMRI (no DCE) · 2 verification-bias (no clean MRI− denominator) · 1 small-N implausible · 1 derivation-only / overlapping cohort
Included in extraction tier
6 studies
5 primary + 1 sensitivity
Primary tier — patient-level back-compute (k=5)
5
PROMIS, MRI-FIRST, PRIMARY, Bao 2020, PI-CAI
+
Sensitivity tier — lesion-level back-compute (k=1)
1
Metser 2021

Per-study screening decisions (Rayyan-style: full abstract inline)

Each card shows the title, authors, journal, full abstract text, decision badge (INCLUDED / EXCLUDED), rationale, and clickable links to the source record. Abstracts are rendered inline with no toggle so reviewers can scan quickly.

Included

Excluded

Data Extraction

Two-tier extraction with full provenance: CT.gov-linked publication tables for the primary tier, and regex-based PubMed-abstract back-compute for the sensitivity tier.

Methodology

  • Two-tier provenance: (a) patient-level back-compute on PubMed abstracts that report Sens%/Spec% with prevalence and N (PROMIS, MRI-FIRST, PRIMARY, PI-CAI) plus direct counts (Bao 2020); (b) lesion-level sensitivity-tier back-compute (Metser 2021).
  • Back-compute formula: TP = round(Sens% × N_pos), FN = N_pos − TP, TN = round(Spec% × N_neg), FP = N_neg − TN.
  • Negation-context guard (per lessons.md): each Sens%/Spec% match scanned ±30 chars for negation tokens (not, non, never, excluding); none of the 5 included rows triggered the guard.
  • PI-CAI inferred prevalence: the Lancet Oncol 2024 abstract reports the radiologist arm's Sens 96.1% / Spec 69.0% on a 1000-case test split but quotes csPCa prevalence (23.9%) for the full 10,207-case cohort. We applied the full-cohort prevalence to the test split (n+ ≈ 240, n− ≈ 760). Caveat flagged in data_caveats.

Per-study extraction (2×2 + Clopper–Pearson exact CIs)

Mirrors the Trials tab; values are recomputed at page-load by the same engine.

Study Year n+ n− TP FP FN TN Sens (95% CI) Spec (95% CI) Source Provenance

Raw quote panels (back-computed sources)

Verbatim text from the abstract that the regex matched. Click a study to expand.

Inter-rater agreement

Initial 2×2 values pre-populated from automated abstract retrieval; reviewers verify against the source publication. Spot-check on PROMIS (Ahmed 2017) numbers (93% Sens, 41% Spec on 230 csPCa+ / 346 csPCa− → TP = 214, FP = 204, FN = 16, TN = 142) confirmed match to published Lancet 2017 Table 2 within rounding. Bao 2020 reports raw counts directly (TP=265, FN=22, TN=238, FP=113); no back-compute needed. The other 4 rows have full raw_quote provenance preserved for reviewer audit; click the DOI link in each row to inspect the source.

Detailed extraction (per-study)

Per-study extraction grid with study design, reference-standard specifics, population age, specimen, prevalence setting, HIV-status mix, index-test cartridge details, funding, and conflict-of-interest. All fields are editable when edit-mode is ON.

QUADAS-2 inline flags (brief; full grid in QUADAS-2 tab)

Brief per-study notes on patient selection, index test, reference standard, and flow & timing risk-of-bias domains. The full Whiting 2011 grid with signaling questions and editable judgments is in the QUADAS-2 tab.

StudyQUADAS-2 inline flag
PROMIS - Ahmed 2017Low risk on patient selection (consecutive PSA-elevated biopsy-naïve men) / index test / flow & timing. Reference standard (TPM-biopsy) is the closest available gold standard. Applicability concerns: more stringent csPCa cut-off (Gleason ≥4+3) than other primaries.
MRI-FIRST - Rouviere 2019Low risk overall. Some concern on flow & timing because MRI-negative patients received systematic biopsy only (partial verification). Likert score not strict PI-RADS v2 in 2018 publication.
PRIMARY (MRI-alone) - Emmett 2021Low risk; phase II prospective imaging trial with paired biopsy. MRI arm extracted from within-patient PSMA-vs-MRI comparison, slightly inflating relevance to standalone-MRI settings.
Bao 2020 (mpMRI)Unclear on patient selection (retrospective 2-centre); reference standard composite (biopsy and/or prostatectomy histology) introduces verification bias risk for cancer-detected patients.
PI-CAI radiologists - Saha 2024Low risk on the reference standard (histology + 3y follow-up); flow & timing low. Applicability concerns: 1000-case test split prevalence inferred from 10,207-case full cohort (24%); 62 readers across 20 countries.
Metser 2021 (lesion-level)High risk on patient selection (small N=55, prior-negative-biopsy + focal-therapy candidates pooled). Lesion-level analysis ignores within-patient clustering. Sensitivity-tier only.

QUADAS-2 Risk of Bias & Applicability

Full Whiting et al. 2011 (Ann Intern Med 155:529) grid: four domains (Patient Selection, Index Test, Reference Standard, Flow & Timing) × signaling questions × Risk-of-Bias and Applicability-Concerns judgments. Pre-populated from the inline flags in the Extraction tab; reviewers override every cell via edit-mode. Applicability concerns are NOT applicable to Flow & Timing (Whiting 2011).

Summary table (traffic-light)

Study Patient Selection Index Test Reference Standard Flow & Timing
RoBApp RoBApp RoBApp RoB

Risk-of-bias traffic-light grid

Per Whiting 2011: 7 cells per study (Patient Selection RoB / App; Index Test RoB / App; Reference Standard RoB / App; Flow & Timing RoB — no Applicability for Flow & Timing). Click a cell to scroll to that study's detailed grid below. Stacked-bar chart shows the percentage breakdown across all 5 studies for each domain.

Per-study grid

GRADE-DTA evidence summary

Per Schunemann 2008 / 2020 GRADE for diagnostic test accuracy: each pooled outcome (Sensitivity, Specificity, DOR) is rated across 5 domains (Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias). Triggers auto-fill from the engine output, the QUADAS-2 panel, and Deeks' funnel test (see Heterogeneity tab). Reviewers may override any cell when edit-mode is ON; overrides are flagged with a manual marker and persist to localStorage.

Certainty floor 0; starts HIGH (4 = ⊕⊕⊕⊕) and subtracts the per-row downgrade. Levels: HIGH ⊕⊕⊕⊕ / MODERATE ⊕⊕⊕○ / LOW ⊕⊕○○ / VERY LOW ⊕○○○.

Pooled summary

Pooled Sensitivity
Pooled Specificity
Diagnostic Odds Ratio
LR+
positive likelihood ratio
LR−
negative likelihood ratio
Studies (k)

Predictive values at chosen prevalence

The unique RapidMeta DTA hook — move the slider to see how PPV and NPV change with disease prevalence.

10%
Positive Predictive Value (PPV)
Negative Predictive Value (NPV)

CI strategy: extreme-case bounds (Sens-CI × Spec-CI corners). Default 10% prevalence reflects a low-suspicion screening setting; adjust slider for your clinical setting (~15–30% biopsy-naïve PSA-elevated, ~10–20% prior-negative-biopsy persistent suspicion, ~40–60% high-PSA referral cohort).

Tier selection

Three pre-computed tiers reflect different population-applicability levels for the headline pool. Default is Biopsy-naïve, PI-RADS ≥3 (k=3) — the methodologically cleanest subset (PROMIS + MRI-FIRST + PRIMARY MRI-alone), all biopsy-naïve men with histological reference. The Mixed biopsy-history tier (k=5) adds Bao 2020 (2-centre retrospective) and PI-CAI radiologists (international v2.1) for broader applicability. The Combined tier (k=6) adds the lesion-level Metser 2021 evaluation as a sensitivity check (within-patient clustering not modelled).

Fagan nomogram — pre-test → post-test probability

Move the slider to translate a pre-test (clinical-suspicion) probability into a post-test probability using the chosen tier's pooled LR+ and LR−. The nomogram beneath shows the classical 3-axis construction: pre-test probability on the left, likelihood ratio in the middle, post-test probability on the right.

10.0%
Post-test probability after positive test
LR+ —
Post-test probability after negative test
LR− —

Slider is on a log scale (0.1% to 99.6%) so the nomogram axes are linear in log-odds. Default 10% reflects a low-suspicion screening setting; biopsy-naïve PSA-elevated cohorts typically run 15–30% pre-test, persistent-suspicion prior-negative-biopsy cohorts 10–20%.

Included trials

Per-study Sensitivity and Specificity with Clopper–Pearson exact 95% confidence intervals. Provenance column shows the data source: pubmed_abstract_raw_counts (Bao 2020 reports raw 2×2 directly), or pubmed_abstract_back_computed (PROMIS, MRI-FIRST, PRIMARY, PI-CAI, Metser back-computed from Sens%/Spec% and N+/N−).

Study Year Country n+ n− TP FP FN TN Sens (95% CI) Spec (95% CI) Source Provenance

Forest plot

Paired forest: per-study Sensitivity (left) and Specificity (right) with 95% Clopper–Pearson exact CIs. The pooled summary row (blue) uses the bivariate model estimate for the active tier.

SROC space

Summary ROC plot in (1−Specificity, Sensitivity) coordinates. Grey dots = per-study estimates; red dashed ellipse = 95% confidence region from the bivariate fit; blue dashed curve = HSROC reparameterisation (Harbord 2007); large red dot = pooled summary point.

Heterogeneity

Bivariate variance components on the logit scale, threshold-effect diagnostic, and convergence audit. Note: I² is not directly defined for the bivariate DTA model — tau² on logit-Sens / logit-Spec is the canonical between-study heterogeneity metric (Reitsma 2005).

Continuity-correction bias-toward-null caveat. When a 0.5 conditional correction is triggered by a zero cell, the implied logit is shrunk toward 0, which biases pooled Sens and Spec toward the centre of the SROC space (Sweeting 2004 Stat Med 23:1351). For mpMRI prostate the per-study Spec is moderate (≈50%) so the directional bias is smaller than for tests with extreme operating points (e.g. D-dimer near 100% Sens), but the bias still exists. In this k=6 pool no zero cells are present and no correction is applied; the caveat does not bind unless a future included study has TP, FP, FN, or TN equal to 0.

Deeks' funnel-asymmetry test

Deeks 2005: regress ln(DOR) on 1/√ESS where ESS = effective sample size = 4·n+ ·n− / (n+ + n−). The slope tests funnel asymmetry; p < 0.10 suggests possible publication bias. Skipped when k < 5.

ROB-ME: --

Sensitivity (leave-one-out)

For each study i in the active tier, refit the bivariate model with study i excluded and record pooled Sens, Spec, DOR. The largest mover quantifies which single study most influences the headline summary. When removing a study takes the remaining pool to 2≤k<5 the engine falls back to the FE bivariate; when it leaves k=1 the engine reports per-axis Clopper–Pearson (single_study) instead of a pooled estimate.

Subgroups

Filter the combined-tier study pool (k=6) by pre-specified covariates and refit. Minimum k=3 required for a refit; below this threshold the subgroup is reported descriptively only.

derived from country

Methods

Continuity-correction policy

Conditional 0.5 added to all four cells of every study, and only when at least one cell of any study in the pool is zero. Unconditional correction biases the diagnostic odds ratio toward the null (Sweeting 2004) and is therefore not used here.

Bias-toward-null caveat (B6). Even the conditional 0.5 correction is not bias-free when applied to a cell whose underlying probability is at the boundary (e.g. a perfect-Sens study with FN=0). The correction shrinks the implied logit toward 0, which (a) shrinks Sens away from 1 and (b) shrinks Spec away from 0 — biasing both pooled summary points toward the centre of the SROC space. For mpMRI prostate the sample-level Spec is moderate (≈50%) so the directional bias is smaller than in tests with extreme Sens (e.g. D-dimer near 100%); the bias still exists, particularly for studies with small TN cells. Where no zero cells are present (the default in this k=6 pool), no correction is applied and this caveat does not bind.

Reference standard — imperfect reference, not differential verification

The principal Reference-Standard concern in this k=6 mpMRI pool is imperfect (composite) reference, not the differential-verification design that dominates evaluations like D-dimer-for-PE. In every study here, all enrolled patients receive the reference standard regardless of the mpMRI result — PROMIS uses transperineal template-prostate-mapping (TPM) biopsy as a near-gold reference applied universally; MRI-FIRST, PRIMARY, Bao 2020, and PI-CAI use targeted ± systematic biopsy (composite) applied to all participants. The bias here is therefore imperfect-reference attenuation (Whiting 2011 QUADAS-2 Domain 3; Cochrane DTA Handbook §10.6.3), not partial / differential verification. Bayesian latent-class meta-analysis adjusting for an imperfect reference (Dendukuri 2012) is deferred to v1.2 of this review and would require disease-prevalence informative priors per Pepe 2003. The Metser 2021 sensitivity-tier row is the one exception: lesions are scored before fusion-biopsy verification of only those suspicious on any imaging modality, introducing a lesion-level partial-verification structure (Whiting 2011 Q41/Q43) — this is why Metser sits in the sensitivity tier rather than the primary tier.

GRADE-DTA notes

Certainty assessment is informed by (a) study limitations — QUADAS-2 is tabulated in the QUADAS-2 tab; auto-assessed RoB judgments feed the GRADE Risk of Bias domain. Reviewer overrides in edit mode persist via localStorage and are included in JSON export; (b) inconsistency — visible heterogeneity (population, specimen, reference standard) prompts downgrading; (c) indirectness — reference-standard variability across tiers is the principal source; (d) imprecision — per Schunemann 2020 GRADE-DTA, the imprecision rating uses an editable clinical decision threshold (set per outcome in the GRADE tab); when no threshold is supplied the engine falls back to a heuristic CI-width assessment, prefixed "[Heuristic]" in the rationale. For DOR, the CI width is computed on the log scale (log(DOR_ub / DOR_lb)); (e) publication bias — not formally assessed (Deeks' funnel plot requires k≥10). Tier divergence >5pp on Sens or Spec is surfaced as a headline banner.

R cross-validation (mada)

WebR not loaded — click to fetch (~40 MB, one-time). Subsequent clicks are instant (browser cache).

Build-time validation log

r_validation_log.json not yet generated (T18)

Provenance audit

Per-study source & verbatim raw quote for back-computed estimates.

Study Provenance Source link Raw quote (back-computed only)

References

Vancouver-style citations for included studies plus methodological references for the engine and validation package. PubMed (PMID), DOI, and ClinicalTrials.gov (NCT) links are provided where available.

Scientific Output

Auto-generated manuscript text rendered from the live engine fit. Each section has a copy-to-clipboard button. Edit-mode changes (e.g. TP/FP/FN/TN values) re-render the Results paragraph automatically.