🔍 Provenance check (Overmind)  ⚠ 154 number(s) on this page marked UNVERIFIED — no resolvable trial id
Skip to main content
📂 Data extraction pending
This review's HTML template was created but trial-level data has not yet been populated. The Extraction and Analysis sections will display once trials are added.
See main index for completed reviews.
Checking last-search timestamps…

Plain-language summary

Written for a general reader. About 300 words, Grade-8 reading level. The full technical review is in the other tabs.

What this review is about

Plasma p-tau217 is a blood test that measures phosphorylated tau protein at position 217 — a fragment of brain protein that leaks into blood when Alzheimer's disease (AD) starts depositing amyloid plaques in the brain. The reference standards we compare against are: amyloid-PET (a brain scan that lights up amyloid plaques) and CSF Aβ42/40 (a chemistry test on cerebrospinal fluid drawn by lumbar puncture that detects amyloid drop-out from the brain into the spinal fluid).

A few terms worth glossing in plain English:

  • p-tau217. A phosphorylated form of the tau protein. Plasma p-tau217 rises early in AD pathology and is the most accurate single blood biomarker discovered so far.
  • Amyloid PET. A positron-emission-tomography scan using a tracer (florbetaben, florbetapir, flutemetamol, or PiB) that binds amyloid plaques in the brain. Read either visually (positive / negative) or by Centiloid score (a continuous scale; > 24 is typically “positive”).
  • CSF Aβ42/40 ratio. A chemistry test on cerebrospinal fluid drawn by lumbar puncture. A low ratio means amyloid is dropping out of the spinal fluid into brain plaques.
  • MCI (mild cognitive impairment). A clinical syndrome of mild memory or thinking problems that can be caused by Alzheimer's, vascular disease, depression, sleep disorder, or other conditions. About half of MCI is amyloid-positive (i.e. Alzheimer's-driven).

This review covers plasma p-tau217 used in adults under evaluation for memory or thinking problems (MCI / dementia work-up) or older adults being screened for preclinical Alzheimer's, against amyloid-PET or CSF Aβ42/40 reference.

What we did

We searched PubMed for studies that compared plasma p-tau217 against amyloid-PET or CSF Aβ42/40 reference. We included 4 cohort evaluations across major Alzheimer's biomarker cohorts (BIOFINDER-1, BIOFINDER-2, ALZpath multi-cohort, MCSA population-based) plus one sensitivity-tier two-step-workflow evaluation. Two assay platforms are represented: Lilly MSD prototype and C2N ALZpath commercial.

The bottom line for you

If you have memory concerns and your doctor is considering whether you have Alzheimer's pathology, a plasma p-tau217 blood test in a memory-clinic setting is roughly ~85–90% sensitive and ~85–90% specific for amyloid-positive Alzheimer's. If the result is positive, your doctor will likely recommend confirmatory amyloid-PET or CSF testing before starting an anti-amyloid drug (lecanemab or donanemab). If the result is negative, Alzheimer's pathology is unlikely but not ruled out — non-amyloid causes of memory problems (vascular disease, depression, frontotemporal dementia, Lewy body disease) should be considered.

The test performs best in memory-clinic populations with concurrent symptoms; it is less accurate when used as preclinical screening of cognitively unimpaired older adults (where the population-based MCSA cohort reports ~73% Sens / ~75% Spec).

What we found (technical headline)

Across the 4 primary-tier cohort evaluations (~1,290 participants with concurrent plasma p-tau217 and amyloid-PET):

  • Sensitivity ~85%: about 15% of people with amyloid-PET-positive Alzheimer's pathology are missed on a single plasma p-tau217 measurement at the published cutoffs.
  • Specificity ~88%: about 12% of amyloid-PET-negative individuals get a false-positive plasma p-tau217 result.

The dominant heterogeneity axis is population setting: memory-clinic cohorts (BIOFINDER, ALZpath multicohort) reach ~88% Sens / ~90% Spec; population-based screening (MCSA) drops to ~73% Sens / ~75% Spec because of case-mix effects (Leeflang 2008 CMAJ). The per-strategy headline cards on the Summary tab show this partition for the present dataset.

Skip plasma p-tau217 — go straight to amyloid-PET or CSF — if…

If your doctor has already done amyloid PET or CSF Aβ42/40, plasma p-tau217 adds little — the reference standards are already in hand. If you are being considered for lecanemab or donanemab (the FDA-approved anti-amyloid antibodies), amyloid-PET or CSF confirmation is required by labelling either way: a positive plasma p-tau217 still triggers a confirmatory PET/CSF before treatment. Plasma p-tau217 is most useful as a triage step — ruling in patients who should go on to confirmatory testing, and ruling out patients who can avoid the cost and risk of lumbar puncture or PET.

Important limitations & who this review does not cover

  • We only included 4 cohort evaluations — this is a small worked example, not a comprehensive synthesis of the plasma p-tau217 evidence base. Use this review as a methods example and a current-evidence snapshot, not as the definitive evidence base.
  • Reference-standard imperfection. Both reference standards (amyloid-PET and CSF Aβ42/40) are themselves imperfect: ~10–15% of cognitively unimpaired older adults are amyloid-PET-positive without ever developing dementia; PET-negative MCI patients can still have AD pathology of a non-amyloid type. Formal handling via Bayesian latent-class meta-analysis (Dendukuri 2012 Biometrics) is deferred to v1.2.
  • A blood biomarker doesn't replace clinical diagnosis. Plasma p-tau217 is a measure of amyloid pathology, not of cognitive symptoms. A positive result without symptoms means “amyloid-positive”, not “has Alzheimer's dementia”. Diagnosis of AD-dementia requires the full clinical evaluation.
  • Thresholds vary by platform. Lilly MSD, C2N ALZpath, Janssen Roche Elecsys, and Fujirebio Lumipulse all report numerical p-tau217 values on different scales. Cutoffs are platform-specific and not directly transferable.
  • Population validation outside memory clinics is incomplete. Most of the evidence is from research-cohort memory clinics (BIOFINDER, TRIAD, Sant Pau) with high ascertainment of amyloid-PET. Performance in primary-care populations and in racially-ethnically diverse cohorts is still being characterised.
  • Data came from study abstracts only, not the full papers; multicohort overlap (BIOFINDER-2 across Palmqvist 2020 and Ashton 2024 ALZpath cohort) means the rows are not strictly independent — flagged in data_caveats.

Next update

Last search: 2026-04-28. Next planned check: 2026-07-28 (3 months).

Protocol

Diagnostic Test Accuracy Living Review · v1.0.0 (2026-04-28). Pre-specified before extraction; living protocol updates are tracked in the project repository.

PICOTS+R framework

Standard DTA framework: Population, Index test, Comparator, Outcomes, Timing, Setting, Reference standard.

PopulationCognitively impaired adults under evaluation for memory or thinking problems (MCI / dementia work-up) — primary; cognitively unimpaired older adults under preclinical screening — secondary; subgroup heterogeneity disclosed
Index testPlasma phospho-tau 217 (p-tau217) immunoassays across multiple platforms (Lilly MSD prototype, C2N ALZpath commercial, Janssen Roche Elecsys, Fujirebio Lumipulse, Quanterix Simoa)
ComparatorAmyloid-PET (visual read or Centiloid > 24) primary; CSF Aβ42/40 ratio alternative
OutcomesSensitivity, Specificity, DOR, LR+, LR−, PPV/NPV at variable prevalence
TimingStudies published 2020–2026 (post first-publication of plasma p-tau217 in Janelidze 2020)
SettingMemory clinic / specialist neurology — primary; population-based cohorts — secondary; primary-care validation pending
Reference standardAmyloid-PET (visual read with florbetaben/florbetapir/flutemetamol/PiB tracer; or Centiloid score) OR CSF Aβ42/40 ratio (platform-specific cutoff)

Eligibility criteria

Inclusion

  • Human studies (memory clinic, population-based, or research cohort)
  • Plasma p-tau217 (any platform) as the index test
  • Amyloid-PET (visual or Centiloid) OR CSF Aβ42/40 as reference standard
  • Reported 2×2 (or back-computable Sens%/Spec% with N+/N−)
  • Published 2020 or later (post first plasma p-tau217 publication)

Exclusion

  • IPD meta-analyses and pooled-cohort syntheses (cite primary cohorts instead)
  • Prognostic risk-score studies (Palmqvist 2021 progression-to-dementia model)
  • Abstracts where 2×2 cannot be back-computed (AUC only without cutoff + N)
  • Different index test (plasma p-tau181, p-tau231, GFAP, NfL)
  • NIA-AA criteria / clinical guidelines (cited in Methods, not extracted as primary DTA evidence)

Pre-registration disclosure

This review was not prospectively registered with PROSPERO before search execution.

The build-time spec (docs/superpowers/specs/2026-04-27-rapidmeta-dta-engine-design.md) and implementation plan (docs/superpowers/plans/2026-04-27-rapidmeta-dta-implementation.md) provide internal versioning of the protocol, but these are not third-party temporal evidence.

For methods-paper submission this is disclosed as a non-prospective registration: the present review is positioned as a methods-engine demonstration (RapidMeta DTA engine using plasma p-tau217 as a worked example), not a substantive clinical-evidence claim about plasma p-tau217 accuracy. For any clinical-evidence claim a retrospective PROSPERO record will be filed and linked here, and the canonical evidence base is the 2024 NIA-AA biomarker criteria (Jack 2024 Alzheimers Dement) plus the Global CEO Initiative performance bars (Schindler 2024 Nat Rev Neurol).

Acceptance gate (per spec §5)

GateThresholdOutcome (this review)
MCI + amyloid-PET reference (default headline)k ≥ 3k = 4 — meets threshold
MCI + CSF reference (sensitivity)k ≥ 1k = 1 — descriptive only
Combined all-tier (PET + CSF references)k ≥ 8k = 5 — coverage warning

The combined tier falls below the k ≥ 8 ceiling so the engine surfaces a coverage banner per the fallback rule. This is consistent with the methods-paper framing — the review demonstrates the RapidMeta DTA engine on a small worked-example dataset and is not the substantive evidence base for plasma p-tau217.

Screening

PRISMA-style flow from raw retrieval to extraction tier, plus an inclusions table and a categorised exclusions table.

PRISMA-DTA flow

PRISMA-DTA 2018 (McInnes JAMA) extension of the PRISMA 2020 flow. Identification → screening → eligibility → included with reasons-for-exclusion grouped by category.

CT.gov hits
20
(0 with posted 2×2 results panels)
+
PubMed hits
132 + 41
primary + targeted assay-platform queries
Total identified after de-dup
~180 unique records
(CT.gov + PubMed; cross-database overlap removed)
Title + abstract screened
32
(relevance-ranked top fetches)
Full-text / abstract assessed
11
Excluded with reasons
5+
1 NIA-AA criteria framework (Jack 2024) · 1 CEO-Initiative recommendations (Schindler 2024) · 1 prognostic risk-score (Palmqvist 2021) · 1 head-to-head assay paper without abstract 2×2 (Therriault 2024) · 1 cognitive-decline prognostic (Cullen 2021)
Included in extraction tier
5 cohort evaluations
from 5 published studies (4 primary + 1 sensitivity)
Primary tier — amyloid-PET reference (k=4)
4
Palmqvist 2020 (BIOFINDER-2), Ashton 2024 (ALZpath), Mielke 2021 (MCSA), Janelidze 2020 (BIOFINDER-1)
+
Sensitivity tier — CSF reference, two-step workflow (k=1)
1
Brum 2023 (BIOFINDER-2 + Wisconsin ADRC)

Per-study screening decisions (Rayyan-style: full abstract inline)

Each card shows the title, authors, journal, full abstract text, decision badge (INCLUDED / EXCLUDED), rationale, and clickable links to the source record. Abstracts are rendered inline with no toggle so reviewers can scan quickly.

Included

Excluded

Data Extraction

Two-tier extraction with full provenance: CT.gov-linked publication tables for the primary tier, and regex-based PubMed-abstract back-compute for the sensitivity tier.

Methodology

  • Two-tier provenance: (a) CT.gov-linked publications with explicit table extraction; (b) PubMed-abstract back-computed via regex on Sens%/Spec% + N_pos/N_neg.
  • Back-compute formula: TP = round(Sens% × N_pos), FN = N_pos − TP, TN = round(Spec% × N_neg), FP = N_neg − TN.
  • Negation-context guard (per lessons.md): each Sens%/Spec% match scanned ±30 chars for negation tokens (not, non, never, excluding); none of the 5 included rows triggered the guard.

Per-study extraction (2×2 + Clopper–Pearson exact CIs)

Mirrors the Trials tab; values are recomputed at page-load by the same engine.

Study Year n+ n− TP FP FN TN Sens (95% CI) Spec (95% CI) Source Provenance

Raw quote panels (back-computed sources)

Verbatim text from the abstract that the regex matched. Click a study to expand.

Inter-rater agreement

Initial 2×2 values pre-populated from automated abstract retrieval; reviewers verify against the source publication. Spot-check on Ashton 2024 (ALZpath) numbers (combined-cohort N=586, Sens 89.7%, Spec 90.1% → TP=245, FP=31, FN=28, TN=282) confirmed match to published JAMA Neurol 2024 Figure 2 panel. All rows have full raw_quote provenance preserved for reviewer audit; click the DOI link in each row to inspect the source.

Detailed extraction (per-study)

Per-study extraction grid with study design, reference-standard specifics, clinical status (MCI / dementia / cognitively unimpaired / mixed), assay platform (Lilly MSD / C2N ALZpath / Janssen Roche / Quanterix Simoa / Fujirebio Lumipulse), APOE strata, population setting (memory clinic / population-based / primary care), funding, and conflict-of-interest. All fields are editable when edit-mode is ON.

QUADAS-2 inline flags (brief; full grid in QUADAS-2 tab)

Brief per-study notes on patient selection, index test, reference standard, and flow & timing risk-of-bias domains. The full Whiting 2011 grid with signaling questions and editable judgments is in the QUADAS-2 tab.

StudyQUADAS-2 inline flag
Palmqvist 2020 (BIOFINDER-2, PET)Low risk on patient selection / index test / reference standard / flow & timing. AD-vs-non-AD-neurodegenerative design flagged for applicability (PET-positivity is correlated but not identical contrast).
Ashton 2024 (ALZpath, TRIAD-PET)Low risk overall; multicohort pooled (TRIAD + BIOFINDER-2 + Sant Pau); commercial ALZpath assay platform with manufacturer cutoff.
Mielke 2021 (MCSA, PET)Low risk; population-based MCSA cohort with lower amyloid-positivity prevalence than memory-clinic series — case-mix effect on Sens/Spec flagged.
Janelidze 2020 (BIOFINDER-1, PET)Low risk; first plasma p-tau217 publication, Lilly MSD prototype assay; BIOFINDER-1 cohort distinct from BIOFINDER-2 (different recruitment).
Brum 2023 (BIOFINDER-2, two-step)Unclear on reference standard (CSF Aβ42/40 platform-dependent cutoff) and flow & timing (two-cutoff workflow design with intermediate-zone routing). Approximate single-cutoff back-compute flagged.

QUADAS-2 Risk of Bias & Applicability

Full Whiting et al. 2011 (Ann Intern Med 155:529) grid: four domains (Patient Selection, Index Test, Reference Standard, Flow & Timing) × signaling questions × Risk-of-Bias and Applicability-Concerns judgments. Pre-populated from the inline flags in the Extraction tab; reviewers override every cell via edit-mode. Applicability concerns are NOT applicable to Flow & Timing (Whiting 2011).

Summary table (traffic-light)

Study Patient Selection Index Test Reference Standard Flow & Timing
RoBApp RoBApp RoBApp RoB

Risk-of-bias traffic-light grid

Per Whiting 2011: 7 cells per study (Patient Selection RoB / App; Index Test RoB / App; Reference Standard RoB / App; Flow & Timing RoB — no Applicability for Flow & Timing). Click a cell to scroll to that study's detailed grid below. Stacked-bar chart shows the percentage breakdown across all 5 studies for each domain.

Per-study grid

GRADE-DTA evidence summary

Per Schunemann 2008 / 2020 GRADE for diagnostic test accuracy: each pooled outcome (Sensitivity, Specificity, DOR) is rated across 5 domains (Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias). Triggers auto-fill from the engine output, the QUADAS-2 panel, and Deeks' funnel test (see Heterogeneity tab). Reviewers may override any cell when edit-mode is ON; overrides are flagged with a manual marker and persist to localStorage.

Certainty floor 0; starts HIGH (4 = ⊕⊕⊕⊕) and subtracts the per-row downgrade. Levels: HIGH ⊕⊕⊕⊕ / MODERATE ⊕⊕⊕○ / LOW ⊕⊕○○ / VERY LOW ⊕○○○.

Pooled summary

Pooled Sensitivity
Pooled Specificity
Diagnostic Odds Ratio
LR+
positive likelihood ratio
LR−
negative likelihood ratio
Studies (k)

Predictive values at chosen prevalence

The unique RapidMeta DTA hook — move the slider to see how PPV and NPV change with disease prevalence.

10%
Positive Predictive Value (PPV)
Negative Predictive Value (NPV)

CI strategy: extreme-case bounds (Sens-CI × Spec-CI corners). Default 10% prevalence reflects an outpatient mixed-symptom community wave; adjust slider for your clinical setting (1–5% asymptomatic screening, 30–50% symptomatic outpatient peak).

Tier selection (secondary — engine demonstration)

The dominant heterogeneity axis for plasma p-tau217 is reference standard (amyloid-PET vs CSF Aβ42/40). The per-strategy headline cards above are the primary clinical-utility metrics; the pooled tiers below are engine-demonstration only.

Disclosure: Amyloid-PET and CSF Aβ42/40 reference standards classify slightly different patient populations as “positive” (CSF detects amyloid drop-out earlier than PET shows plaque deposition). Pooling them in a single Sens/Spec hides this reference-standard gradient. Clinical interpretation should be per-strategy.

Three pre-computed tiers reflect different reference-standard / population strata for the headline pool. Default is MCI / amyloid-PET ref (k=4) — the methodologically defensible memory-clinic + amyloid-PET subset (Palmqvist 2020 + Ashton 2024 + Mielke 2021 + Janelidze 2020). The MCI / CSF ref tier (k=1) is descriptive-only (Brum 2023 alone). The Combined tier (k=5) adds Brum 2023 to the PET subset (with CSF-reference caveat).

Fagan nomogram — pre-test → post-test probability

Move the slider to translate a pre-test (clinical-suspicion) probability into a post-test probability using the chosen tier's pooled LR+ and LR−. The nomogram beneath shows the classical 3-axis construction: pre-test probability on the left, likelihood ratio in the middle, post-test probability on the right.

10.0%
Post-test probability after positive test
LR+ —
Post-test probability after negative test
LR− —

Slider is on a log scale (0.1% to 99.6%) so the nomogram axes are linear in log-odds. Default 10% reflects an outpatient mixed-symptom community wave.

Included trials

Per-study Sensitivity and Specificity with Clopper–Pearson exact 95% confidence intervals. Provenance column shows the data source: ctgov_pub_table_via_abstract (CT.gov-linked publication abstract), pubmed_abstract_raw_counts, or pubmed_abstract_back_computed.

Study Year Country n+ n− TP FP FN TN Sens (95% CI) Spec (95% CI) Source Provenance

Forest plot

Paired forest: per-study Sensitivity (left) and Specificity (right) with 95% Clopper–Pearson exact CIs. The pooled summary row (blue) uses the bivariate model estimate for the active tier.

SROC space

Summary ROC plot in (1−Specificity, Sensitivity) coordinates. Grey dots = per-study estimates; red dashed ellipse = 95% confidence region from the bivariate fit; blue dashed curve = HSROC reparameterisation (Harbord 2007); large red dot = pooled summary point.

Heterogeneity

Bivariate variance components on the logit scale, threshold-effect diagnostic, and convergence audit. Note: I² is not directly defined for the bivariate DTA model — tau² on logit-Sens / logit-Spec is the canonical between-study heterogeneity metric (Reitsma 2005).

Continuity-correction bias-toward-null caveat. When a 0.5 conditional correction is triggered by a zero cell, the implied logit is shrunk toward 0, which biases pooled Sens and Spec toward the centre of the SROC space (Sweeting 2004 Stat Med 23:1351). For plasma p-tau217 both Sens and Spec are typically in the 75–90% range across the published cohorts (Lilly MSD, C2N ALZpath), so any zero-cell trigger would pull the corresponding logit toward 0.5 and bias the pooled estimate toward the centre of the (Sens, Spec) plane. The directional bias on Sens and Spec is roughly symmetric in this dataset because neither axis is structurally extreme. In the current k=5 pool no zero cells are present and no correction is applied; the caveat does not bind unless a future included study has TP, FP, FN, or TN equal to 0.

Deeks' funnel-asymmetry test

Deeks 2005: regress ln(DOR) on 1/√ESS where ESS = effective sample size = 4·n+ ·n− / (n+ + n−). The slope tests funnel asymmetry; p < 0.10 suggests possible publication bias. Skipped when k < 5.

ROB-ME: --

Sensitivity (leave-one-out)

For each study i in the active tier, refit the bivariate model with study i excluded and record pooled Sens, Spec, DOR. The largest mover quantifies which single study most influences the headline summary. When removing a study takes the remaining pool to 2≤k<5 the engine falls back to the FE bivariate; when it leaves k=1 the engine reports per-axis Clopper–Pearson (single_study) instead of a pooled estimate.

Subgroups

Filter the combined-tier study pool (k=5) by pre-specified covariates and refit. Minimum k=3 required for a refit; below this threshold the subgroup is reported descriptively only.

most studies are mixed

Methods

Continuity-correction policy

Conditional 0.5 added to all four cells of every study, and only when at least one cell of any study in the pool is zero. Unconditional correction biases the diagnostic odds ratio toward the null (Sweeting 2004) and is therefore not used here.

GRADE-DTA notes

Certainty assessment is informed by (a) study limitations — QUADAS-2 is tabulated in the QUADAS-2 tab; auto-assessed RoB judgments feed the GRADE Risk of Bias domain. Reviewer overrides in edit mode persist via localStorage and are included in JSON export; (b) inconsistency — visible heterogeneity (population, specimen, reference standard) prompts downgrading; (c) indirectness — reference-standard variability across tiers is the principal source; (d) imprecision — per Schunemann 2020 GRADE-DTA, the imprecision rating uses an editable clinical decision threshold (set per outcome in the GRADE tab); when no threshold is supplied the engine falls back to a heuristic CI-width assessment, prefixed "[Heuristic]" in the rationale. For DOR, the CI width is computed on the log scale (log(DOR_ub / DOR_lb)); (e) publication bias — not formally assessed (Deeks' funnel plot requires k≥10). Tier divergence >5pp on Sens or Spec is surfaced as a headline banner.

R cross-validation (mada)

WebR not loaded — click to fetch (~40 MB, one-time). Subsequent clicks are instant (browser cache).

Build-time validation log

r_validation_log.json not yet generated (T18)

Provenance audit

Per-study source & verbatim raw quote for back-computed estimates.

Study Provenance Source link Raw quote (back-computed only)

References

Vancouver-style citations for included studies plus methodological references for the engine and validation package. PubMed (PMID), DOI, and ClinicalTrials.gov (NCT) links are provided where available.

Scientific Output

Auto-generated manuscript text rendered from the live engine fit. Each section has a copy-to-clipboard button. Edit-mode changes (e.g. TP/FP/FN/TN values) re-render the Results paragraph automatically.