Plain-language summary
Written for a general reader. About 300 words, Grade-8 reading level. The full technical review is in the other tabs.
What this review is about
Plasma p-tau217 is a blood test that measures phosphorylated tau protein at position 217 — a fragment of brain protein that leaks into blood when Alzheimer's disease (AD) starts depositing amyloid plaques in the brain. The reference standards we compare against are: amyloid-PET (a brain scan that lights up amyloid plaques) and CSF Aβ42/40 (a chemistry test on cerebrospinal fluid drawn by lumbar puncture that detects amyloid drop-out from the brain into the spinal fluid).
A few terms worth glossing in plain English:
- p-tau217. A phosphorylated form of the tau protein. Plasma p-tau217 rises early in AD pathology and is the most accurate single blood biomarker discovered so far.
- Amyloid PET. A positron-emission-tomography scan using a tracer (florbetaben, florbetapir, flutemetamol, or PiB) that binds amyloid plaques in the brain. Read either visually (positive / negative) or by Centiloid score (a continuous scale; > 24 is typically “positive”).
- CSF Aβ42/40 ratio. A chemistry test on cerebrospinal fluid drawn by lumbar puncture. A low ratio means amyloid is dropping out of the spinal fluid into brain plaques.
- MCI (mild cognitive impairment). A clinical syndrome of mild memory or thinking problems that can be caused by Alzheimer's, vascular disease, depression, sleep disorder, or other conditions. About half of MCI is amyloid-positive (i.e. Alzheimer's-driven).
This review covers plasma p-tau217 used in adults under evaluation for memory or thinking problems (MCI / dementia work-up) or older adults being screened for preclinical Alzheimer's, against amyloid-PET or CSF Aβ42/40 reference.
What we did
We searched PubMed for studies that compared plasma p-tau217 against amyloid-PET or CSF Aβ42/40 reference. We included 4 cohort evaluations across major Alzheimer's biomarker cohorts (BIOFINDER-1, BIOFINDER-2, ALZpath multi-cohort, MCSA population-based) plus one sensitivity-tier two-step-workflow evaluation. Two assay platforms are represented: Lilly MSD prototype and C2N ALZpath commercial.
The bottom line for you
If you have memory concerns and your doctor is considering whether you have Alzheimer's pathology, a plasma p-tau217 blood test in a memory-clinic setting is roughly ~85–90% sensitive and ~85–90% specific for amyloid-positive Alzheimer's. If the result is positive, your doctor will likely recommend confirmatory amyloid-PET or CSF testing before starting an anti-amyloid drug (lecanemab or donanemab). If the result is negative, Alzheimer's pathology is unlikely but not ruled out — non-amyloid causes of memory problems (vascular disease, depression, frontotemporal dementia, Lewy body disease) should be considered.
The test performs best in memory-clinic populations with concurrent symptoms; it is less accurate when used as preclinical screening of cognitively unimpaired older adults (where the population-based MCSA cohort reports ~73% Sens / ~75% Spec).
What we found (technical headline)
Across the 4 primary-tier cohort evaluations (~1,290 participants with concurrent plasma p-tau217 and amyloid-PET):
- Sensitivity ~85%: about 15% of people with amyloid-PET-positive Alzheimer's pathology are missed on a single plasma p-tau217 measurement at the published cutoffs.
- Specificity ~88%: about 12% of amyloid-PET-negative individuals get a false-positive plasma p-tau217 result.
The dominant heterogeneity axis is population setting: memory-clinic cohorts (BIOFINDER, ALZpath multicohort) reach ~88% Sens / ~90% Spec; population-based screening (MCSA) drops to ~73% Sens / ~75% Spec because of case-mix effects (Leeflang 2008 CMAJ). The per-strategy headline cards on the Summary tab show this partition for the present dataset.
Skip plasma p-tau217 — go straight to amyloid-PET or CSF — if…
If your doctor has already done amyloid PET or CSF Aβ42/40, plasma p-tau217 adds little — the reference standards are already in hand. If you are being considered for lecanemab or donanemab (the FDA-approved anti-amyloid antibodies), amyloid-PET or CSF confirmation is required by labelling either way: a positive plasma p-tau217 still triggers a confirmatory PET/CSF before treatment. Plasma p-tau217 is most useful as a triage step — ruling in patients who should go on to confirmatory testing, and ruling out patients who can avoid the cost and risk of lumbar puncture or PET.
Important limitations & who this review does not cover
- We only included 4 cohort evaluations — this is a small worked example, not a comprehensive synthesis of the plasma p-tau217 evidence base. Use this review as a methods example and a current-evidence snapshot, not as the definitive evidence base.
- Reference-standard imperfection. Both reference standards (amyloid-PET and CSF Aβ42/40) are themselves imperfect: ~10–15% of cognitively unimpaired older adults are amyloid-PET-positive without ever developing dementia; PET-negative MCI patients can still have AD pathology of a non-amyloid type. Formal handling via Bayesian latent-class meta-analysis (Dendukuri 2012 Biometrics) is deferred to v1.2.
- A blood biomarker doesn't replace clinical diagnosis. Plasma p-tau217 is a measure of amyloid pathology, not of cognitive symptoms. A positive result without symptoms means “amyloid-positive”, not “has Alzheimer's dementia”. Diagnosis of AD-dementia requires the full clinical evaluation.
- Thresholds vary by platform. Lilly MSD, C2N ALZpath, Janssen Roche Elecsys, and Fujirebio Lumipulse all report numerical p-tau217 values on different scales. Cutoffs are platform-specific and not directly transferable.
- Population validation outside memory clinics is incomplete. Most of the evidence is from research-cohort memory clinics (BIOFINDER, TRIAD, Sant Pau) with high ascertainment of amyloid-PET. Performance in primary-care populations and in racially-ethnically diverse cohorts is still being characterised.
- Data came from study abstracts only, not the full papers; multicohort overlap (BIOFINDER-2 across Palmqvist 2020 and Ashton 2024 ALZpath cohort) means the rows are not strictly independent — flagged in
data_caveats.
Next update
Last search: 2026-04-28. Next planned check: 2026-07-28 (3 months).
Protocol
Diagnostic Test Accuracy Living Review · v1.0.0 (2026-04-28). Pre-specified before extraction; living protocol updates are tracked in the project repository.
PICOTS+R framework
Standard DTA framework: Population, Index test, Comparator, Outcomes, Timing, Setting, Reference standard.
| Population | Cognitively impaired adults under evaluation for memory or thinking problems (MCI / dementia work-up) — primary; cognitively unimpaired older adults under preclinical screening — secondary; subgroup heterogeneity disclosed |
| Index test | Plasma phospho-tau 217 (p-tau217) immunoassays across multiple platforms (Lilly MSD prototype, C2N ALZpath commercial, Janssen Roche Elecsys, Fujirebio Lumipulse, Quanterix Simoa) |
| Comparator | Amyloid-PET (visual read or Centiloid > 24) primary; CSF Aβ42/40 ratio alternative |
| Outcomes | Sensitivity, Specificity, DOR, LR+, LR−, PPV/NPV at variable prevalence |
| Timing | Studies published 2020–2026 (post first-publication of plasma p-tau217 in Janelidze 2020) |
| Setting | Memory clinic / specialist neurology — primary; population-based cohorts — secondary; primary-care validation pending |
| Reference standard | Amyloid-PET (visual read with florbetaben/florbetapir/flutemetamol/PiB tracer; or Centiloid score) OR CSF Aβ42/40 ratio (platform-specific cutoff) |
Eligibility criteria
Inclusion
- Human studies (memory clinic, population-based, or research cohort)
- Plasma p-tau217 (any platform) as the index test
- Amyloid-PET (visual or Centiloid) OR CSF Aβ42/40 as reference standard
- Reported 2×2 (or back-computable Sens%/Spec% with N+/N−)
- Published 2020 or later (post first plasma p-tau217 publication)
Exclusion
- IPD meta-analyses and pooled-cohort syntheses (cite primary cohorts instead)
- Prognostic risk-score studies (Palmqvist 2021 progression-to-dementia model)
- Abstracts where 2×2 cannot be back-computed (AUC only without cutoff + N)
- Different index test (plasma p-tau181, p-tau231, GFAP, NfL)
- NIA-AA criteria / clinical guidelines (cited in Methods, not extracted as primary DTA evidence)
Pre-registration disclosure
This review was not prospectively registered with PROSPERO before search execution.
The build-time spec (docs/superpowers/specs/2026-04-27-rapidmeta-dta-engine-design.md) and implementation plan (docs/superpowers/plans/2026-04-27-rapidmeta-dta-implementation.md) provide internal versioning of the protocol, but these are not third-party temporal evidence.
For methods-paper submission this is disclosed as a non-prospective registration: the present review is positioned as a methods-engine demonstration (RapidMeta DTA engine using plasma p-tau217 as a worked example), not a substantive clinical-evidence claim about plasma p-tau217 accuracy. For any clinical-evidence claim a retrospective PROSPERO record will be filed and linked here, and the canonical evidence base is the 2024 NIA-AA biomarker criteria (Jack 2024 Alzheimers Dement) plus the Global CEO Initiative performance bars (Schindler 2024 Nat Rev Neurol).
Acceptance gate (per spec §5)
| Gate | Threshold | Outcome (this review) |
|---|---|---|
| MCI + amyloid-PET reference (default headline) | k ≥ 3 | k = 4 — meets threshold |
| MCI + CSF reference (sensitivity) | k ≥ 1 | k = 1 — descriptive only |
| Combined all-tier (PET + CSF references) | k ≥ 8 | k = 5 — coverage warning |
The combined tier falls below the k ≥ 8 ceiling so the engine surfaces a coverage banner per the fallback rule. This is consistent with the methods-paper framing — the review demonstrates the RapidMeta DTA engine on a small worked-example dataset and is not the substantive evidence base for plasma p-tau217.
Search Strategy
Search executed 2026-04-28 via automated retrieval against the ClinicalTrials.gov and PubMed APIs. Full provenance preserved in
ctgov_ptau217_ad_pack_2026-04-28.json and pubmed_ptau217_ad_abstracts_2026-04-28.json. Reviewers verify and finalize all included records.
CT.gov search
Condition: "Alzheimer disease" OR "mild cognitive impairment"
Intervention: "p-tau217" OR "phospho-tau 217" OR "ALZpath" OR "PrecivityAD" OR "Lumipulse p-tau217"
Date filter: study start ≥ 2018-01-01
Results: 0 trials with posted 2×2 results panels CT.gov returned no eligible primary-tier registrations — plasma p-tau217 DTA evidence is overwhelmingly published directly to PubMed via cohort biomarker papers (BIOFINDER, ADNI, MCSA, A4)
PubMed search
Primary query
("p-tau217" OR "phospho-tau 217" OR "phosphorylated tau 217") AND ("Alzheimer" OR "amyloid PET" OR "CSF") AND ("sensitivity" AND "specificity")
Dates: 2020–2026 · Total: 132 abstracts Top by relevance: 20 fetched
Targeted assay-platform query
("plasma p-tau217" OR "ALZpath" OR "Lilly MSD p-tau217" OR "Lumipulse plasma p-tau217") AND ("diagnostic accuracy" OR "AUC" OR "sensitivity")
Dates: 2020–2026 · Total: 41 abstracts Top fetched: 12
Search execution
| Date | 2026-04-28 |
| Searcher / role | Automated initial retrieval; reviewers verify and finalize the included record set. |
| Databases | ClinicalTrials.gov (NIH/NLM API v2) + PubMed (NCBI E-utilities via claude.ai PubMed MCP) |
| Records identified | 14 + 6 (CT.gov, no extractable panels) + 132 + 41 (PubMed) = 193 raw hits before de-duplication |
PRISMA-NMA Flow Diagram
Screening
PRISMA-style flow from raw retrieval to extraction tier, plus an inclusions table and a categorised exclusions table.
PRISMA-DTA flow
PRISMA-DTA 2018 (McInnes JAMA) extension of the PRISMA 2020 flow. Identification → screening → eligibility → included with reasons-for-exclusion grouped by category.
Per-study screening decisions (Rayyan-style: full abstract inline)
Each card shows the title, authors, journal, full abstract text, decision badge (INCLUDED / EXCLUDED), rationale, and clickable links to the source record. Abstracts are rendered inline with no toggle so reviewers can scan quickly.
Included
Excluded
Data Extraction
Two-tier extraction with full provenance: CT.gov-linked publication tables for the primary tier, and regex-based PubMed-abstract back-compute for the sensitivity tier.
Methodology
- Two-tier provenance: (a) CT.gov-linked publications with explicit table extraction; (b) PubMed-abstract back-computed via regex on Sens%/Spec% + N_pos/N_neg.
- Back-compute formula:
TP = round(Sens% × N_pos),FN = N_pos − TP,TN = round(Spec% × N_neg),FP = N_neg − TN. - Negation-context guard (per
lessons.md): each Sens%/Spec% match scanned ±30 chars for negation tokens (not, non, never, excluding); none of the 5 included rows triggered the guard.
Per-study extraction (2×2 + Clopper–Pearson exact CIs)
Mirrors the Trials tab; values are recomputed at page-load by the same engine.
| Study | Year | n+ | n− | TP | FP | FN | TN | Sens (95% CI) | Spec (95% CI) | Source | Provenance |
|---|
Raw quote panels (back-computed sources)
Verbatim text from the abstract that the regex matched. Click a study to expand.
Inter-rater agreement
raw_quote provenance preserved for reviewer audit; click the DOI link in each row to inspect the source.
Detailed extraction (per-study)
Per-study extraction grid with study design, reference-standard specifics, clinical status (MCI / dementia / cognitively unimpaired / mixed), assay platform (Lilly MSD / C2N ALZpath / Janssen Roche / Quanterix Simoa / Fujirebio Lumipulse), APOE strata, population setting (memory clinic / population-based / primary care), funding, and conflict-of-interest. All fields are editable when edit-mode is ON.
QUADAS-2 inline flags (brief; full grid in QUADAS-2 tab)
Brief per-study notes on patient selection, index test, reference standard, and flow & timing risk-of-bias domains. The full Whiting 2011 grid with signaling questions and editable judgments is in the QUADAS-2 tab.
| Study | QUADAS-2 inline flag |
|---|---|
| Palmqvist 2020 (BIOFINDER-2, PET) | Low risk on patient selection / index test / reference standard / flow & timing. AD-vs-non-AD-neurodegenerative design flagged for applicability (PET-positivity is correlated but not identical contrast). |
| Ashton 2024 (ALZpath, TRIAD-PET) | Low risk overall; multicohort pooled (TRIAD + BIOFINDER-2 + Sant Pau); commercial ALZpath assay platform with manufacturer cutoff. |
| Mielke 2021 (MCSA, PET) | Low risk; population-based MCSA cohort with lower amyloid-positivity prevalence than memory-clinic series — case-mix effect on Sens/Spec flagged. |
| Janelidze 2020 (BIOFINDER-1, PET) | Low risk; first plasma p-tau217 publication, Lilly MSD prototype assay; BIOFINDER-1 cohort distinct from BIOFINDER-2 (different recruitment). |
| Brum 2023 (BIOFINDER-2, two-step) | Unclear on reference standard (CSF Aβ42/40 platform-dependent cutoff) and flow & timing (two-cutoff workflow design with intermediate-zone routing). Approximate single-cutoff back-compute flagged. |
QUADAS-2 Risk of Bias & Applicability
Full Whiting et al. 2011 (Ann Intern Med 155:529) grid: four domains (Patient Selection, Index Test, Reference Standard, Flow & Timing) × signaling questions × Risk-of-Bias and Applicability-Concerns judgments. Pre-populated from the inline flags in the Extraction tab; reviewers override every cell via edit-mode. Applicability concerns are NOT applicable to Flow & Timing (Whiting 2011).
Summary table (traffic-light)
| Study | Patient Selection | Index Test | Reference Standard | Flow & Timing | |||
|---|---|---|---|---|---|---|---|
| RoB | App | RoB | App | RoB | App | RoB | |
Risk-of-bias traffic-light grid
Per Whiting 2011: 7 cells per study (Patient Selection RoB / App; Index Test RoB / App; Reference Standard RoB / App; Flow & Timing RoB — no Applicability for Flow & Timing). Click a cell to scroll to that study's detailed grid below. Stacked-bar chart shows the percentage breakdown across all 5 studies for each domain.
Per-study grid
GRADE-DTA evidence summary
Per Schunemann 2008 / 2020 GRADE for diagnostic test accuracy: each pooled outcome (Sensitivity, Specificity, DOR) is rated across 5 domains (Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias). Triggers auto-fill from the engine output, the QUADAS-2 panel, and Deeks' funnel test (see Heterogeneity tab). Reviewers may override any cell when edit-mode is ON; overrides are flagged with a manual marker and persist to localStorage.
Certainty floor 0; starts HIGH (4 = ⊕⊕⊕⊕) and subtracts the per-row downgrade. Levels: HIGH ⊕⊕⊕⊕ / MODERATE ⊕⊕⊕○ / LOW ⊕⊕○○ / VERY LOW ⊕○○○.
Pooled summary
Predictive values at chosen prevalence
The unique RapidMeta DTA hook — move the slider to see how PPV and NPV change with disease prevalence.
CI strategy: extreme-case bounds (Sens-CI × Spec-CI corners). Default 10% prevalence reflects an outpatient mixed-symptom community wave; adjust slider for your clinical setting (1–5% asymptomatic screening, 30–50% symptomatic outpatient peak).
Tier selection (secondary — engine demonstration)
The dominant heterogeneity axis for plasma p-tau217 is reference standard (amyloid-PET vs CSF Aβ42/40). The per-strategy headline cards above are the primary clinical-utility metrics; the pooled tiers below are engine-demonstration only.
Disclosure: Amyloid-PET and CSF Aβ42/40 reference standards classify slightly different patient populations as “positive” (CSF detects amyloid drop-out earlier than PET shows plaque deposition). Pooling them in a single Sens/Spec hides this reference-standard gradient. Clinical interpretation should be per-strategy.
Three pre-computed tiers reflect different reference-standard / population strata for the headline pool. Default is MCI / amyloid-PET ref (k=4) — the methodologically defensible memory-clinic + amyloid-PET subset (Palmqvist 2020 + Ashton 2024 + Mielke 2021 + Janelidze 2020). The MCI / CSF ref tier (k=1) is descriptive-only (Brum 2023 alone). The Combined tier (k=5) adds Brum 2023 to the PET subset (with CSF-reference caveat).
Fagan nomogram — pre-test → post-test probability
Move the slider to translate a pre-test (clinical-suspicion) probability into a post-test probability using the chosen tier's pooled LR+ and LR−. The nomogram beneath shows the classical 3-axis construction: pre-test probability on the left, likelihood ratio in the middle, post-test probability on the right.
Slider is on a log scale (0.1% to 99.6%) so the nomogram axes are linear in log-odds. Default 10% reflects an outpatient mixed-symptom community wave.
Included trials
Per-study Sensitivity and Specificity with Clopper–Pearson exact 95% confidence intervals. Provenance column shows the data source: ctgov_pub_table_via_abstract (CT.gov-linked publication abstract), pubmed_abstract_raw_counts, or pubmed_abstract_back_computed.
| Study | Year | Country | n+ | n− | TP | FP | FN | TN | Sens (95% CI) | Spec (95% CI) | Source | Provenance |
|---|
Forest plot
Paired forest: per-study Sensitivity (left) and Specificity (right) with 95% Clopper–Pearson exact CIs. The pooled summary row (blue) uses the bivariate model estimate for the active tier.
SROC space
Summary ROC plot in (1−Specificity, Sensitivity) coordinates. Grey dots = per-study estimates; red dashed ellipse = 95% confidence region from the bivariate fit; blue dashed curve = HSROC reparameterisation (Harbord 2007); large red dot = pooled summary point.
Heterogeneity
Bivariate variance components on the logit scale, threshold-effect diagnostic, and convergence audit. Note: I² is not directly defined for the bivariate DTA model — tau² on logit-Sens / logit-Spec is the canonical between-study heterogeneity metric (Reitsma 2005).
Continuity-correction bias-toward-null caveat. When a 0.5 conditional correction is triggered by a zero cell, the implied logit is shrunk toward 0, which biases pooled Sens and Spec toward the centre of the SROC space (Sweeting 2004 Stat Med 23:1351). For plasma p-tau217 both Sens and Spec are typically in the 75–90% range across the published cohorts (Lilly MSD, C2N ALZpath), so any zero-cell trigger would pull the corresponding logit toward 0.5 and bias the pooled estimate toward the centre of the (Sens, Spec) plane. The directional bias on Sens and Spec is roughly symmetric in this dataset because neither axis is structurally extreme. In the current k=5 pool no zero cells are present and no correction is applied; the caveat does not bind unless a future included study has TP, FP, FN, or TN equal to 0.
Deeks' funnel-asymmetry test
Deeks 2005: regress ln(DOR) on 1/√ESS where ESS = effective sample size = 4·n+ ·n− / (n+ + n−). The slope tests funnel asymmetry; p < 0.10 suggests possible publication bias. Skipped when k < 5.
Sensitivity (leave-one-out)
For each study i in the active tier, refit the bivariate model with study i excluded and record pooled Sens, Spec, DOR. The largest mover quantifies which single study most influences the headline summary. When removing a study takes the remaining pool to 2≤k<5 the engine falls back to the FE bivariate; when it leaves k=1 the engine reports per-axis Clopper–Pearson (single_study) instead of a pooled estimate.
Subgroups
Filter the combined-tier study pool (k=5) by pre-specified covariates and refit. Minimum k=3 required for a refit; below this threshold the subgroup is reported descriptively only.
Methods
Continuity-correction policy
Conditional 0.5 added to all four cells of every study, and only when at least one cell of any study in the pool is zero. Unconditional correction biases the diagnostic odds ratio toward the null (Sweeting 2004) and is therefore not used here.
GRADE-DTA notes
Certainty assessment is informed by (a) study limitations — QUADAS-2 is tabulated in the QUADAS-2 tab; auto-assessed RoB judgments feed the GRADE Risk of Bias domain. Reviewer overrides in edit mode persist via localStorage and are included in JSON export; (b) inconsistency — visible heterogeneity (population, specimen, reference standard) prompts downgrading; (c) indirectness — reference-standard variability across tiers is the principal source; (d) imprecision — per Schunemann 2020 GRADE-DTA, the imprecision rating uses an editable clinical decision threshold (set per outcome in the GRADE tab); when no threshold is supplied the engine falls back to a heuristic CI-width assessment, prefixed "[Heuristic]" in the rationale. For DOR, the CI width is computed on the log scale (log(DOR_ub / DOR_lb)); (e) publication bias — not formally assessed (Deeks' funnel plot requires k≥10). Tier divergence >5pp on Sens or Spec is surfaced as a headline banner.
R cross-validation (mada)
Build-time validation log
r_validation_log.json not yet generated (T18)
Provenance audit
Per-study source & verbatim raw quote for back-computed estimates.
| Study | Provenance | Source link | Raw quote (back-computed only) |
|---|
References
Vancouver-style citations for included studies plus methodological references for the engine and validation package. PubMed (PMID), DOI, and ClinicalTrials.gov (NCT) links are provided where available.
Scientific Output
Auto-generated manuscript text rendered from the live engine fit. Each section has a copy-to-clipboard button. Edit-mode changes (e.g. TP/FP/FN/TN values) re-render the Results paragraph automatically.