Plain-language summary
Written for a general reader. About 300 words, Grade-8 reading level. The full technical review is in the other tabs.
The bottom line for you
If you are being tested for tuberculosis (TB) and the GeneXpert Ultra is positive, you almost certainly have TB — about 96% specific in adults, so very few people without TB get a false-positive. Your clinician will likely start treatment and confirm with culture in the background.
If the result is negative AND you have classic TB symptoms (cough > 2 weeks, weight loss, night sweats, fevers) — do not stop there. Repeat sampling, ask about clinical or empirical treatment, and continue follow-up. Ultra misses about 1 in 8–12 culture-positive adult cases on a single test, and more in children and people with HIV who produce less of the bacteria (paucibacillary disease). A negative test is NOT a clean bill of health if your symptoms fit.
Key terms in plain English
- GeneXpert Ultra
- An automated molecular test that detects TB DNA in sputum (or stool / gastric aspirate). Result in ~90 minutes — much faster than culture, which takes weeks.
- Culture (LJ / MGIT)
- Growing TB bacteria from a sample in special liquid (MGIT) or solid (Löwenstein–Jensen) media — the gold-standard but slow (~2–6 weeks).
- Smear-positive vs smear-negative
- AFB smear: looking at sputum under a microscope after staining for acid-fast bacilli. Smear-positive = high bacterial load; smear-negative = paucibacillary disease (harder to diagnose).
- HIV co-infection
- People living with HIV often have paucibacillary TB (lower bacterial load) and can have non-pulmonary TB; both make TB diagnosis harder.
- Sensitivity / specificity
- Sensitivity = of people who really have TB, the % the test catches. Specificity = of people who really do not have TB, the % the test correctly clears.
What this review is about
GeneXpert Ultra is a rapid molecular test for tuberculosis (TB), endorsed by the World Health Organization in 2017. It detects TB DNA in sputum (or other samples) within about 90 minutes — much faster than traditional culture, which can take weeks.
Skip-pathway: If you have signs of TB meningitis, miliary TB, severe immunocompromise, or you are very ill, your clinician may start TB treatment empirically rather than waiting for any test result — Ultra is part of TB diagnosis but not the only tool.
What we did
We searched ClinicalTrials.gov and PubMed for studies that tested GeneXpert Ultra in people who might have lung TB, between 2017 and 2026. We found 5 studies that we could include with extractable accuracy numbers.
What we found — by who is being tested
The test works very differently in different groups. Three useful slices:
- Adult sputum vs culture (gold standard, k=2 studies, ~2,100 people): Sensitivity ~89% · Specificity ~96%. About 1 in 10 culture-positive adults are missed on first test; about 4 in 100 healthy adults get a false-positive.
- Pediatric sputum or gastric aspirate (k=2 studies): Sensitivity drops considerably (children have lower bacterial loads). Specificity remains high. A negative test in a symptomatic child does NOT rule out TB — clinical follow-up is essential.
- Stool sample in children (k=1 study): Sensitivity ~59%, Specificity ~88%. Useful when sputum cannot be collected, but performance is the lowest of the three slices.
The Summary tab shows these three slices as per-strategy headline cards; the pooled headline number averaging across them is methodologically convenient but clinically misleading.
What this means for you
If your doctor uses Ultra and the result is positive, the chance you really have TB is high — but you should still have a confirmatory culture, especially if you've had TB before (false positives are more common then).
If the result is negative, that mostly rules out TB in adults without classic symptoms. With classic TB symptoms, or in children, or with HIV co-infection, a negative test does NOT rule out TB — repeat sampling and clinical follow-up are needed.
Important limitations
- Methods example, not the substantive evidence base. We only included 5 studies — much smaller than the Cochrane review of GeneXpert Ultra (Zifodya 2022 / Horne 2025, ~70 study evaluations). For clinical decisions, defer to that Cochrane review.
- Pulmonary TB only. This review covers GeneXpert Ultra in adults and children with pulmonary TB symptoms. It does NOT cover extrapulmonary TB, drug-resistance detection beyond rifampicin, or Ultra in pleural fluid / CSF / lymph node aspirate. For those, see Kohli 2021 Cochrane CD012768.
- LMIC high-burden settings dominate. Most included studies are from low- and middle-income high-burden countries (Brazil, Bangladesh, China, Uganda) — performance in low-burden high-income-country settings may differ (lower pre-test probability changes positive predictive value).
- "Trace" results excluded. The "trace" result category (very low bacterial load, only Ultra detects) was excluded from this review's primary analysis. WHO 2017 advises treating trace as positive in HIV-positive adults and children; otherwise repeat the test.
- Abstract-only data. Data came from study abstracts only, not the full papers.
- Imperfect reference standard. Three of the five studies use a composite (clinical + microbiological) reference standard rather than culture, which is itself imperfect for paucibacillary disease — this biases the estimates and is best handled by Bayesian latent-class meta-analysis (deferred to v1.2).
Next update
Last search: 2026-04-27. Next planned check: 2026-07-27 (3 months).
Protocol
Diagnostic Test Accuracy Living Review · v1.0.0 (2026-04-27). Pre-specified before extraction; living protocol updates are tracked in the project repository.
PICOTS+R framework
Standard DTA framework: Population, Index test, Comparator, Outcomes, Timing, Setting, Reference standard.
| Population | Patients with clinical suspicion of pulmonary tuberculosis (adults and children) |
| Index test | GeneXpert MTB/RIF Ultra (Cepheid; WHO-endorsed 2017) |
| Comparator | Standard-of-care diagnostic workup (where reported) |
| Outcomes | Sensitivity, Specificity, DOR, LR+, LR−, PPV/NPV at variable prevalence |
| Timing | Studies published 2017–2026 (post-WHO endorsement of Ultra) |
| Setting | Mixed — outpatient/inpatient/community-screening; LMIC and HIC |
| Reference standard | Mycobacterial culture (LJ or MGIT) preferred; composite microbiological + clinical accepted as sensitivity-tier |
Eligibility criteria
Inclusion
- Human studies
- Xpert Ultra as the index test
- Reported 2×2 (or back-computable Sens%/Spec% with N+/N−)
- Clear reference standard documented
- Published 2017 or later
Exclusion
- Cochrane reviews / meta-analyses (cite primary studies instead)
- Different test (TB-EASY, Xpert MTB/RIF, Xpert XDR, Xpert HR)
- Methodology-only papers without results
- Abstracts where 2×2 cannot be back-computed (Sens% reported but no N+/N− breakdown)
- Population subgroup not target (e.g. trace-result interpretation only)
Pre-registration disclosure
This review was not prospectively registered with PROSPERO before search execution.
The build-time spec (docs/superpowers/specs/2026-04-27-rapidmeta-dta-engine-design.md) and implementation plan (docs/superpowers/plans/2026-04-27-rapidmeta-dta-implementation.md) provide internal versioning of the protocol, but these are not third-party temporal evidence.
For methods-paper submission this is disclosed as a non-prospective registration: the present review is positioned as a methods-engine demonstration (RapidMeta DTA engine using GeneXpert Ultra as a worked example), not a substantive clinical-evidence claim about Ultra. For any clinical-evidence claim a retrospective PROSPERO record will be filed and linked here.
Acceptance gate (per spec §5)
| Gate | Threshold | Outcome (this review) |
|---|---|---|
| Sputum + culture (default headline) | k ≥ 5 | k = 2 — coverage warning |
| All sputum / respiratory | k ≥ 5 | k = 4 — coverage warning |
| Combined all-tier (disclosure only) | k ≥ 8 | k = 5 — coverage warning |
All three tiers fall below the k ≥ 8 ceiling, so the engine ships under a coverage banner per the fallback rule. This is consistent with the methods-paper framing — the review demonstrates the RapidMeta DTA engine on a small worked-example dataset and is not the substantive evidence base for Xpert Ultra (the Cochrane review of Ultra has 70+ studies).
Search Strategy
Search executed 2026-04-27 via automated retrieval against the ClinicalTrials.gov and PubMed APIs. Full provenance preserved in
ctgov_genexpert_ultra_pack_2026-04-27.json and pubmed_genexpert_ultra_abstracts_2026-04-27.json. Reviewers verify and finalize all included records.
CT.gov search
Condition: "tuberculosis" OR "pulmonary tuberculosis"
Intervention: "Xpert Ultra" OR "MTB/RIF Ultra" OR "GeneXpert Ultra"
Date filter: study start ≥ 2015-01-01 (Ultra was first WHO-endorsed in 2017)
Results: 32 trials returned 5 deep-dived 0 with 2×2 results posted to CT.gov
PubMed search
Primary query
("Xpert Ultra" OR "MTB/RIF Ultra" OR "GeneXpert Ultra") AND ("sensitivity" OR "specificity" OR "diagnostic accuracy") AND humans[Filter]
Dates: 2017–2026 · Total: 390 abstracts Top by relevance: 30 fetched
Targeted query
"Xpert Ultra" AND "sputum" AND "culture" AND "pulmonary" NOT review NOT meta-analysis
Dates: 2018–2025 · Total: 81 abstracts Top fetched: 6
Search execution
| Date | 2026-04-27 |
| Searcher / role | Automated initial retrieval; reviewers verify and finalize the included record set. |
| Databases | ClinicalTrials.gov (NIH/NLM API v2) + PubMed (NCBI E-utilities) |
| Records identified | 32 (CT.gov) + 390 + 81 (PubMed) = 503 raw hits before de-duplication |
PRISMA-NMA Flow Diagram
Screening
PRISMA-style flow from raw retrieval to extraction tier, plus an inclusions table and a categorised exclusions table.
PRISMA-DTA flow
PRISMA-DTA 2018 (McInnes JAMA) extension of the PRISMA 2020 flow. Identification → screening → eligibility → included with reasons-for-exclusion grouped by category.
Per-study screening decisions (Rayyan-style: full abstract inline)
Each card shows the title, authors, journal, full abstract text, decision badge (INCLUDED / EXCLUDED), rationale, and clickable links to the source record. Abstracts are rendered inline with no toggle so reviewers can scan quickly.
Included
Excluded
Data Extraction
Two-tier extraction with full provenance: CT.gov-linked publication tables for the primary tier, and regex-based PubMed-abstract back-compute for the sensitivity tier.
Methodology
- Two-tier provenance: (a) CT.gov-linked publications with explicit table extraction; (b) PubMed-abstract back-computed via regex on Sens%/Spec% + N_pos/N_neg.
- Back-compute formula:
TP = round(Sens% × N_pos),FN = N_pos − TP,TN = round(Spec% × N_neg),FP = N_neg − TN. - Negation-context guard (per
lessons.md): each Sens%/Spec% match scanned ±30 chars for negation tokens (not, non, never, excluding); none of the 5 included rows triggered the guard.
Per-study extraction (2×2 + Clopper–Pearson exact CIs)
Mirrors the Trials tab; values are recomputed at page-load by the same engine.
| Study | Year | n+ | n− | TP | FP | FN | TN | Sens (95% CI) | Spec (95% CI) | Source | Provenance |
|---|
Raw quote panels (back-computed sources)
Verbatim text from the abstract that the regex matched. Click a study to expand.
Inter-rater agreement
raw_quote provenance preserved for reviewer audit; click the DOI link in each row to inspect the source.
Detailed extraction (per-study)
Per-study extraction grid with study design, reference-standard specifics, population age, specimen, prevalence setting, HIV-status mix, index-test cartridge details, funding, and conflict-of-interest. All fields are editable when edit-mode is ON.
QUADAS-2 inline flags (brief; full grid in QUADAS-2 tab)
Brief per-study notes on patient selection, index test, reference standard, and flow & timing risk-of-bias domains. The full Whiting 2011 grid with signaling questions and editable judgments is in the QUADAS-2 tab.
| Study | QUADAS-2 inline flag |
|---|---|
| Dorman 2018 | Low risk on patient selection / index test / reference standard / flow & timing. |
| Andama 2021 | Low risk; HIV-mixed cohort flagged for applicability. |
| Quan 2023 | Unclear on patient selection (single-centre tertiary referral); composite reference standard flagged. |
| Pereira Battaglia 2025 | High risk on small sample (n = 41); composite reference (clinical-radiological + treatment response) flagged. |
| Kabir 2021 | Alternative-specimen (stool vs sputum) flagged for applicability; reference is bacteriologically-confirmed-on-induced-sputum (not pure culture). |
QUADAS-2 Risk of Bias & Applicability
Full Whiting et al. 2011 (Ann Intern Med 155:529) grid: four domains (Patient Selection, Index Test, Reference Standard, Flow & Timing) × signaling questions × Risk-of-Bias and Applicability-Concerns judgments. Pre-populated from the inline flags in the Extraction tab; reviewers override every cell via edit-mode. Applicability concerns are NOT applicable to Flow & Timing (Whiting 2011).
Summary table (traffic-light)
| Study | Patient Selection | Index Test | Reference Standard | Flow & Timing | |||
|---|---|---|---|---|---|---|---|
| RoB | App | RoB | App | RoB | App | RoB | |
Risk-of-bias traffic-light grid
Per Whiting 2011: 7 cells per study (Patient Selection RoB / App; Index Test RoB / App; Reference Standard RoB / App; Flow & Timing RoB — no Applicability for Flow & Timing). Click a cell to scroll to that study's detailed grid below. Stacked-bar chart shows the percentage breakdown across all 5 studies for each domain.
Per-study grid
GRADE-DTA evidence summary
Per Schunemann 2008 / 2020 GRADE for diagnostic test accuracy: each pooled outcome (Sensitivity, Specificity, DOR) is rated across 5 domains (Risk of bias, Inconsistency, Indirectness, Imprecision, Publication bias). Triggers auto-fill from the engine output, the QUADAS-2 panel, and Deeks' funnel test (see Heterogeneity tab). Reviewers may override any cell when edit-mode is ON; overrides are flagged with a manual marker and persist to localStorage.
Certainty floor 0; starts HIGH (4 = ⊕⊕⊕⊕) and subtracts the per-row downgrade. Levels: HIGH ⊕⊕⊕⊕ / MODERATE ⊕⊕⊕○ / LOW ⊕⊕○○ / VERY LOW ⊕○○○.
Pooled summary
Predictive values at chosen prevalence
The unique RapidMeta DTA hook — move the slider to see how PPV and NPV change with disease prevalence.
CI strategy: extreme-case bounds (Sens-CI × Spec-CI corners). Default 10% prevalence reflects outpatient LMIC presumptive PTB; adjust slider for your clinical setting.
Tier selection — engine demonstration only
Three pre-computed tiers reflect different methodological defensibility levels for the headline pool. Default is Sputum + culture (k=2) — the methodologically defensible adult sputum-vs-culture subset. The Combined tier (k=5) is provided for completeness with a disclosure banner; it pools studies that differ on population (adult vs. pediatric), specimen (sputum vs. stool), and reference standard (culture vs. composite/clinical) and so should not be used directly for clinical decisions. The per-strategy cards above are the primary clinical-utility view; the tier radio below is engine-demonstration only.
Fagan nomogram — pre-test → post-test probability
Move the slider to translate a pre-test (clinical-suspicion) probability into a post-test probability using the chosen tier's pooled LR+ and LR−. The nomogram beneath shows the classical 3-axis construction: pre-test probability on the left, likelihood ratio in the middle, post-test probability on the right.
Slider is on a log scale (0.1% to 99.6%) so the nomogram axes are linear in log-odds. Default 10% reflects outpatient LMIC presumptive PTB.
Included trials
Per-study Sensitivity and Specificity with Clopper–Pearson exact 95% confidence intervals. Provenance column shows the data source: ctgov_pub_table_via_abstract (CT.gov-linked publication abstract), pubmed_abstract_raw_counts, or pubmed_abstract_back_computed.
| Study | Year | Country | n+ | n− | TP | FP | FN | TN | Sens (95% CI) | Spec (95% CI) | Source | Provenance |
|---|
Forest plot
Paired forest: per-study Sensitivity (left) and Specificity (right) with 95% Clopper–Pearson exact CIs. The pooled summary row (blue) uses the bivariate model estimate for the active tier.
SROC space
Summary ROC plot in (1−Specificity, Sensitivity) coordinates. Grey dots = per-study estimates; red dashed ellipse = 95% confidence region from the bivariate fit; blue dashed curve = HSROC reparameterisation (Harbord 2007); large red dot = pooled summary point.
Heterogeneity
Bivariate variance components on the logit scale, threshold-effect diagnostic, and convergence audit. Note: I² is not directly defined for the bivariate DTA model — tau² on logit-Sens / logit-Spec is the canonical between-study heterogeneity metric (Reitsma 2005).
Deeks' funnel-asymmetry test
Deeks 2005: regress ln(DOR) on 1/√ESS where ESS = effective sample size = 4·n+ ·n− / (n+ + n−). The slope tests funnel asymmetry; p < 0.10 suggests possible publication bias. Skipped when k < 5.
Sensitivity (leave-one-out)
For each study i in the active tier, refit the bivariate model with study i excluded and record pooled Sens, Spec, DOR. The largest mover quantifies which single study most influences the headline summary. When removing a study takes the remaining pool to 2≤k<5 the engine falls back to the FE bivariate; when it leaves k=1 the engine reports per-axis Clopper–Pearson (single_study) instead of a pooled estimate.
Subgroups
Filter the combined-tier study pool (k=5) by pre-specified covariates and refit. Minimum k=3 required for a refit; below this threshold the subgroup is reported descriptively only.
Methods
Continuity-correction policy
Conditional 0.5 added to all four cells of every study, and only when at least one cell of any study in the pool is zero. Unconditional correction biases the diagnostic odds ratio toward the null (Sweeting 2004) and is therefore not used here.
GRADE-DTA notes
Certainty assessment is informed by (a) study limitations — QUADAS-2 is tabulated in the QUADAS-2 tab; auto-assessed RoB judgments feed the GRADE Risk of Bias domain. Reviewer overrides in edit mode persist via localStorage and are included in JSON export; (b) inconsistency — visible heterogeneity (population, specimen, reference standard) prompts downgrading; (c) indirectness — reference-standard variability across tiers is the principal source; (d) imprecision — per Schunemann 2020 GRADE-DTA, the imprecision rating uses an editable clinical decision threshold (set per outcome in the GRADE tab); when no threshold is supplied the engine falls back to a heuristic CI-width assessment, prefixed "[Heuristic]" in the rationale. For DOR, the CI width is computed on the log scale (log(DOR_ub / DOR_lb)); (e) publication bias — not formally assessed (Deeks' funnel plot requires k≥10). Tier divergence >5pp on Sens or Spec is surfaced as a headline banner.
R cross-validation (mada)
Build-time validation log
r_validation_log_genexpert_ultra_tb.json not yet generated (T18)
Provenance audit
Per-study source & verbatim raw quote for back-computed estimates.
| Study | Provenance | Source link | Raw quote (back-computed only) |
|---|
References
Vancouver-style citations for included studies plus methodological references for the engine and validation package. PubMed (PMID), DOI, and ClinicalTrials.gov (NCT) links are provided where available.
Scientific Output
Auto-generated manuscript text rendered from the live engine fit. Each section has a copy-to-clipboard button. Edit-mode changes (e.g. TP/FP/FN/TN values) re-render the Results paragraph automatically.