🔍 Provenance check (Overmind)  ⚠ 61 number(s) on this page marked UNVERIFIED — no resolvable trial id
AUTOMATED OUTPUT β€” NOT a validated meta-analysis. This page pools fewer than two trials with poolable data and has not been externally benchmarked or provenance-verified. See the validated portfolio at the index.
Skip to main content
πŸ“‚ Data extraction pending
This review's HTML template was created but trial-level data has not yet been populated. The Extraction and Analysis sections will display once trials are added.
See main index for completed reviews.

Pooled C-statistic

Across … external-validation cohorts the pooled 5-year C-statistic of the 4-variable KFRE on the logit scale was back-transformed to give the value below. Paule-Mandel REML τ² is reported on the logit-C scale; the HKSJ multiplier reflects the Cochrane v6.5 floor max(1, Q/(kβˆ’1)) with tkβˆ’1 critical value.

Heterogeneity

PROBAST overall risk-of-bias rollup

Tangri 2016 IPDMA consistency check

…

Subgroup contrast (derivation vs external)

All six cohorts in this subset are external validations of the KFRE β€” the derivation cohort (Tangri et al. 2011, Canadian community-CKD) is not re-included here. Subgroup tabulation is shown to demonstrate the dev-vs-external split capability of the engine.

BucketkPooled C95% CIτ² (logit)IΒ²

Included cohorts

Six external-validation cohorts of the 4-variable KFRE at the 5-year horizon. Each row links to PubMed via PMID; the DOI is also given. C-statistic and 95% CI are quoted verbatim from each publication's abstract; SE(C) is derived from the CI assuming Wald symmetry (SE = (CIhigh βˆ’ CIlow) / (2 Γ— 1.96)) for input to the logit-scale pool.

Cohort & setting PMID Year Country N total Events / FU (y) C (95% CI) SE(C)

Cohort detail

Discrimination — C-statistic forest

Per-cohort C-statistic with Wald 95% CI (back-transformed from logit-C). Pooled diamond is Paule-Mandel REML + HKSJ with the Cochrane v6.5 small-sample floor. The dashed orange line is the 95% prediction interval (tk-1; the expected range for the C-statistic of a notional seventh external validation drawn from the same underlying population of validation studies).

Figure 1. Pooled C-statistic of the 4-variable Kidney Failure Risk Equation at the 5-year prediction horizon across six external validation cohorts. Blue squares are per-cohort point estimates with 95% CIs. The cyan diamond is the pooled estimate by Paule-Mandel REML τ² with HKSJ small-sample CI (Cochrane v6.5 floor max(1, Q/(kβˆ’1)); Student-tk-1 critical value). The dashed orange bracket below the diamond is the 95% prediction interval (Higgins 2009; Riley 2011), reflecting the between-cohort variability in the underlying logit-C population. Heterogeneity is high (IΒ² β‰ˆ 97%; τ² point estimate on logit-C scale β‰ˆ 0.68 with q-profile 95% CI [β‰ˆ 0.25, β‰ˆ 4.42], Viechtbauer 2007), driven by a wide spread of validation settings: primary care (C = 0.93 in Major 2019), advanced CKD (C = 0.77–0.82 in Ali 2021 and Gallego-Valcarce 2024), kidney transplant recipients (C = 0.73 in Tangri 2020 β€” a different prediction context to the derivation cohort), and population data linkage (C = 0.96 in Irish 2023). The wide HKSJ-adjusted CI is a deliberate honest reflection of this heterogeneity, not a numerical artefact. The heterogeneity-fragility flag (parity with the pairwise-pool engine) is agree here: both the pooled CI and the prediction interval lie above logit(0.5) = 0, so the conclusion of useful discrimination is not heterogeneity-fragile.

Variance derivation: SE(C) is taken from the published 95% CI assuming Wald symmetry. Delta-method conversion to the logit scale: SE(logit C) = SE(C) / [C(1βˆ’C)]. Pooling on the logit-C scale follows Snell et al. 2018 (BMC Med Res Methodol) and Debray et al. 2017 (BMJ tutorial). Back-transformation is via the inverse-logit (sigmoid) function.

Calibration findings (narrative)

The six included cohorts reported calibration narratively without machine-extractable numerical intercepts, slopes, or O/E ratios. The aggregate qualitative pattern across studies is consistent: the original (non-recalibrated, non-North American calibrated) 4-variable KFRE over-predicts kidney failure risk in lower-risk validation cohorts and under-predicts in higher-risk advanced-CKD cohorts. Both directional miscalibrations are reported in the source publications.

CohortCalibration finding (quoted or paraphrased from abstract)
Major 2019 (UK primary care) Original KFRE over-predicted in lower (<20%) risk groups; recalibration via baseline-risk adjustment improved calibration (baseline survival 0.9832 β†’ 0.9878 at 2 y; 0.9365 β†’ 0.9570 at 5 y).
Lennartz 2016 (CARE FOR HOMe, Germany) Original KFRE "well calibrated" per visual calibration plots; numerical calibration intercept/slope not provided in abstract.
Tangri 2020 (N. American transplant) Numerical calibration not reported in abstract. Subgroup analysis in patients with eGFR < 45 mL/min/1.73 mΒ² gave higher 2 y C (0.88 vs 0.81) suggesting better discrimination near event horizon.
Ali 2021 (Salford advanced CKD) "Underestimation of risk at 2 years and overestimation of risk at 5 years, especially in high-risk patients"; underestimation specifically reported in ADPKD aetiology subgroup.
Gallego-Valcarce 2024 (Spain advanced CKD) "Excellent calibration"; Brier score < 0.20; Hosmer-Lemeshow goodness-of-fit p > 0.05.
Irish 2023 (Tasmania, Australia) Calibration curves indicated systematic over-prediction; Brier score 0.01 at 5 y. Recalibration factors developed (baseline survival 0.9365 β†’ 0.9633 at 5 y).

Why calibration was not pooled

Numerical pooling of calibration intercept, calibration slope, and observed/expected ratio requires (i) the point estimate, (ii) the standard error, and (iii) the analysis scale, to be reported in the source publication. None of the six abstracts surfaced all three quantities simultaneously for any single calibration metric. Pooling the qualitative narrative would require unverifiable transformation. The engine's calibration pooling primitives are exercised in the unit-test suite (Paule-Mandel REML on raw-scale intercept, raw-scale slope, and log-O/E with back-transformation) and documented in tests/prediction_fixtures/r_metafor_baselines.json; for a real-cohort calibration pool, individual-participant data or full-text supplementary tables from the source studies are required.

PROBAST risk-of-bias

Per-cohort PROBAST judgement across the four standard domains (participants, predictors, outcome, analysis) following the Wolff et al. 2019 manual (Annals of Internal Medicine, PMID 30596875). Overall judgement: low only when every domain is rated low; high when any domain is rated high; unclear otherwise. Domain ratings are pragmatic assessments based on each cohort's published methods (sample size, multi-centre vs single-centre design, completeness of follow-up, completeness of predictor measurement).

Cohort
Participants
Predictors
Outcome
Analysis
Overall

Domain-by-domain rationale

Participants

The three large multi-centre / population-linkage cohorts (Major 2019, Tangri 2020, Irish 2023) drew participants from defined source populations with explicit inclusion criteria and minimal selection bias β€” rated low. Lennartz 2016 enrolled a tertiary nephrology referral population (selection toward advanced CKD; rated unclear). Ali 2021 and Gallego-Valcarce 2024 are single-centre advanced-CKD cohorts with small samples; Ali 2021 split into aetiology subgroups with small per-subgroup N raising bias concerns, and Gallego-Valcarce 2024 has N = 339 β€” rated high for participants.

Predictors

The 4-variable KFRE predictors (age, sex, eGFR, urine ACR) are routinely measured in all included cohorts with standardised laboratory methods. All six rated low.

Outcome

Kidney failure (initiation of kidney replacement therapy or estimated GFR < 15 mL/min/1.73 mΒ² in sustained measurements) is an objective hard endpoint with standardised ascertainment in registry-linked or EHR-linked cohorts. All six rated low.

Analysis

Large cohorts with high events-per-variable (Major 2019 with 429 events, Tangri 2020 pooled, Irish 2023 with 285 KRT events) rated low. Lennartz 2016 had moderate event count in a small sample (rated unclear). Ali 2021 (small per-aetiology subgroups) rated unclear. Gallego-Valcarce 2024 (N = 339, abstract does not specify event count) raises concern about events-per-variable and is rated high for analysis pending full-text events-per-variable verification.

Overall rollup

Three cohorts rated overall low (Major 2019, Tangri 2020, Irish 2023), two rated high (Ali 2021, Gallego-Valcarce 2024), one rated unclear (Lennartz 2016). A sensitivity analysis restricted to the three low-RoB cohorts would re-pool over a much narrower range of validation contexts (large-scale primary care, multi-centre transplant, population data linkage) and is recommended as a follow-up.

Caveat. These domain ratings are this review's interpretation of each cohort's published methods and are intended for engine demonstration, not as a definitive PROBAST assessment. A definitive PROBAST application would require two independent reviewers consulting each full-text article, with adjudication of disagreements.

Methods

Eligibility

External validations of the 4-variable Kidney Failure Risk Equation (Tangri 2011; age + sex + eGFR + urine ACR predicting kidney replacement therapy) that report a numerical C-statistic with 95% CI at the 5-year horizon. Both the original KFRE and "non-North-American-calibrated" KFRE were included; recalibrated KFRE versions are noted but not re-pooled. Cohorts in kidney-transplant recipients are included as a clinically-distinct validation context, with the caveat that the prediction target (graft loss vs native-kidney KRT) differs from the derivation cohort.

Search and selection

Six independently-published external validation cohorts were identified through anchor-search on the Tangri 2016 multinational IPDMA (JAMA, PMID 26757465) and subsequent KFRE-validation literature. This is a methods-engine demonstration on a verifiable subset of the broader KFRE-validation literature, not an exhaustive systematic review. For comprehensive coverage, the 31-cohort IPDMA (Tangri 2016) and the recent 59-cohort assessment with CKD-EPI 2021 eGFR (PMC10103205) remain authoritative.

Data extraction

For each cohort the C-statistic and its 95% CI were extracted verbatim from the publication's PubMed abstract. SE(C) was derived from the CI assuming Wald symmetry: SE = (CIupper βˆ’ CIlower) / (2 Γ— 1.96). Cohort N, country, setting, and follow-up duration were also extracted where reported. Calibration metrics (intercept, slope, O/E ratio, Brier score) were extracted where numerically reported with associated SE; absence of full numerical calibration triples meant calibration was not pooled (see Calibration tab).

Statistical pooling

The 4-variable KFRE 5-year C-statistic was pooled on the logit scale (Snell 2018; Debray 2017). For each cohort:

Random-effects pooling was via Paule-Mandel REML β€” an exact REML estimator for the single-axis case that iterates τ² to solve Ξ£ wi(yi βˆ’ ΞΌw)Β² = k βˆ’ 1 with wi = 1/(vi + τ²). Advanced-stats.md guidance: PM is preferred over DerSimonian-Laird for k < 10 because DL is downward-biased at small k. The pooled mean was the inverse-variance-weighted average at the converged τ²; SE was sqrt(1 / Ξ£ wi).

Small-sample CI used the Hartung-Knapp-Sidik-Jonkman adjustment with the Cochrane Handbook v6.5 (Nov 2024 Β§10.10.4.3) floor: HKSJ scale factor = max(1, Q / (k βˆ’ 1)), with Student-tk βˆ’ 1 critical value rather than the normal-z. This protects against CI under-coverage in the presence of heterogeneity. A 95% prediction interval (Higgins 2009; Riley 2011) was reported using tk βˆ’ 1 Γ— √(τ² + SEΞΌΒ²), defined for k β‰₯ 3.

In addition to the τ² point estimate the engine reports a 95% q-profile confidence interval for τ² (Viechtbauer 2007, Statistics in Medicine), obtained by bisection-inverting Cochran's Q against χ²(kβˆ’1) at Ξ±/2 and 1 βˆ’ Ξ±/2. This is the same algorithm used by the DTA engine's qProfileTau2CI. A wide τ² CI is informative: it confirms that the heterogeneity estimate itself is uncertain at small k, separately from the τ² magnitude. A heterogeneity-fragility flag (parity with the pairwise-pool engine's piGap) reports whether the pooled CI and prediction interval disagree about whether a clinically-meaningful reference value lies inside the interval β€” reference values are logit(0.5) = 0 for the C-statistic, 0 for calibration intercept, 1 for calibration slope, and log(1) = 0 for the O/E ratio. The flag takes three values: agree, ci_excludes_pi_includes (pooled CI claims discrimination but a future cohort could be at chance β€” heterogeneity-fragile), or ci_includes_pi_excludes (CI is null but PI is shifted).

All pooling primitives, back-transformations, and the PROBAST rollup are implemented in the open-source rapidmeta-prediction-engine-v1.js (~ 826 lines, MIT licence, no external dependencies). The engine is validated by a 64-assertion unit-test suite (tests/test_prediction_engine.mjs) including: closed-form τ² = 0 case (hand-derived against inverse-variance algebra), heterogeneous k = 3 case where PM converges to a value identical to DerSimonian-Laird (verified by hand on symmetric balanced data), the 6-cohort KFRE real-data frozen baseline, the Hanley-McNeil 1982 variance against the 1982 Radiology paper Table II hand-calculation, Student-t and chi-squared quantiles against Abramowitz & Stegun, and various edge-case stress paths.

Risk-of-bias assessment

PROBAST (Wolff et al. 2019, PMID 30596875) was applied across the four prediction-model-specific domains: participants (selection bias, source population), predictors (assessment and definition), outcome (ascertainment and blinding), and analysis (events-per-variable, handling of missing data, statistical method). Each domain was rated low / high / unclear; overall: low only if all four domains low, high if any domain high, unclear otherwise. Domain rationale is provided on the PROBAST tab.

Cross-validation against R

The engine's pooling primitives (Paule-Mandel iteration; HKSJ floor with tk βˆ’ 1; logit / inv-logit / Hanley-McNeil variance) are the same primitives metafor::rma(yi, vi, method='PM', test='knha') uses for univariate meta-analyses, and the same primitives metamisc::valmeta() uses for prediction-model meta-analyses. Three frozen baseline cases are recorded in tests/prediction_fixtures/r_metafor_baselines.json: (i) homogeneous τ² = 0 closed-form, (ii) heterogeneous k = 3 case where PM converges to the DL value 0.24 by analytic derivation, (iii) the 6-cohort KFRE real-data pool with all key intermediates frozen to engine-output precision. Bit-reproducibility against metafor on identical inputs is the planned v1.1 build-time check.

Discussion

Principal findings

Across six independently-published external validations the pooled 5-year C-statistic of the 4-variable KFRE on the logit scale was 1.95 (back-transformed C = …; 95% HKSJ CI …; 95% prediction interval …). The Tangri 2016 multinational IPDMA (31 cohorts, n = 721,357) reported a pooled 5-year C of 0.88 (95% CI 0.86–0.90); the present 6-cohort subset is consistent in central tendency (…) but produces a substantially wider HKSJ-adjusted CI because of the heterogeneity inherent in pooling validations across very different prediction contexts.

Heterogeneity

Between-cohort heterogeneity is high (τ² β‰ˆ 0.68 on the logit-C scale; IΒ² β‰ˆ 97%). The principal driver is the diversity of validation contexts. At one extreme, Major 2019 (UK primary care, 35,539 participants) and Irish 2023 (Tasmania population data linkage, 8,182 participants at 5 y) report C β‰ˆ 0.92–0.96 β€” large unselected cohorts where the KFRE's calibration on age, sex, eGFR, and ACR maps efficiently to a kidney-failure rate that is rare in the overall population. At the other extreme, the kidney-transplant cohort (Tangri 2020, C = 0.73) and the advanced-CKD subgroups (Ali 2021 in Salford and Gallego-Valcarce 2024 in Spain, C = 0.77–0.82) report substantially lower discrimination, consistent with the principle that discrimination shrinks when the case-mix narrows (Riley et al. 2019, BMJ; Vergouwe et al. 2010). In these populations most participants are at high baseline kidney-failure risk and the variation in predicted risk between participants is small, so distinguishing future events from non-events on the basis of the same predictors is harder.

This pattern is the textbook example of why prediction-model meta-analysis must pool with random effects and report the prediction interval alongside the pooled CI: a notional seventh validation cohort would be expected to land somewhere in the prediction interval (…), not within the much-narrower CI. The HKSJ adjustment with the Cochrane v6.5 small-sample floor is what produces this honest wider CI; using a normal-z CI without the small-sample correction would produce a misleadingly tight pooled estimate.

Risk-of-bias and applicability

Three of six cohorts were rated overall low risk of bias (Major 2019, Tangri 2020 transplant, Irish 2023). Two were rated high (Ali 2021 with small per-aetiology subgroups; Gallego-Valcarce 2024 with N = 339 and events-per-variable concerns). Lennartz 2016 was rated unclear (tertiary referral selection). A sensitivity analysis restricted to the three low-RoB cohorts would re-pool over primary-care, transplant, and population-linkage settings only β€” narrower applicability but with greater methodological coherence.

Implications for clinical risk stratification

The KFRE was developed and validated for native-kidney CKD populations (Tangri 2011, JAMA). Its application to kidney-transplant recipients (Tangri 2020) gives substantially lower discrimination (C = 0.73 vs β‰ˆ 0.88 in native CKD), reflecting that graft-loss biology differs from native-kidney-failure biology. For Finrenone-relevant CKD populations β€” i.e. native-kidney chronic kidney disease with type 2 diabetes, where Finrenone is indicated β€” the KFRE's primary-care discrimination is excellent (Major 2019: C = 0.92) and its advanced-CKD discrimination is good (Ali 2021: C = 0.77). The KFRE is therefore well suited for Finrenone-eligible-cohort risk stratification at the primary-care to early-CKD-clinic transition, but a recalibration step (Major 2019 baseline-risk adjustment; Irish 2023 calibration factors) is required outside North American populations.

Strengths

Limitations

Conclusion

The 4-variable KFRE shows good-to-excellent 5-year discrimination across diverse external validation contexts when pooled honestly with Paule-Mandel REML + HKSJ and a Cochrane v6.5 prediction interval. The wide pooled CI reflects real between-cohort variability driven by validation-context heterogeneity (primary-care vs advanced-CKD vs transplant), not engine pathology. For Finrenone-relevant native-kidney CKD populations the primary-care KFRE discrimination is excellent. Calibration assessment requires individual-cohort full-text data and was outside the scope of this abstract-level methods demonstration.

TRIPOD-Cluster reporting checklist

TRIPOD-Cluster (Debray et al. 2023, BMJ; PMID 36750236) extends TRIPOD to multi-cluster / multi-cohort prediction-model evaluation studies. Items below mark this review's compliance for engine-demonstration purposes; a clinical-evidence submission would require full-text extraction beyond abstract-level data.

ItemDescriptionStatus
1aIdentify as prediction-model meta-analysisβœ“ Title states KFRE external-validation meta-analysis
1bAbstract: structured summaryβœ“ See Overview tab
2Background and objectivesβœ“ Methods tab
3aSource of data per cohortβœ“ Cohorts tab (PMID, DOI, setting, N, country, follow-up)
3bInclusion / exclusion criteria across cohortsβœ“ Methods β†’ Eligibility
4Outcome definitionβœ“ Kidney replacement therapy / sustained eGFR < 15 (per cohort sources)
5Predictors and their measurementβœ“ 4-variable KFRE: age, sex, eGFR, urine ACR (Tangri 2011)
6Sample size per cohortβœ“ Cohorts table
7Missing-data handlingN/A at abstract-level extraction
8Statistical pooling methodβœ“ Methods β†’ Statistical pooling (PM REML + HKSJ)
9Risk-of-bias assessmentβœ“ PROBAST tab (per-domain rationale)
10Discrimination metric pooledβœ“ C-statistic on logit scale
11Calibration metric pooledNarrative only β€” see Calibration tab for rationale
12Heterogeneity (τ², IΒ², Q)βœ“ Overview β†’ Heterogeneity grid
13Prediction intervalβœ“ Overview + Discrimination figure caption
14Subgroup / sensitivity analysesβœ“ Derivation-vs-external subgroup (Overview); low-RoB sensitivity recommended (Discussion)
15Limitationsβœ“ Discussion β†’ Limitations
16Funding / conflictsN/A (engine demonstration; no clinical-funding declarations)
17Code / data availabilityβœ“ Engine source: rapidmeta-prediction-engine-v1.js; cohort fixture: tests/prediction_fixtures/kfre_5y_external_validations.json; R baseline: tests/prediction_fixtures/r_metafor_baselines.json

Note: this checklist is a self-assessment for engine-demonstration transparency. A clinical-evidence submission would require independent two-reviewer extraction from full-text PDFs and a definitive PROBAST adjudication.

References

Vancouver-style citations with PubMed and DOI links.

Included cohorts (n = 6)

  1. Major RW, Shepherd D, Medcalf JF, Xu G, Gray LJ, Brunskill NJ. The Kidney Failure Risk Equation for prediction of end stage renal disease in UK primary care: An external validation and clinical impact projection cohort study. PLoS Med. 2019;16(11):e1002955. PMID 31693662 Β· DOI 10.1371/journal.pmed.1002955
  2. Lennartz CS, Pickering JW, Seiler-Mußler S, Bauer L, Untersteller K, Emrich IE, Zawada AM, Radermacher J, Tangri N, Fliser D, Heine GH. External Validation of the Kidney Failure Risk Equation and Re-Calibration with Addition of Ultrasound Parameters. Clin J Am Soc Nephrol. 2016;11(4):609–615. PMID 26787778 Β· DOI 10.2215/CJN.08110715
  3. Tangri N, Ferguson TW, Bamforth RJ, Sereda G, Kumar M, Alkurd I, Tageldin T, Hingwala J, Knoll GA, Komenda P, Rigatto C. Validation of the Kidney Failure Risk Equation in Kidney Transplant Recipients. Can J Kidney Health Dis. 2020;7:2054358120922627. PMID 32549052 Β· DOI 10.1177/2054358120922627
  4. Ali I, Donne RL, Kalra PA. A validation study of the kidney failure risk equation in advanced chronic kidney disease according to disease aetiology with evaluation of discrimination, calibration and clinical utility. BMC Nephrol. 2021;22(1):194. PMID 34030639 Β· DOI 10.1186/s12882-021-02402-1
  5. Gallego-Valcarce E, Cobo-Caso MA, SΓ‘nchez-Horrillo A, Ortega-Cerrato A, PΓ©rez-MartΓ­nez J, PΓ©rez-GarcΓ­a R. External validation of the KFRE and Grams prediction models for kidney failure and death in a Spanish cohort of patients with advanced chronic kidney disease. J Nephrol. 2024;37(2):393–405. PMID 38060108 Β· DOI 10.1007/s40620-023-01819-1
  6. Irish GL, Cuthbertson L, Kitsos A, Saunder T, Clayton PA, Jose MD. The kidney failure risk equation predicts kidney failure: Validation in an Australian cohort. Nephrology (Carlton). 2023;28(6):328–335. PMID 37076122 Β· DOI 10.1111/nep.14160

KFRE derivation and anchor IPDMA

  1. Tangri N, Stevens LA, Griffith J, Tighiouart H, Djurdjev O, Naimark D, Levin A, Levey AS. A predictive model for progression of chronic kidney disease to kidney failure. JAMA. 2011;305(15):1553–1559. PMID 21482743 Β· DOI 10.1001/jama.2011.451
  2. Tangri N, Grams ME, Levey AS, Coresh J, Appel LJ, Astor BC, et al; CKD Prognosis Consortium. Multinational Assessment of Accuracy of Equations for Predicting Risk of Kidney Failure: A Meta-analysis. JAMA. 2016;315(2):164–174. PMID 26757465 Β· DOI 10.1001/jama.2015.18202

Methodological references

  1. Snell KIE, Ensor J, Debray TPA, Moons KGM, Riley RD. Meta-analysis of prediction model performance across multiple studies: which scale helps ensure between-study normality for the C-statistic and calibration measures? BMC Med Res Methodol. 2018;18(1):84. DOI 10.1186/s12874-018-0535-5
  2. Debray TPA, Damen JAAG, Snell KIE, Ensor J, Hooft L, Reitsma JB, Riley RD, Moons KGM. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460. DOI 10.1136/bmj.i6460
  3. Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, Reitsma JB, Kleijnen J, Mallett S; PROBAST Group. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann Intern Med. 2019;170(1):51–58. PMID 30596875 Β· DOI 10.7326/M18-1376
  4. Debray TPA, Collins GS, Riley RD, Snell KIE, Van Calster B, Reitsma JB, Moons KGM. Transparent reporting of multivariable prediction models developed or validated using clustered data (TRIPOD-Cluster): explanation and elaboration. BMJ. 2023;380:e071058. PMID 36750236 Β· DOI 10.1136/bmj-2022-071058
  5. Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29–36. DOI 10.1148/radiology.143.1.7063747
  6. Steyerberg EW, Vergouwe Y. Towards better clinical prediction models: seven steps for development and an ABCD for validation. Eur Heart J. 2014;35(29):1925–1931. DOI 10.1093/eurheartj/ehu207
  7. Riley RD, Ensor J, Snell KIE, Debray TPA, Altman DG, Moons KGM, Collins GS. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges. BMJ. 2019;353:i3140. DOI 10.1136/bmj.i3140
  8. Paule RC, Mandel J. Consensus values and weighting factors. J Res Natl Bur Stand. 1982;87(5):377–385. (Paule-Mandel τ² estimator.)
  9. Hartung J, Knapp G. A refined method for the meta-analysis of controlled clinical trials with binary outcome. Stat Med. 2001;20(24):3875–3889. DOI 10.1002/sim.1009
  10. Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA (editors). Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (updated August 2024). Β§10.10.4.3 β€” prediction-interval convention tk-1.
  11. Riley RD, Higgins JPT, Deeks JJ. Interpretation of random effects meta-analyses. BMJ. 2011;342:d549. DOI 10.1136/bmj.d549
  12. Vergouwe Y, Moons KG, Steyerberg EW. External validity of risk models: use of benchmark values to disentangle a case-mix effect from incorrect coefficients. Am J Epidemiol. 2010;172(8):971–980. DOI 10.1093/aje/kwq223
  13. Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1–48. DOI 10.18637/jss.v036.i03
  14. Viechtbauer W. Confidence intervals for the amount of heterogeneity in meta-analysis. Stat Med. 2007;26(1):37–52. (q-profile τ² CI by inverting Cochran's Q against χ².) DOI 10.1002/sim.2514
  15. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley, 2009. Chapter 19 (subgroup analyses; Q_between as a Wald contrast on subgroup means).