Pooled C-statistic
Across β¦ external-validation cohorts the pooled 5-year C-statistic of the 4-variable KFRE on the logit scale was back-transformed to give the value below. Paule-Mandel REML ΟΒ² is reported on the logit-C scale; the HKSJ multiplier reflects the Cochrane v6.5 floor max(1, Q/(kβ1)) with tkβ1 critical value.
Heterogeneity
PROBAST overall risk-of-bias rollup
Tangri 2016 IPDMA consistency check
β¦
Subgroup contrast (derivation vs external)
All six cohorts in this subset are external validations of the KFRE β the derivation cohort (Tangri et al. 2011, Canadian community-CKD) is not re-included here. Subgroup tabulation is shown to demonstrate the dev-vs-external split capability of the engine.
| Bucket | k | Pooled C | 95% CI | ΟΒ² (logit) | IΒ² |
|---|
Included cohorts
Six external-validation cohorts of the 4-variable KFRE at the 5-year horizon. Each row links to PubMed via PMID; the DOI is also given. C-statistic and 95% CI are quoted verbatim from each publication's abstract; SE(C) is derived from the CI assuming Wald symmetry (SE = (CIhigh β CIlow) / (2 Γ 1.96)) for input to the logit-scale pool.
| Cohort & setting | PMID | Year | Country | N total | Events / FU (y) | C (95% CI) | SE(C) |
|---|
Cohort detail
Discrimination — C-statistic forest
Per-cohort C-statistic with Wald 95% CI (back-transformed from logit-C). Pooled diamond is Paule-Mandel REML + HKSJ with the Cochrane v6.5 small-sample floor. The dashed orange line is the 95% prediction interval (tk-1; the expected range for the C-statistic of a notional seventh external validation drawn from the same underlying population of validation studies).
Figure 1. Pooled C-statistic of the 4-variable Kidney Failure Risk Equation at the 5-year prediction horizon across six external validation cohorts. Blue squares are per-cohort point estimates with 95% CIs. The cyan diamond is the pooled estimate by Paule-Mandel REML ΟΒ² with HKSJ small-sample CI (Cochrane v6.5 floor max(1, Q/(kβ1)); Student-tk-1 critical value). The dashed orange bracket below the diamond is the 95% prediction interval (Higgins 2009; Riley 2011), reflecting the between-cohort variability in the underlying logit-C population. Heterogeneity is high (IΒ² β 97%; ΟΒ² point estimate on logit-C scale β 0.68 with q-profile 95% CI [β 0.25, β 4.42], Viechtbauer 2007), driven by a wide spread of validation settings: primary care (C = 0.93 in Major 2019), advanced CKD (C = 0.77β0.82 in Ali 2021 and Gallego-Valcarce 2024), kidney transplant recipients (C = 0.73 in Tangri 2020 β a different prediction context to the derivation cohort), and population data linkage (C = 0.96 in Irish 2023). The wide HKSJ-adjusted CI is a deliberate honest reflection of this heterogeneity, not a numerical artefact. The heterogeneity-fragility flag (parity with the pairwise-pool engine) is agree here: both the pooled CI and the prediction interval lie above logit(0.5) = 0, so the conclusion of useful discrimination is not heterogeneity-fragile.
Variance derivation: SE(C) is taken from the published 95% CI assuming Wald symmetry. Delta-method conversion to the logit scale: SE(logit C) = SE(C) / [C(1βC)]. Pooling on the logit-C scale follows Snell et al. 2018 (BMC Med Res Methodol) and Debray et al. 2017 (BMJ tutorial). Back-transformation is via the inverse-logit (sigmoid) function.
Calibration findings (narrative)
The six included cohorts reported calibration narratively without machine-extractable numerical intercepts, slopes, or O/E ratios. The aggregate qualitative pattern across studies is consistent: the original (non-recalibrated, non-North American calibrated) 4-variable KFRE over-predicts kidney failure risk in lower-risk validation cohorts and under-predicts in higher-risk advanced-CKD cohorts. Both directional miscalibrations are reported in the source publications.
| Cohort | Calibration finding (quoted or paraphrased from abstract) |
|---|---|
| Major 2019 (UK primary care) | Original KFRE over-predicted in lower (<20%) risk groups; recalibration via baseline-risk adjustment improved calibration (baseline survival 0.9832 β 0.9878 at 2 y; 0.9365 β 0.9570 at 5 y). |
| Lennartz 2016 (CARE FOR HOMe, Germany) | Original KFRE "well calibrated" per visual calibration plots; numerical calibration intercept/slope not provided in abstract. |
| Tangri 2020 (N. American transplant) | Numerical calibration not reported in abstract. Subgroup analysis in patients with eGFR < 45 mL/min/1.73 mΒ² gave higher 2 y C (0.88 vs 0.81) suggesting better discrimination near event horizon. |
| Ali 2021 (Salford advanced CKD) | "Underestimation of risk at 2 years and overestimation of risk at 5 years, especially in high-risk patients"; underestimation specifically reported in ADPKD aetiology subgroup. |
| Gallego-Valcarce 2024 (Spain advanced CKD) | "Excellent calibration"; Brier score < 0.20; Hosmer-Lemeshow goodness-of-fit p > 0.05. |
| Irish 2023 (Tasmania, Australia) | Calibration curves indicated systematic over-prediction; Brier score 0.01 at 5 y. Recalibration factors developed (baseline survival 0.9365 β 0.9633 at 5 y). |
Why calibration was not pooled
Numerical pooling of calibration intercept, calibration slope, and observed/expected ratio requires (i) the point estimate, (ii) the standard error, and (iii) the analysis scale, to be reported in the source publication. None of the six abstracts surfaced all three quantities simultaneously for any single calibration metric. Pooling the qualitative narrative would require unverifiable transformation. The engine's calibration pooling primitives are exercised in the unit-test suite (Paule-Mandel REML on raw-scale intercept, raw-scale slope, and log-O/E with back-transformation) and documented in tests/prediction_fixtures/r_metafor_baselines.json; for a real-cohort calibration pool, individual-participant data or full-text supplementary tables from the source studies are required.
PROBAST risk-of-bias
Per-cohort PROBAST judgement across the four standard domains (participants, predictors, outcome, analysis) following the Wolff et al. 2019 manual (Annals of Internal Medicine, PMID 30596875). Overall judgement: low only when every domain is rated low; high when any domain is rated high; unclear otherwise. Domain ratings are pragmatic assessments based on each cohort's published methods (sample size, multi-centre vs single-centre design, completeness of follow-up, completeness of predictor measurement).
Domain-by-domain rationale
Participants
The three large multi-centre / population-linkage cohorts (Major 2019, Tangri 2020, Irish 2023) drew participants from defined source populations with explicit inclusion criteria and minimal selection bias β rated low. Lennartz 2016 enrolled a tertiary nephrology referral population (selection toward advanced CKD; rated unclear). Ali 2021 and Gallego-Valcarce 2024 are single-centre advanced-CKD cohorts with small samples; Ali 2021 split into aetiology subgroups with small per-subgroup N raising bias concerns, and Gallego-Valcarce 2024 has N = 339 β rated high for participants.
Predictors
The 4-variable KFRE predictors (age, sex, eGFR, urine ACR) are routinely measured in all included cohorts with standardised laboratory methods. All six rated low.
Outcome
Kidney failure (initiation of kidney replacement therapy or estimated GFR < 15 mL/min/1.73 mΒ² in sustained measurements) is an objective hard endpoint with standardised ascertainment in registry-linked or EHR-linked cohorts. All six rated low.
Analysis
Large cohorts with high events-per-variable (Major 2019 with 429 events, Tangri 2020 pooled, Irish 2023 with 285 KRT events) rated low. Lennartz 2016 had moderate event count in a small sample (rated unclear). Ali 2021 (small per-aetiology subgroups) rated unclear. Gallego-Valcarce 2024 (N = 339, abstract does not specify event count) raises concern about events-per-variable and is rated high for analysis pending full-text events-per-variable verification.
Overall rollup
Three cohorts rated overall low (Major 2019, Tangri 2020, Irish 2023), two rated high (Ali 2021, Gallego-Valcarce 2024), one rated unclear (Lennartz 2016). A sensitivity analysis restricted to the three low-RoB cohorts would re-pool over a much narrower range of validation contexts (large-scale primary care, multi-centre transplant, population data linkage) and is recommended as a follow-up.
Caveat. These domain ratings are this review's interpretation of each cohort's published methods and are intended for engine demonstration, not as a definitive PROBAST assessment. A definitive PROBAST application would require two independent reviewers consulting each full-text article, with adjudication of disagreements.
Methods
Eligibility
External validations of the 4-variable Kidney Failure Risk Equation (Tangri 2011; age + sex + eGFR + urine ACR predicting kidney replacement therapy) that report a numerical C-statistic with 95% CI at the 5-year horizon. Both the original KFRE and "non-North-American-calibrated" KFRE were included; recalibrated KFRE versions are noted but not re-pooled. Cohorts in kidney-transplant recipients are included as a clinically-distinct validation context, with the caveat that the prediction target (graft loss vs native-kidney KRT) differs from the derivation cohort.
Search and selection
Six independently-published external validation cohorts were identified through anchor-search on the Tangri 2016 multinational IPDMA (JAMA, PMID 26757465) and subsequent KFRE-validation literature. This is a methods-engine demonstration on a verifiable subset of the broader KFRE-validation literature, not an exhaustive systematic review. For comprehensive coverage, the 31-cohort IPDMA (Tangri 2016) and the recent 59-cohort assessment with CKD-EPI 2021 eGFR (PMC10103205) remain authoritative.
Data extraction
For each cohort the C-statistic and its 95% CI were extracted verbatim from the publication's PubMed abstract. SE(C) was derived from the CI assuming Wald symmetry: SE = (CIupper β CIlower) / (2 Γ 1.96). Cohort N, country, setting, and follow-up duration were also extracted where reported. Calibration metrics (intercept, slope, O/E ratio, Brier score) were extracted where numerically reported with associated SE; absence of full numerical calibration triples meant calibration was not pooled (see Calibration tab).
Statistical pooling
The 4-variable KFRE 5-year C-statistic was pooled on the logit scale (Snell 2018; Debray 2017). For each cohort:
- logit(C) = log(C / (1 β C));
- SE on the logit scale via delta method: SE(logit C) = SE(C) / [C(1 β C)];
- vi = [SE(logit C)]Β² .
Random-effects pooling was via Paule-Mandel REML β an exact REML estimator for the single-axis case that iterates ΟΒ² to solve Ξ£ wi(yi β ΞΌw)Β² = k β 1 with wi = 1/(vi + ΟΒ²). Advanced-stats.md guidance: PM is preferred over DerSimonian-Laird for k < 10 because DL is downward-biased at small k. The pooled mean was the inverse-variance-weighted average at the converged ΟΒ²; SE was sqrt(1 / Ξ£ wi).
Small-sample CI used the Hartung-Knapp-Sidik-Jonkman adjustment with the Cochrane Handbook v6.5 (Nov 2024 Β§10.10.4.3) floor: HKSJ scale factor = max(1, Q / (k β 1)), with Student-tk β 1 critical value rather than the normal-z. This protects against CI under-coverage in the presence of heterogeneity. A 95% prediction interval (Higgins 2009; Riley 2011) was reported using tk β 1 Γ β(ΟΒ² + SEΞΌΒ²), defined for k β₯ 3.
In addition to the ΟΒ² point estimate the engine reports a 95% q-profile confidence interval for ΟΒ² (Viechtbauer 2007, Statistics in Medicine), obtained by bisection-inverting Cochran's Q against ΟΒ²(kβ1) at Ξ±/2 and 1 β Ξ±/2. This is the same algorithm used by the DTA engine's qProfileTau2CI. A wide ΟΒ² CI is informative: it confirms that the heterogeneity estimate itself is uncertain at small k, separately from the ΟΒ² magnitude. A heterogeneity-fragility flag (parity with the pairwise-pool engine's piGap) reports whether the pooled CI and prediction interval disagree about whether a clinically-meaningful reference value lies inside the interval β reference values are logit(0.5) = 0 for the C-statistic, 0 for calibration intercept, 1 for calibration slope, and log(1) = 0 for the O/E ratio. The flag takes three values: agree, ci_excludes_pi_includes (pooled CI claims discrimination but a future cohort could be at chance β heterogeneity-fragile), or ci_includes_pi_excludes (CI is null but PI is shifted).
All pooling primitives, back-transformations, and the PROBAST rollup are implemented in the open-source rapidmeta-prediction-engine-v1.js (~ 826 lines, MIT licence, no external dependencies). The engine is validated by a 64-assertion unit-test suite (tests/test_prediction_engine.mjs) including: closed-form ΟΒ² = 0 case (hand-derived against inverse-variance algebra), heterogeneous k = 3 case where PM converges to a value identical to DerSimonian-Laird (verified by hand on symmetric balanced data), the 6-cohort KFRE real-data frozen baseline, the Hanley-McNeil 1982 variance against the 1982 Radiology paper Table II hand-calculation, Student-t and chi-squared quantiles against Abramowitz & Stegun, and various edge-case stress paths.
Risk-of-bias assessment
PROBAST (Wolff et al. 2019, PMID 30596875) was applied across the four prediction-model-specific domains: participants (selection bias, source population), predictors (assessment and definition), outcome (ascertainment and blinding), and analysis (events-per-variable, handling of missing data, statistical method). Each domain was rated low / high / unclear; overall: low only if all four domains low, high if any domain high, unclear otherwise. Domain rationale is provided on the PROBAST tab.
Cross-validation against R
The engine's pooling primitives (Paule-Mandel iteration; HKSJ floor with tk β 1; logit / inv-logit / Hanley-McNeil variance) are the same primitives metafor::rma(yi, vi, method='PM', test='knha') uses for univariate meta-analyses, and the same primitives metamisc::valmeta() uses for prediction-model meta-analyses. Three frozen baseline cases are recorded in tests/prediction_fixtures/r_metafor_baselines.json: (i) homogeneous ΟΒ² = 0 closed-form, (ii) heterogeneous k = 3 case where PM converges to the DL value 0.24 by analytic derivation, (iii) the 6-cohort KFRE real-data pool with all key intermediates frozen to engine-output precision. Bit-reproducibility against metafor on identical inputs is the planned v1.1 build-time check.
Discussion
Principal findings
Across six independently-published external validations the pooled 5-year C-statistic of the 4-variable KFRE on the logit scale was 1.95 (back-transformed C = β¦; 95% HKSJ CI β¦; 95% prediction interval β¦). The Tangri 2016 multinational IPDMA (31 cohorts, n = 721,357) reported a pooled 5-year C of 0.88 (95% CI 0.86β0.90); the present 6-cohort subset is consistent in central tendency (β¦) but produces a substantially wider HKSJ-adjusted CI because of the heterogeneity inherent in pooling validations across very different prediction contexts.
Heterogeneity
Between-cohort heterogeneity is high (ΟΒ² β 0.68 on the logit-C scale; IΒ² β 97%). The principal driver is the diversity of validation contexts. At one extreme, Major 2019 (UK primary care, 35,539 participants) and Irish 2023 (Tasmania population data linkage, 8,182 participants at 5 y) report C β 0.92β0.96 β large unselected cohorts where the KFRE's calibration on age, sex, eGFR, and ACR maps efficiently to a kidney-failure rate that is rare in the overall population. At the other extreme, the kidney-transplant cohort (Tangri 2020, C = 0.73) and the advanced-CKD subgroups (Ali 2021 in Salford and Gallego-Valcarce 2024 in Spain, C = 0.77β0.82) report substantially lower discrimination, consistent with the principle that discrimination shrinks when the case-mix narrows (Riley et al. 2019, BMJ; Vergouwe et al. 2010). In these populations most participants are at high baseline kidney-failure risk and the variation in predicted risk between participants is small, so distinguishing future events from non-events on the basis of the same predictors is harder.
This pattern is the textbook example of why prediction-model meta-analysis must pool with random effects and report the prediction interval alongside the pooled CI: a notional seventh validation cohort would be expected to land somewhere in the prediction interval (β¦), not within the much-narrower CI. The HKSJ adjustment with the Cochrane v6.5 small-sample floor is what produces this honest wider CI; using a normal-z CI without the small-sample correction would produce a misleadingly tight pooled estimate.
Risk-of-bias and applicability
Three of six cohorts were rated overall low risk of bias (Major 2019, Tangri 2020 transplant, Irish 2023). Two were rated high (Ali 2021 with small per-aetiology subgroups; Gallego-Valcarce 2024 with N = 339 and events-per-variable concerns). Lennartz 2016 was rated unclear (tertiary referral selection). A sensitivity analysis restricted to the three low-RoB cohorts would re-pool over primary-care, transplant, and population-linkage settings only β narrower applicability but with greater methodological coherence.
Implications for clinical risk stratification
The KFRE was developed and validated for native-kidney CKD populations (Tangri 2011, JAMA). Its application to kidney-transplant recipients (Tangri 2020) gives substantially lower discrimination (C = 0.73 vs β 0.88 in native CKD), reflecting that graft-loss biology differs from native-kidney-failure biology. For Finrenone-relevant CKD populations β i.e. native-kidney chronic kidney disease with type 2 diabetes, where Finrenone is indicated β the KFRE's primary-care discrimination is excellent (Major 2019: C = 0.92) and its advanced-CKD discrimination is good (Ali 2021: C = 0.77). The KFRE is therefore well suited for Finrenone-eligible-cohort risk stratification at the primary-care to early-CKD-clinic transition, but a recalibration step (Major 2019 baseline-risk adjustment; Irish 2023 calibration factors) is required outside North American populations.
Strengths
- All cohort-level numerical values are extracted verbatim from PubMed-indexed publications with PMIDs and DOIs; no synthetic or fabricated data.
- Pooling uses Paule-Mandel REML (preferred for k < 10 per advanced-stats.md) with HKSJ small-sample CI on the logit-C scale (Snell 2018; Debray 2017).
- Prediction interval is reported alongside the CI, surfacing the between-cohort variability that the CI alone hides.
- PROBAST is applied per-domain with explicit rationale for each cohort.
- The engine is open-source under MIT with a 64-assertion unit-test suite, including a frozen baseline against the 6-cohort real-data pool for regression testing.
Limitations
- Calibration metrics are not pooled because the six abstracts do not jointly report numerical calibration intercept / slope / O/E with SEs. The narrative qualitative pattern (over-prediction in low-risk groups, under-prediction in advanced CKD) is documented but cannot be turned into a forest plot from the source data available.
- The six cohorts are a subset of the broader KFRE-validation literature; the Tangri 2016 IPDMA (31 cohorts) is the authoritative pooled estimate, and the present review is positioned as a methods-engine demonstration, not as a replacement.
- SE(C) is derived from the published CI assuming Wald symmetry, which is approximate; this is the same approximation used by metafor::rma and metamisc::valmeta when only CIs are reported.
- The Tangri 2020 transplant cohort validates the KFRE in a clinically-distinct population (graft loss / death-censored), which differs from native-kidney KFRE prediction; its inclusion contributes to the high IΒ² and the wider HKSJ CI.
- PROBAST domain ratings are this review's interpretation of each cohort's published methods, not a definitive two-reviewer adjudicated assessment.
- Bit-reproducibility against
metafor::rma(method='PM', test='knha')andmetamisc::valmeta()is documented as the planned v1.1 build-time check; the engine's pooling primitives match metafor's by construction (Paule-Mandel iteration, HKSJ floor with tk β 1), but a build-time WebR or Rscript invocation has not yet been wired.
Conclusion
The 4-variable KFRE shows good-to-excellent 5-year discrimination across diverse external validation contexts when pooled honestly with Paule-Mandel REML + HKSJ and a Cochrane v6.5 prediction interval. The wide pooled CI reflects real between-cohort variability driven by validation-context heterogeneity (primary-care vs advanced-CKD vs transplant), not engine pathology. For Finrenone-relevant native-kidney CKD populations the primary-care KFRE discrimination is excellent. Calibration assessment requires individual-cohort full-text data and was outside the scope of this abstract-level methods demonstration.
TRIPOD-Cluster reporting checklist
TRIPOD-Cluster (Debray et al. 2023, BMJ; PMID 36750236) extends TRIPOD to multi-cluster / multi-cohort prediction-model evaluation studies. Items below mark this review's compliance for engine-demonstration purposes; a clinical-evidence submission would require full-text extraction beyond abstract-level data.
| Item | Description | Status |
|---|---|---|
| 1a | Identify as prediction-model meta-analysis | β Title states KFRE external-validation meta-analysis |
| 1b | Abstract: structured summary | β See Overview tab |
| 2 | Background and objectives | β Methods tab |
| 3a | Source of data per cohort | β Cohorts tab (PMID, DOI, setting, N, country, follow-up) |
| 3b | Inclusion / exclusion criteria across cohorts | β Methods β Eligibility |
| 4 | Outcome definition | β Kidney replacement therapy / sustained eGFR < 15 (per cohort sources) |
| 5 | Predictors and their measurement | β 4-variable KFRE: age, sex, eGFR, urine ACR (Tangri 2011) |
| 6 | Sample size per cohort | β Cohorts table |
| 7 | Missing-data handling | N/A at abstract-level extraction |
| 8 | Statistical pooling method | β Methods β Statistical pooling (PM REML + HKSJ) |
| 9 | Risk-of-bias assessment | β PROBAST tab (per-domain rationale) |
| 10 | Discrimination metric pooled | β C-statistic on logit scale |
| 11 | Calibration metric pooled | Narrative only β see Calibration tab for rationale |
| 12 | Heterogeneity (ΟΒ², IΒ², Q) | β Overview β Heterogeneity grid |
| 13 | Prediction interval | β Overview + Discrimination figure caption |
| 14 | Subgroup / sensitivity analyses | β Derivation-vs-external subgroup (Overview); low-RoB sensitivity recommended (Discussion) |
| 15 | Limitations | β Discussion β Limitations |
| 16 | Funding / conflicts | N/A (engine demonstration; no clinical-funding declarations) |
| 17 | Code / data availability | β Engine source: rapidmeta-prediction-engine-v1.js; cohort fixture: tests/prediction_fixtures/kfre_5y_external_validations.json; R baseline: tests/prediction_fixtures/r_metafor_baselines.json |
Note: this checklist is a self-assessment for engine-demonstration transparency. A clinical-evidence submission would require independent two-reviewer extraction from full-text PDFs and a definitive PROBAST adjudication.
References
Vancouver-style citations with PubMed and DOI links.
Included cohorts (n = 6)
- Major RW, Shepherd D, Medcalf JF, Xu G, Gray LJ, Brunskill NJ. The Kidney Failure Risk Equation for prediction of end stage renal disease in UK primary care: An external validation and clinical impact projection cohort study. PLoS Med. 2019;16(11):e1002955. PMID 31693662 Β· DOI 10.1371/journal.pmed.1002955
- Lennartz CS, Pickering JW, Seiler-MuΓler S, Bauer L, Untersteller K, Emrich IE, Zawada AM, Radermacher J, Tangri N, Fliser D, Heine GH. External Validation of the Kidney Failure Risk Equation and Re-Calibration with Addition of Ultrasound Parameters. Clin J Am Soc Nephrol. 2016;11(4):609β615. PMID 26787778 Β· DOI 10.2215/CJN.08110715
- Tangri N, Ferguson TW, Bamforth RJ, Sereda G, Kumar M, Alkurd I, Tageldin T, Hingwala J, Knoll GA, Komenda P, Rigatto C. Validation of the Kidney Failure Risk Equation in Kidney Transplant Recipients. Can J Kidney Health Dis. 2020;7:2054358120922627. PMID 32549052 Β· DOI 10.1177/2054358120922627
- Ali I, Donne RL, Kalra PA. A validation study of the kidney failure risk equation in advanced chronic kidney disease according to disease aetiology with evaluation of discrimination, calibration and clinical utility. BMC Nephrol. 2021;22(1):194. PMID 34030639 Β· DOI 10.1186/s12882-021-02402-1
- Gallego-Valcarce E, Cobo-Caso MA, SΓ‘nchez-Horrillo A, Ortega-Cerrato A, PΓ©rez-MartΓnez J, PΓ©rez-GarcΓa R. External validation of the KFRE and Grams prediction models for kidney failure and death in a Spanish cohort of patients with advanced chronic kidney disease. J Nephrol. 2024;37(2):393β405. PMID 38060108 Β· DOI 10.1007/s40620-023-01819-1
- Irish GL, Cuthbertson L, Kitsos A, Saunder T, Clayton PA, Jose MD. The kidney failure risk equation predicts kidney failure: Validation in an Australian cohort. Nephrology (Carlton). 2023;28(6):328β335. PMID 37076122 Β· DOI 10.1111/nep.14160
KFRE derivation and anchor IPDMA
- Tangri N, Stevens LA, Griffith J, Tighiouart H, Djurdjev O, Naimark D, Levin A, Levey AS. A predictive model for progression of chronic kidney disease to kidney failure. JAMA. 2011;305(15):1553β1559. PMID 21482743 Β· DOI 10.1001/jama.2011.451
- Tangri N, Grams ME, Levey AS, Coresh J, Appel LJ, Astor BC, et al; CKD Prognosis Consortium. Multinational Assessment of Accuracy of Equations for Predicting Risk of Kidney Failure: A Meta-analysis. JAMA. 2016;315(2):164β174. PMID 26757465 Β· DOI 10.1001/jama.2015.18202
Methodological references
- Snell KIE, Ensor J, Debray TPA, Moons KGM, Riley RD. Meta-analysis of prediction model performance across multiple studies: which scale helps ensure between-study normality for the C-statistic and calibration measures? BMC Med Res Methodol. 2018;18(1):84. DOI 10.1186/s12874-018-0535-5
- Debray TPA, Damen JAAG, Snell KIE, Ensor J, Hooft L, Reitsma JB, Riley RD, Moons KGM. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460. DOI 10.1136/bmj.i6460
- Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, Reitsma JB, Kleijnen J, Mallett S; PROBAST Group. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann Intern Med. 2019;170(1):51β58. PMID 30596875 Β· DOI 10.7326/M18-1376
- Debray TPA, Collins GS, Riley RD, Snell KIE, Van Calster B, Reitsma JB, Moons KGM. Transparent reporting of multivariable prediction models developed or validated using clustered data (TRIPOD-Cluster): explanation and elaboration. BMJ. 2023;380:e071058. PMID 36750236 Β· DOI 10.1136/bmj-2022-071058
- Hanley JA, McNeil BJ. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology. 1982;143(1):29β36. DOI 10.1148/radiology.143.1.7063747
- Steyerberg EW, Vergouwe Y. Towards better clinical prediction models: seven steps for development and an ABCD for validation. Eur Heart J. 2014;35(29):1925β1931. DOI 10.1093/eurheartj/ehu207
- Riley RD, Ensor J, Snell KIE, Debray TPA, Altman DG, Moons KGM, Collins GS. External validation of clinical prediction models using big datasets from e-health records or IPD meta-analysis: opportunities and challenges. BMJ. 2019;353:i3140. DOI 10.1136/bmj.i3140
- Paule RC, Mandel J. Consensus values and weighting factors. J Res Natl Bur Stand. 1982;87(5):377β385. (Paule-Mandel ΟΒ² estimator.)
- Hartung J, Knapp G. A refined method for the meta-analysis of controlled clinical trials with binary outcome. Stat Med. 2001;20(24):3875β3889. DOI 10.1002/sim.1009
- Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA (editors). Cochrane Handbook for Systematic Reviews of Interventions, version 6.5 (updated August 2024). Β§10.10.4.3 β prediction-interval convention tk-1.
- Riley RD, Higgins JPT, Deeks JJ. Interpretation of random effects meta-analyses. BMJ. 2011;342:d549. DOI 10.1136/bmj.d549
- Vergouwe Y, Moons KG, Steyerberg EW. External validity of risk models: use of benchmark values to disentangle a case-mix effect from incorrect coefficients. Am J Epidemiol. 2010;172(8):971β980. DOI 10.1093/aje/kwq223
- Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1β48. DOI 10.18637/jss.v036.i03
- Viechtbauer W. Confidence intervals for the amount of heterogeneity in meta-analysis. Stat Med. 2007;26(1):37β52. (q-profile ΟΒ² CI by inverting Cochran's Q against ΟΒ².) DOI 10.1002/sim.2514
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley, 2009. Chapter 19 (subgroup analyses; Q_between as a Wald contrast on subgroup means).