Skip to main content

What changed — rapidmeta-finerenone data-integrity audit

Per-round summary of the automated integrity sweep · 15+ rounds · 0 human-reviewer hours · ← back to review index
6,295
Total fixes
26
Audit rounds
225
Flagged (no auto-fix)
735
Reviews audited

Latest sweep — 2026-05-25 → 2026-05-26 (4 commits)

Performance + encoding-correctness pass. Every change went through the existing CI pipeline (idempotency + sentinel + puppeteer smoke).

Performance

Encoding correctness

Idempotency + tests

Honest deferrals

Prior sweep — 2026-05-24 → 2026-05-25 (16 commits)

Two-day sweep across the whole portfolio. Every change shipped behind R-metafor parity validation and a 3-job CI pipeline (idempotency + sentinel + puppeteer smoke).

Reliability & correctness

Statistical methodology

Content & usability

Internal docs & tests

Pre-session state (preserved for historical reference)

532
Trustworthy (<0.30)
106
Low concern (0.30-0.50)
92
Manual review (0.50-0.70)
5
Quarantined (≥0.70)

Post-session: see audit_table.html for the current banded counts (Trustworthy 2,005 / LowConcern 82 / ManualReview 67 / Quarantined 0 as of 2026-05-25).

Per-trial identity-confidence score (R33 — 1,655 trials)

Composite of 25%·NCT-in-AACT + 20%·drug-in-AACT-intvs + 15%·PMID-present + 10%·year + 10%·name + 20%·evidence[]-non-empty.

718
Excellent (80-100)
545
Good (60-79)
268
Caution (40-59)
119
Poor (20-39)
5
Bad (0-19)

Audit rounds (chronological)

R1-init Initial 12-method audit + 8-agent scan + deterministic fixes
768 fixes
Sources: deterministic, agent
R1-multi R1 multi-agent (3 lenses, 77 trials)
11 fixes
Sources: multi-agent
R2 Multi-agent R2 (249 trials, 3 lenses)
204 fixes
Sources: multi-agent
R3 Multi-agent R3 (184 trials, fab-review deep audit)
230 fixes
Sources: multi-agent, quarantines
R4 R4 + risk classifier + aggressive cleanup
38 fixes
Sources: multi-agent, classifier
R5 R5 fidelity vs published landmark trials (16 NCTs)
1 fixes
Sources: landmark trials, calibration
R5b R5b pool reproduction vs 20 published reference MAs
Sources: published MAs, calibration
8-agent 8-agent blinded TruthCert audit (157 reviews, 1,100 trials)
57 fixes
Sources: 8 blinded agents, TruthCert HMAC
8-agent-ext 8-agent extended fixer + MED + carryforward
53 fixes
Sources: multi-agent
R6abc Internal-consistency suite (9 checks: direction, copy-paste, baseline-N…)
264 fixes
Sources: internal
R7 AACT cross-verify (LOW+MANUAL band, 1,063 trials)
100 fixes
Sources: AACT 2026-04-12
R7c AACT verify on OK band (248 trustworthy reviews)
139 fixes
Sources: AACT
R8 PMID-year era heuristic (PubMed indexing rate)
31 fixes
Sources: PubMed indexing
R9 AACT baseline_counts per-arm enrollment check
110 fixes
Sources: AACT baseline_counts
R10 AACT outcome_measurements event-count cross-check
554 fixes
Sources: AACT outcome_measurements
R19 AACT primary_completion_date vs extracted year
197 fixes
Sources: AACT completion dates
R20 Stress-test (re-run R5+R5b on cleaned corpus)
Sources: regression test
R22 MeSH-anchored review-topic vs AACT-conditions
422 fixes
Sources: AACT, curated MeSH dict
R23 Trial acronym vs AACT brief_title + id_information
249 fixes
Sources: AACT id_information
Path-A PMC OAI full-text evidence injection (29.8% PMC OA hit rate)
12 fixes
Sources: PMC OAI, NCBI E-utilities
R24 AACT design_outcomes type classification (flag-only)
225 flagged
Sources: AACT design_outcomes
R24b PubMed-directed R24 estimand auto-fix (conservative gate)
10 fixes
Sources: PubMed abstracts
R30 Fully-null trial-row cleanup
Sources: internal
R31-33 Polish suite (Benford after cleanup + decimal precision + confidence score)
Sources: audit metrics
R25-RW Retraction Watch DB cross-check (Crossref/GitLab, 70,147 retractions)
Sources: Retraction Watch
R26-perf terser engine minification (1,524 pages, 569 MB saved) + em-dash mojibake repair (1,321 pages) + Sentinel R7/R8 + clone_dashboard auto-minify hook
2,845 fixes
Sources: terser CLI (mangle disabled) · UTF-8 byte-level repair · per-rule WCAG validation

Ground-truth sources

Files