Literature Verification Ledger

AMiner tells you what others claimed. SwarmLabs tells you which claims hold up — and where they break. Every entry below was re-run through the SwarmLabs virtual-experiment engine (GP surrogate + uncertainty quantification). Each carries a verification badge, measured 95% interval coverage, and an OOD red-zone flag.

How to read a badge

PASS+PASS MARGINALUNVALIDATED
Coverage = share of held-out points falling inside the 95% prediction interval (target 0.95, marked by the violet line). Green ≥ 0.9 · Amber 0.5–0.9 · Red < 0.5 = OOD red zone (interval is overconfident; do not extrapolate here without real experiments). Anything not listed in this ledger is UNVALIDATED by default — we never imply verification we did not run.

Trust surface by kinetic family — one claim, cross-checked across papers

Verification is only as strong as its corroboration. Each card below groups every benchmark that shares a kinetic model, so you can see whether a model is backed by several independent papers or a single source — and whether the family as a whole survives our 3% noise floor. A family where all members fall in the OOD red zone is not yet trustworthy, no matter how often it is cited.

Why coverage matters more than accuracy. A model can score a high R² and still be dangerously overconfident: several benchmarks below fit the trend (R² > 0.9) while their 95% interval covers almost nothing. SwarmLabs normalizes the noise floor to 3% and never lowers it — lowering it would manufacture confidence.
How SwarmLabs differs from other AI research assistants →

Prove it live — the same OOD guard we apply

Paste your own training data and query conditions. SwarmLabs runs the deployed /api/v3/guard on the spot and returns the same pass / controlled / reject verdict — with calibrated uncertainty — that we use to flag the red zones above. No screenshot, no hand-wave. The prefilled example is the Monod benchmark; try 30.0 as a query to watch it land in the red zone.