Methodology

← Back to the board

What we include (and what we don't)

How rulings are made

Met / Missed / Partial are human judgments, each backed by one or more public sources (an obligation source that carries the verbatim promise, and where relevant a fulfillment source proving what was or wasn't delivered). Overdue / Upcoming are computed automatically from the deadline versus today, so those timers stay current on their own. Genuinely disputable rulings are flagged ⚠ contested and phrased as open questions.

Rating rubric

How we keep it current

The live overdue / upcoming countdowns recompute in your browser, so they are never stale. Separately, an automated job re-checks every cited source daily: that the link still resolves, that the verbatim quote is still on the page (catching silent edits), and it archives each source to the Wayback Machine so a ruling's basis can't quietly disappear. It also flags rulings that haven't had a human re-review in a while (30 days for contested/unresolved, 180 for settled) and watches lab policy pages for new dated promises. When a check needs attention, the affected row shows an ⚠ under review badge — we surface our own drift rather than hide it. The automation only detects and proposes; every Met/Missed/Partial verdict is still issued by a human. PDF sources are read with a text extractor so their clauses can be drift-checked too; a few obligations summarized from aspirational or multi-part commitments (with no single verbatim clause) are labeled summarized and exempt from quote-drift.

Editorial principles

Phrasing is strictly factual and neutral: we state the deadline and what shipped by it, and avoid editorializing verbs. A ruling ships only with a rock-solid source. Statuses are curated as of the "updated" date on the board; only the timers update live.

Data, license & how to cite

The dataset (src/data/*, /commitments.json, /commitments.csv) is licensed CC BY 4.0republish freely with attribution. The code is MIT. Nonprofits, journalists, and researchers are welcome to reuse it.

Suggested citation:

Overdue: Frontier AI Safety Commitment Tracker. https://overduetracker.org (retrieved 2026-08-05).

Download: JSON · CSV. Explore: table view. When a ruling changes or an error is fixed, it is logged on Corrections.

Spotted an error? Open a GitHub issue.

How Overdue differs (related trackers)

Overdue complements existing work rather than replacing it. The Midas Project's Seoul Tracker grades one collective deadline, and its Watchtower flags quiet policy changes; METR's index catalogs policy documents; the FLI AI Safety Index and SaferAI grade overall posture. AI Lab Watch compiled the broadest commitment list, but has been unmaintained since September 2025. Overdue's contribution is to bring those regimes together with a live, per-promise deadline clock: many individual dated commitments — RSPs, frontier safety frameworks, the Seoul and White House commitments — each with a status, an explicit upcoming / overdue countdown, and one source. It is not the first accountability project.