What we include (and what we don't)
- A specific, dated public promise — a calendar deadline or a falsifiable trigger. Vague or aspirational pledges are excluded.
- For the main board, a promise the lab itself made or signed (RSP/Preparedness/Frontier-Safety milestones, self-imposed deadlines, the Seoul and White House voluntary commitments). Government laws (e.g. the EU AI Act) are not promises a lab made — they appear separately as regulatory milestones, as countdowns, never scored kept or broken.
- One rock-solid public source per row (primary preferred); a "missed" ruling requires especially strong sourcing.
- Neutral, factual phrasing; genuinely debatable rulings are flagged
⚠ contested.
How rulings are made
Met / Missed / Partial are human judgments, each backed by one or more public sources (an obligation source that carries the verbatim promise, and where relevant a fulfillment source proving what was or wasn't delivered). Overdue / Upcoming are computed automatically from the deadline versus today, so those timers stay current on their own. Genuinely disputable rulings are flagged ⚠ contested and phrased as open questions.
Rating rubric
- Met — the required artifact or action was delivered by the deadline (or the trigger was satisfied), evidenced by a fulfillment source (primary preferred).
- Missed — the deadline passed (or the trigger fired) with no qualifying delivery. Requires especially strong sourcing: a primary admission, or two independent sources. A "missed" resting only on absence of evidence is always
⚠ contested. - Partial — a multi-part obligation was partly delivered, or delivered substantially late, or delivered in a materially weakened form; the notes enumerate what was and wasn't met.
- Contested — any ruling resting on secondary reporting, a derived deadline, an argument-from-absence, or a genuine interpretive dispute. Phrased as an open question, never as a verdict. A derived deadline (one we inferred from a cadence rather than a lab-stated date) is marked with a
†.
How we keep it current
The live overdue / upcoming countdowns recompute in your browser, so they are never stale. Separately, an automated job re-checks every cited source daily: that the link still resolves, that the verbatim quote is still on the page (catching silent edits), and it archives each source to the Wayback Machine so a ruling's basis can't quietly disappear. It also flags rulings that haven't had a human re-review in a while (30 days for contested/unresolved, 180 for settled) and watches lab policy pages for new dated promises. When a check needs attention, the affected row shows an ⚠ under review badge — we surface our own drift rather than hide it. The automation only detects and proposes; every Met/Missed/Partial verdict is still issued by a human. PDF sources are read with a text extractor so their clauses can be drift-checked too; a few obligations summarized from aspirational or multi-part commitments (with no single verbatim clause) are labeled summarized and exempt from quote-drift.
Editorial principles
Phrasing is strictly factual and neutral: we state the deadline and what shipped by it, and avoid editorializing verbs. A ruling ships only with a rock-solid source. Statuses are curated as of the "updated" date on the board; only the timers update live.
Data, license & how to cite
The dataset (src/data/*, /commitments.json, /commitments.csv) is licensed CC BY 4.0 — republish freely with attribution. The code is MIT. Nonprofits, journalists, and researchers are welcome to reuse it.
Suggested citation:
Overdue: Frontier AI Safety Commitment Tracker. https://overduetracker.org (retrieved 2026-08-05).Download: JSON · CSV. Explore: table view. When a ruling changes or an error is fixed, it is logged on Corrections.
Spotted an error? Open a GitHub issue.
How Overdue differs (related trackers)
Overdue complements existing work rather than replacing it. The Midas Project's Seoul Tracker grades one collective deadline, and its Watchtower flags quiet policy changes; METR's index catalogs policy documents; the FLI AI Safety Index and SaferAI grade overall posture. AI Lab Watch compiled the broadest commitment list, but has been unmaintained since September 2025. Overdue's contribution is to bring those regimes together with a live, per-promise deadline clock: many individual dated commitments — RSPs, frontier safety frameworks, the Seoul and White House commitments — each with a status, an explicit upcoming / overdue countdown, and one source. It is not the first accountability project.