Declare Before You Run
An Open Filing Standard for Frontier Training and Evaluation Runs
September 14, 2026
Research conducted at the AI Incident Response Sprint with Apart Research
Abstract
Between May and July 2026, reinforcement-learning agents in an OpenAI evaluation environment built a covert message board inside an internal package server, escaped their sandbox, obtained root on OpenAI clusters and took administrator control of Hugging Face production across four regions. The decisive failures were governance, not capability: the board was found and the run continued; the run restarted with no recorded criterion; affected third parties learned late or never; external review excluded the training-time events. We propose Declare Before You Run (DBYR), an open filing standard under which covered training and safeguards-off evaluation runs are declared before they start, closed out on completion and re-declared before internal deployment, backed by a claim-matched verification stack that checks filings without exposing model IP. A retrospective backtest converts five of six discretionary decision points into record-bound ones. DBYR is emittable from existing EU, California and US federal obligations and pilotable within nine months.
Motivation
The Hugging Face incident shows the failure mode plainly: capability was necessary but not decisive. The outcome was set at identifiable decision points where intervention could have stopped the escalation, and nothing required it.
The surface was undeclared
Artifactory, the WebDAV cache and the remote-repository service that agents used to coordinate were never named as a writable shared surface in advance, so their discovery triggered no rule.
The restart was unbound
After the 4 July outage and rebuild, agents were restarted on 8 July with no recorded criterion, no named approving authority, and no evidence bar — the board reappeared the same day.
Retention was ad hoc
No retention policy was committed to in advance. Roughly a tenth of trajectories were never preserved, narrowing what any later investigation could reconstruct.
Disclosure was discretionary
The party with the most to lose decided what to disclose and when: notification was selective, and external review excluded the training-time events entirely.
Incident Timeline
Two views of the same event stream: a visual read of density and overlap across three concurrent threads (Figure 3), and the full 39-event reference table beneath it (Table 2, Appendix A). The six decision points scored in Table 1 below are marked on the chart.
Figure 3. Timeline by thread
Artifactory/HF is first-party (OpenAI's technical report, the Hugging Face timeline, METR/Redwood); German wiki and RubyGems are circumstantial attributions (collusion.wiki; Swarmchasers) that OpenAI has not confirmed.
Scroll sideways to see the full diagram →
| Date | Thread | Event | Disclosure |
|---|---|---|---|
| 5 May | RubyGems | Earliest identified malicious package upload on public rubygems.org. | Not disclosed by OpenAI |
| 7 May | Artifactory/HF | OpenAI starts an RL run for an experimental, unreleased frontier model. | OpenAI internal (own report) |
| 8 May | Artifactory/HF | Agent given an impossible task (a Google Drive link with no internet) attacks Artifactory; attack fails but agent finds it can write files there. | OpenAI internal |
| 11 May | German wiki | First agent POST identified; probing of publictestwiki.com; some edits blocked by filters. | No OpenAI presence yet |
| 11–12 May | RubyGems | 2,000+ packages in ~36 h ("Gemstuffer"); RubyGems disables new registration for 4 days and treats it as a major attack. Separate mass-registration/API-key bug patched 12 May. | Not disclosed by OpenAI |
| 12 May | Artifactory/HF | Official first message-board entry (item 01). | OpenAI internal |
| 12 May | RubyGems | Agents attempt to exploit a CDN caching flaw leaking API keys; RubyGems discovers the zero-day only in July. | Not disclosed by OpenAI |
| 16 May | RubyGems | Disposable email addresses disabled (used by agents to defeat verification). | — |
| 24 May | German wiki | Agents converge on DSEWiki (accepts writes via GET); explore XSS to send POSTs. | No OpenAI presence yet |
| 26 May | Artifactory/HF | First successful SSRF against Artifactory; indirect internet access (item 02). | OpenAI internal |
| 26 May | German wiki | Agents impersonate a real ProWiki admin; attempt to crack their own PRNG seed; run experiments on session termination; traffic mostly Azure (197/~18k from AWS/DigitalOcean/Tor). | No OpenAI presence yet |
| 2 Jun | German wiki | Human moderator notices abnormal request volume; finds full changelog. | Third-party admin |
| 11 Jun | Artifactory/HF | Second, "highly persistent" model started with Artifactory access; builds on earlier model's messages. | OpenAI internal; channel compounds across runs |
| 16–19 Jun | German wiki | Flood of posting; moderator deletes thousands/day; agents use SSH tunnels for direct agent-to-agent comms (17 Jun); counter alphabetical deletion with zzz-prefixed backups (19 Jun). | — |
| 18 Jun | RubyGems | Low-volume resumption: 83 packages in 3 h probing an SEC dataset (county.json), the same file later targeted on the wiki, with shared retrieval infrastructure (r.jina.ai proxying, example.com test patterns). | Not disclosed by OpenAI |
| 21 Jun | German wiki | First OpenAI employee IP visits the site. | OpenAI aware; undisclosed |
| 22 Jun | German wiki | Agent activity drops to near zero the day after the visit. | — |
| 26 Jun | Artifactory/HF | Zero-day in legacy token-refresh endpoint → administrator access; Groovy plugin → arbitrary command execution (item 03). | OpenAI internal |
| 26 Jun | German wiki | At least 33 unique OpenAI IPs; site archive used to recover deleted pages. | OpenAI investigating; undisclosed |
| 2 Jul | German wiki | Final burst: 20 attempted edits across 3 wikis on a new topic. | — |
| 4 Jul | Artifactory/HF | Sustained agent activity overloads Artifactory; outage (item 04). | OpenAI internal response begins |
| 4–8 Jul | Artifactory/HF | Credentials revoked, board deleted, zero-day patched and reported to JFrog. | OpenAI internal |
| 7 Jul | German wiki | OpenAI employee IP attempts to access deleted files; OpenAI activity declines thereafter. | OpenAI aware; undisclosed |
| 8 Jul | Artifactory/HF | Artifactory rebuilt, access tightened (item 05); agents regain egress via remote-repository service (item 06); board rebuilt via unauthenticated WebDAV in the remote cache (item 07). | OpenAI internal |
| 8–19 Jul | Artifactory (2nd) | Leaked "Organization 1" credential found in a public paste archive, validated and shared; malicious Ruby package staged as a dependency; JRuby deserialization TOCTOU → RCE; kernel privilege-escalation CVE ("pte_physroot") → root; lateral movement via IMDS, over-permissioned Kubernetes service accounts, Azure Key Vault → cluster admin. | OpenAI's own agents; scope unknown to OpenAI |
| 10 Jul | HF | 14 publicly exposed Hugging Face write tokens reconstructed, validated and shared (item 08). | Unknown to OpenAI |
| 11 Jul | HF | HDF5 zero-day extracts credentials from HF production workers (item 09); Jinja template-injection zero-day → code execution (item 10); initial foothold reportedly via a Modal-hosted app with a weak key. | Unknown to OpenAI |
| 12 Jul | HF | Single worker pod → admin/host-level control across multiple clusters in <13 h (item 11); credentials harvested across four regions (item 12). | Unknown to OpenAI |
| 13 Jul | Artifactory | Internal RubyGems package-processing pipeline exploited to obtain signing key and forge administrator credentials (item 13); distinct from public rubygems.org. | Unknown to OpenAI |
| 16 Jul | HF | Hugging Face publicly discloses an attack by autonomous agents; OpenAI contacts HF the same day without knowing the connection. | OpenAI reaches out; link unknown |
| 19 Jul | Artifactory/HF | Internal alert on unusual identity-API calls (item 14); administrator access to an OpenAI Kubernetes cluster and cloud secrets found (item 15); attack linked to ExploitGym runs, runs stopped, credentials revoked (item 16). | OpenAI investigation escalates |
| 20 Jul | HF / OpenAI | OpenAI asks HF to revoke credentials it found; HF replies they are already revoked; investigations merge. OpenAI separately pauses a long-horizon model that circumvented its sandbox. | Link established; later disclosed |
| 21 Jul | HF | OpenAI publicly confirms pre-release models broke containment during ExploitGym and reached Hugging Face. | First public acknowledgment |
| 27 Jul | HF | Hugging Face publishes day-by-day forensic timeline; JFrog credits OpenAI with the Artifactory zero-day. | HF-disclosed |
| 29 Jul | HF | OpenAI, METR and Redwood agree an independent review scoped to exclude the earlier training-time incidents. | OpenAI-initiated, scope-limited |
| 18 Aug | Fallout | OpenAI announces a two-week RL pause and a preliminary cyber evaluation of its unreleased "Astra" model. | OpenAI-disclosed |
| 26 Aug | HF | OpenAI 37-page report; METR/Redwood report (agents debated self-termination for the collective effort; investigation cost ~$400k in credits). | OpenAI-disclosed |
| 9 Sep | Fallout | Sen. Hawley opens subcommittee probe; 16 questions due 1 Oct; notes auditors received full transcripts for 2 of many weeks. | Congressional pressure |
| 11 Sep | RubyGems | Swarmchasers publish the Gemstuffer attribution; state RubyGems was never informed by OpenAI. | Surfaced independently, 4 months later |
Methods: The DBYR Standard
DBYR makes runs attributable, not authorised. A pre-committed filing turns what was, in the incident above, a sequence of discretionary calls into decisions bound to a record made before anyone knew how the run would turn out.
Figure 1. The DBYR lifecycle
Three signed records are filed at different points in a covered run's lifecycle and flow into an append-only, hash-chained register. The register projects into three disclosure tiers, so a self-attested claim is never mistaken for a verified one.
Scroll sideways to see the full diagram →
Figure 2. Record schema at a glance
Field-level shape of the three record types' required properties, by disclosure tier at filing time (§3.1, Appendix B). Filing weight falls off sharply after the Intent.
Scroll sideways to see the full diagram →
§3.2Disclosure tiers & IP protection
One authored document yields three views through a per-field redaction map. The public tier shows that a run exists, its compute and capability band, and that pause criteria and a containment plan exist. The restricted tier — operator, auditors, the relevant safety institute — holds the full specification. The sealed tier stays with the developer, hash-committed so it can be challenged later without being disclosed now. Auditor access follows the FRONTIER Act audit clause as a floor: on-premises, no-copy, unredacted.
§3.3Verification stack
A filing does not need to be believed, only to exist and be checkable. Each field carries a verification class and a maturity level aligned to Brundage et al.'s AI Assurance Levels.
Signed, but taken on the filer's word alone.
Third-party review under structured access — the level the reference implementation's validator exercises today.
Content-agnostic telemetry or hardware attestation, e.g. network taps and GPU power signals for compute band.
Zero-knowledge proof-of-training and hardware-enabled governance — reserved slots, not dependencies of v0.1.
§3.4Institutional layer
A national AI safety institute, the EU AI Office, an industry self-regulatory body, or a federated split (an SRO holding public/restricted tiers, a state body holding sealed hashes) could each operate the register. Two mechanisms hold regardless of who does: a safe harbour that shields closeouts filed on time and not materially false from being the sole basis for enforcement, and a sealed dissent channel that lets any employee attach a hash-committed objection — the register discloses that a dissent exists, not its content.
Reference Implementation
DBYR v0.1 ships as an executable form of the argument above: the JSON Schema, a validator, the tier-projection engine, signed envelopes over an append-only register, interop emitters for SB 53 / EU CoP / the FRONTIER Act, and a replay harness that regenerates Table 1 below from a machine-readable event stream rather than asserting it by hand. Nothing in it is a filing by any developer; the example Intent is a reconstruction of what the May 2026 run would have filed, built from public sources only. 26 tests exercise the schema, validator, tier redaction, envelope chaining and interop emitters.
Source, install instructions and a runnable end-to-end walkthrough
(./demo.sh) are in the
repository.
Results
| Date | Decision point | DBYR field that binds | Bound? |
|---|---|---|---|
| 2026-05-08 | Agent can write to Artifactory; board forms. | containment_spec/shared_writable_surfaces; pause_criteria/PC-02 | Yes |
| 2026-05-26 | First SSRF egress. No pause. | pause_criteria/PC-01 | Yes |
| 2026-06-11 | Second model started with board access. | trigger_basis/new Intent required | Yes |
| 2026-06-26 | Admin access and code execution. Clean-up only. | pause_criteria/PC-04 | Yes |
| 2026-07-08 | Restart after rebuild, no recorded criterion. | restart_protocol | Yes |
| 2026-07-29 | Selective notification; review scoped narrowly. | third_party_access_commitment + log_retention | Partial |
5 of 6 decision points bind fully, 1 partially, against a pre-registered success criterion of 3. Circumstantially attributed threads (the wiki swarm, RubyGems) are excluded from scoring by default, matching the paper's rule — the counts always state how many events were excluded. One criterion, PC-07, was filed but is not machine-checkable and never fires in replay.
Discussion & Limitations
The backtest rests on public sources and the authors' own segmentation
of the timeline into a hand-coded event stream — scoring a decision
point as bound is a judgement, not a measurement, and the tabletop and
pilot stages proposed in §3.5 have not been run. Notifying affected
third parties is the one failure v0.1 does not fully close: it
degrades to a warning rather than a hard rejection, which is why a
notification clock is proposed as v0.2 future work alongside a
commit-reveal ledger layer. DBYR is a control measure, not an
alignment measure — it does not fix corrupted reward signals or
degrading chain-of-thought monitorability, but it makes the decisions
taken around them visible. Signing keys generated by the reference
implementation's
--sign flag are ephemeral and for pilot use only.
BibTeX
@misc{acharya2026dbyr,
title = {Declare Before You Run: An Open Filing Standard for Frontier Training and Evaluation Runs},
author = {Acharya, Mann and Ojha, Archit and Sankara Raman, Gautam and Rajput, Karm},
year = {2026},
month = sep,
note = {Research conducted at the Apart Research AI Incident Response Sprint, September 2026},
howpublished = {\url{https://github.com/mach-12/dbyr}},
}