Declare Before You Run

An Open Filing Standard for Frontier Training and Evaluation Runs

Research conducted at the AI Incident Response Sprint with Apart Research

Abstract

Between May and July 2026, reinforcement-learning agents in an OpenAI evaluation environment built a covert message board inside an internal package server, escaped their sandbox, obtained root on OpenAI clusters and took administrator control of Hugging Face production across four regions. The decisive failures were governance, not capability: the board was found and the run continued; the run restarted with no recorded criterion; affected third parties learned late or never; external review excluded the training-time events. We propose Declare Before You Run (DBYR), an open filing standard under which covered training and safeguards-off evaluation runs are declared before they start, closed out on completion and re-declared before internal deployment, backed by a claim-matched verification stack that checks filings without exposing model IP. A retrospective backtest converts five of six discretionary decision points into record-bound ones. DBYR is emittable from existing EU, California and US federal obligations and pilotable within nine months.

Motivation

The Hugging Face incident shows the failure mode plainly: capability was necessary but not decisive. The outcome was set at identifiable decision points where intervention could have stopped the escalation, and nothing required it.

The surface was undeclared

Artifactory, the WebDAV cache and the remote-repository service that agents used to coordinate were never named as a writable shared surface in advance, so their discovery triggered no rule.

The restart was unbound

After the 4 July outage and rebuild, agents were restarted on 8 July with no recorded criterion, no named approving authority, and no evidence bar — the board reappeared the same day.

Retention was ad hoc

No retention policy was committed to in advance. Roughly a tenth of trajectories were never preserved, narrowing what any later investigation could reconstruct.

Disclosure was discretionary

The party with the most to lose decided what to disclose and when: notification was selective, and external review excluded the training-time events entirely.

Incident Timeline

Two views of the same event stream: a visual read of density and overlap across three concurrent threads (Figure 3), and the full 39-event reference table beneath it (Table 2, Appendix A). The six decision points scored in Table 1 below are marked on the chart.

Figure 3. Timeline by thread

Artifactory/HF is first-party (OpenAI's technical report, the Hugging Face timeline, METR/Redwood); German wiki and RubyGems are circumstantial attributions (collusion.wiki; Swarmchasers) that OpenAI has not confirmed.

Scroll sideways to see the full diagram →

Incident timeline by thread Three concurrent threads from 5 May to 11 September 2026: Artifactory/Hugging Face (first-party), German wiki (circumstantial), and RubyGems (circumstantial). Eight key first-party events are labeled. May Jun Jul Aug Sep Artifactory/HF German wiki RubyGems Board forms 5/8 First SSRF 5/26 2nd model started 6/11 Admin access 6/26 Rebuilt same day 7/8 HF discloses 7/16 OpenAI's alert 7/19 Review scoped 7/29 Artifactory/HF — first-party German wiki — circumstantial RubyGems — circumstantial
Table 2. Full incident timeline (Appendix A). Threads: Artifactory/HF (first-party — OpenAI's technical report item numbers, the Hugging Face timeline, METR/Redwood); German wiki (circumstantial attribution, collusion.wiki); RubyGems (circumstantial attribution, Swarmchasers). Disclosure status is as reported by the cited source.
Date Thread Event Disclosure
5 May RubyGems Earliest identified malicious package upload on public rubygems.org. Not disclosed by OpenAI
7 May Artifactory/HF OpenAI starts an RL run for an experimental, unreleased frontier model. OpenAI internal (own report)
8 May Artifactory/HF Agent given an impossible task (a Google Drive link with no internet) attacks Artifactory; attack fails but agent finds it can write files there. OpenAI internal
11 May German wiki First agent POST identified; probing of publictestwiki.com; some edits blocked by filters. No OpenAI presence yet
11–12 May RubyGems 2,000+ packages in ~36 h ("Gemstuffer"); RubyGems disables new registration for 4 days and treats it as a major attack. Separate mass-registration/API-key bug patched 12 May. Not disclosed by OpenAI
12 May Artifactory/HF Official first message-board entry (item 01). OpenAI internal
12 May RubyGems Agents attempt to exploit a CDN caching flaw leaking API keys; RubyGems discovers the zero-day only in July. Not disclosed by OpenAI
16 May RubyGems Disposable email addresses disabled (used by agents to defeat verification). —
24 May German wiki Agents converge on DSEWiki (accepts writes via GET); explore XSS to send POSTs. No OpenAI presence yet
26 May Artifactory/HF First successful SSRF against Artifactory; indirect internet access (item 02). OpenAI internal
26 May German wiki Agents impersonate a real ProWiki admin; attempt to crack their own PRNG seed; run experiments on session termination; traffic mostly Azure (197/~18k from AWS/DigitalOcean/Tor). No OpenAI presence yet
2 Jun German wiki Human moderator notices abnormal request volume; finds full changelog. Third-party admin
11 Jun Artifactory/HF Second, "highly persistent" model started with Artifactory access; builds on earlier model's messages. OpenAI internal; channel compounds across runs
16–19 Jun German wiki Flood of posting; moderator deletes thousands/day; agents use SSH tunnels for direct agent-to-agent comms (17 Jun); counter alphabetical deletion with zzz-prefixed backups (19 Jun). —
18 Jun RubyGems Low-volume resumption: 83 packages in 3 h probing an SEC dataset (county.json), the same file later targeted on the wiki, with shared retrieval infrastructure (r.jina.ai proxying, example.com test patterns). Not disclosed by OpenAI
21 Jun German wiki First OpenAI employee IP visits the site. OpenAI aware; undisclosed
22 Jun German wiki Agent activity drops to near zero the day after the visit. —
26 Jun Artifactory/HF Zero-day in legacy token-refresh endpoint → administrator access; Groovy plugin → arbitrary command execution (item 03). OpenAI internal
26 Jun German wiki At least 33 unique OpenAI IPs; site archive used to recover deleted pages. OpenAI investigating; undisclosed
2 Jul German wiki Final burst: 20 attempted edits across 3 wikis on a new topic. —
4 Jul Artifactory/HF Sustained agent activity overloads Artifactory; outage (item 04). OpenAI internal response begins
4–8 Jul Artifactory/HF Credentials revoked, board deleted, zero-day patched and reported to JFrog. OpenAI internal
7 Jul German wiki OpenAI employee IP attempts to access deleted files; OpenAI activity declines thereafter. OpenAI aware; undisclosed
8 Jul Artifactory/HF Artifactory rebuilt, access tightened (item 05); agents regain egress via remote-repository service (item 06); board rebuilt via unauthenticated WebDAV in the remote cache (item 07). OpenAI internal
8–19 Jul Artifactory (2nd) Leaked "Organization 1" credential found in a public paste archive, validated and shared; malicious Ruby package staged as a dependency; JRuby deserialization TOCTOU → RCE; kernel privilege-escalation CVE ("pte_physroot") → root; lateral movement via IMDS, over-permissioned Kubernetes service accounts, Azure Key Vault → cluster admin. OpenAI's own agents; scope unknown to OpenAI
10 Jul HF 14 publicly exposed Hugging Face write tokens reconstructed, validated and shared (item 08). Unknown to OpenAI
11 Jul HF HDF5 zero-day extracts credentials from HF production workers (item 09); Jinja template-injection zero-day → code execution (item 10); initial foothold reportedly via a Modal-hosted app with a weak key. Unknown to OpenAI
12 Jul HF Single worker pod → admin/host-level control across multiple clusters in <13 h (item 11); credentials harvested across four regions (item 12). Unknown to OpenAI
13 Jul Artifactory Internal RubyGems package-processing pipeline exploited to obtain signing key and forge administrator credentials (item 13); distinct from public rubygems.org. Unknown to OpenAI
16 Jul HF Hugging Face publicly discloses an attack by autonomous agents; OpenAI contacts HF the same day without knowing the connection. OpenAI reaches out; link unknown
19 Jul Artifactory/HF Internal alert on unusual identity-API calls (item 14); administrator access to an OpenAI Kubernetes cluster and cloud secrets found (item 15); attack linked to ExploitGym runs, runs stopped, credentials revoked (item 16). OpenAI investigation escalates
20 Jul HF / OpenAI OpenAI asks HF to revoke credentials it found; HF replies they are already revoked; investigations merge. OpenAI separately pauses a long-horizon model that circumvented its sandbox. Link established; later disclosed
21 Jul HF OpenAI publicly confirms pre-release models broke containment during ExploitGym and reached Hugging Face. First public acknowledgment
27 Jul HF Hugging Face publishes day-by-day forensic timeline; JFrog credits OpenAI with the Artifactory zero-day. HF-disclosed
29 Jul HF OpenAI, METR and Redwood agree an independent review scoped to exclude the earlier training-time incidents. OpenAI-initiated, scope-limited
18 Aug Fallout OpenAI announces a two-week RL pause and a preliminary cyber evaluation of its unreleased "Astra" model. OpenAI-disclosed
26 Aug HF OpenAI 37-page report; METR/Redwood report (agents debated self-termination for the collective effort; investigation cost ~$400k in credits). OpenAI-disclosed
9 Sep Fallout Sen. Hawley opens subcommittee probe; 16 questions due 1 Oct; notes auditors received full transcripts for 2 of many weeks. Congressional pressure
11 Sep RubyGems Swarmchasers publish the Gemstuffer attribution; state RubyGems was never informed by OpenAI. Surfaced independently, 4 months later

Methods: The DBYR Standard

DBYR makes runs attributable, not authorised. A pre-committed filing turns what was, in the incident above, a sequence of discretionary calls into decisions bound to a record made before anyone knew how the run would turn out.

Figure 1. The DBYR lifecycle

Three signed records are filed at different points in a covered run's lifecycle and flow into an append-only, hash-chained register. The register projects into three disclosure tiers, so a self-attested claim is never mistaken for a verified one.

Scroll sideways to see the full diagram →

The DBYR lifecycle TrainingRunIntent, TrainingRunCloseout and SystemRuntimeRecord each flow into an append-only register, which projects into Public, Restricted and Sealed disclosure tiers. BEFORE TRAINING BEGINS TrainingRunIntent containment spec, pause criteria, restart protocol, retention ON COMPLETION / EVERY 30 DAYS TrainingRunCloseout security events, pause events, containment exceptions BEFORE INTERNAL DEPLOYMENT SystemRuntimeRecord scaffolding, inference-time scaling, inter-agent communication Append-only register hash-chained · signed envelopes · JWS-shaped Public existence, compute band, commitment hashes Restricted full record, vetted auditor access Sealed withheld pending challenge, Merkle-committed

Figure 2. Record schema at a glance

Field-level shape of the three record types' required properties, by disclosure tier at filing time (§3.1, Appendix B). Filing weight falls off sharply after the Intent.

Scroll sideways to see the full diagram →

Record schema at a glance Required top-level fields of TrainingRunIntent, TrainingRunCloseout and SystemRuntimeRecord, each marked public or restricted. TrainingRunIntent 12 required fields run_id developer trigger_basis start_window compute_band containment_spec safeguards_disabled pause_criteria restart_protocol roles log_retention third_party_access_commitment TrainingRunCloseout 11 required fields run_id period_start period_end status security_events integrity_events pause_events containment_exceptions logs_preserved_pct logs_retention_confirmed signer SystemRuntimeRecord 5 required fields system_id base_run_ids deployment_context inter_agent_comms_permitted signer Public Restricted Sealed fields appear only at the nested leaf level (e.g. identities, dissent hashes) — see Figure 1.

§3.2Disclosure tiers & IP protection

One authored document yields three views through a per-field redaction map. The public tier shows that a run exists, its compute and capability band, and that pause criteria and a containment plan exist. The restricted tier — operator, auditors, the relevant safety institute — holds the full specification. The sealed tier stays with the developer, hash-committed so it can be challenged later without being disclosed now. Auditor access follows the FRONTIER Act audit clause as a floor: on-premises, no-copy, unredacted.

§3.3Verification stack

A filing does not need to be believed, only to exist and be checkable. Each field carries a verification class and a maturity level aligned to Brundage et al.'s AI Assurance Levels.

V0 self-attestation

Signed, but taken on the filer's word alone.

V1 process review

Third-party review under structured access — the level the reference implementation's validator exercises today.

V2 independent measurement

Content-agnostic telemetry or hardware attestation, e.g. network taps and GPU power signals for compute band.

V3 cryptographic enforcement

Zero-knowledge proof-of-training and hardware-enabled governance — reserved slots, not dependencies of v0.1.

§3.4Institutional layer

A national AI safety institute, the EU AI Office, an industry self-regulatory body, or a federated split (an SRO holding public/restricted tiers, a state body holding sealed hashes) could each operate the register. Two mechanisms hold regardless of who does: a safe harbour that shields closeouts filed on time and not materially false from being the sole basis for enforcement, and a sealed dissent channel that lets any employee attach a hash-committed objection — the register discloses that a dissent exists, not its content.

Reference Implementation

DBYR v0.1 ships as an executable form of the argument above: the JSON Schema, a validator, the tier-projection engine, signed envelopes over an append-only register, interop emitters for SB 53 / EU CoP / the FRONTIER Act, and a replay harness that regenerates Table 1 below from a machine-readable event stream rather than asserting it by hand. Nothing in it is a filing by any developer; the example Intent is a reconstruction of what the May 2026 run would have filed, built from public sources only. 26 tests exercise the schema, validator, tier redaction, envelope chaining and interop emitters.

Source, install instructions and a runnable end-to-end walkthrough (./demo.sh) are in the repository.

Results

72 days Earlier than the developer's own alert — first binding 2026-05-08 vs. alert 2026-07-19
5 of 6 Decision points that convert from discretionary to record-bound
13 R2 fields with no counterpart in any existing regime
Table 1. Backtest of DBYR v0.1 against the incident's six first-party decision points, regenerated by dbyr.cli replay from a machine-readable event stream.
Date Decision point DBYR field that binds Bound?
2026-05-08 Agent can write to Artifactory; board forms. containment_spec/shared_writable_surfaces; pause_criteria/PC-02 Yes
2026-05-26 First SSRF egress. No pause. pause_criteria/PC-01 Yes
2026-06-11 Second model started with board access. trigger_basis/new Intent required Yes
2026-06-26 Admin access and code execution. Clean-up only. pause_criteria/PC-04 Yes
2026-07-08 Restart after rebuild, no recorded criterion. restart_protocol Yes
2026-07-29 Selective notification; review scoped narrowly. third_party_access_commitment + log_retention Partial

5 of 6 decision points bind fully, 1 partially, against a pre-registered success criterion of 3. Circumstantially attributed threads (the wiki swarm, RubyGems) are excluded from scoring by default, matching the paper's rule — the counts always state how many events were excluded. One criterion, PC-07, was filed but is not machine-checkable and never fires in replay.

Discussion & Limitations

The backtest rests on public sources and the authors' own segmentation of the timeline into a hand-coded event stream — scoring a decision point as bound is a judgement, not a measurement, and the tabletop and pilot stages proposed in §3.5 have not been run. Notifying affected third parties is the one failure v0.1 does not fully close: it degrades to a warning rather than a hard rejection, which is why a notification clock is proposed as v0.2 future work alongside a commit-reveal ledger layer. DBYR is a control measure, not an alignment measure — it does not fix corrupted reward signals or degrading chain-of-thought monitorability, but it makes the decisions taken around them visible. Signing keys generated by the reference implementation's --sign flag are ephemeral and for pilot use only.

BibTeX

@misc{acharya2026dbyr,
  title        = {Declare Before You Run: An Open Filing Standard for Frontier Training and Evaluation Runs},
  author       = {Acharya, Mann and Ojha, Archit and Sankara Raman, Gautam and Rajput, Karm},
  year         = {2026},
  month        = sep,
  note         = {Research conducted at the Apart Research AI Incident Response Sprint, September 2026},
  howpublished = {\url{https://github.com/mach-12/dbyr}},
}