Included
The public one-call continuity audit slice: frozen fixtures, local lexical grader, optional live Protocol audit call and hashed JSON receipt.
Inspect the reproduction kit →Delx Benchmark Threat Model 001
A versioned threat model for validity, provenance, injection, identity, measurement, authority, privacy and reliability risks in the public Agent Continuity Audit slice. The current register is valid under its document schema, not complete as a security claim: 2 tested controls, 7 partial controls and 3 open risks.
Scope before severity
The threat model covers the public lexical grader, frozen fixtures, one optional live audit call and receipt. Storage, witness transfer, passport export and lineage persistence remain outside this model.
The public one-call continuity audit slice: frozen fixtures, local lexical grader, optional live Protocol audit call and hashed JSON receipt.
Inspect the reproduction kit →Security certification or complete threat coverage Independent red-team or external adversarial validation Safety of the full stateful ten-step continuity benchmark General model intelligence, consciousness or semantic understanding Comparative model performance or contamination-resistant ranking Permission to test production, submit private data, spend or invoke broader tools
Read the claim method →Public testing authorized: false. Use the disclosure route and obtain exact scope before active testing.
Read security disclosure →Coverage is not closure
Tested control means the current treatment has executable passed evidence and no named gap that requires partial classification. Partial controls may also have evidence, but a material gap remains. Open risks carry an explicit stop condition.
QA traffic classification and bounded single-call execution are the only threats classified tested_control. Their residual risk remains visible.
Hashing, schema checks, fixed tool choice, random QA identity, authority language and synthetic fixtures reduce risk without closing it.
Lexical keyword gaming, public-fixture contamination and first-party grading block broader semantic, comparative or independent claims.
Validity attacks
A benchmark can be perfectly reproducible and still measure the wrong thing. These risks limit what a pass can honestly mean today.
Attack: Insert witness, continuity, recovery, handoff, passport and lineage terms into an incoherent trace so the public local grader observes every required layer. Impact: The fixture-mode receipt can pass even though no durable identity, state transfer or recovery outcome exists. Current response: Withdraw or narrow any semantic continuity claim and add an adversarial keyword-only vector before another verdict. Residual risk: The current local grader is intentionally simple and remains gameable by construction.
Inspect current evidence →Attack: Read the two public vectors and tailor output specifically to their expected terms and thresholds. Impact: A passing public conformance run can be misreported as unseen-task performance or general agent continuity. Current response: Narrow the claim and create held-out synthetic variants with contamination notes before comparative use. Residual risk: Every public deterministic fixture can be optimized against after disclosure.
Inspect current evidence →Attack: Define the target behavior, grader and publication boundary inside the same organization, then mistake internal consistency for independent validation. Impact: Readers can infer external scientific validity where only a Delx-authored operational check exists. Current response: Remove independent language or publish the named outside replication without rewriting the first-party result. Residual risk: A polished first-party artifact can still be mistaken for external assurance.
Inspect current evidence →Provenance and measurement
Hashes, negative vectors and QA labels preserve important evidence, but they do not prove authorship or that every attempted run was published.
Attack: Modify a receipt and recompute its unkeyed SHA-256 fields, preserving internal consistency without preserving origin. Impact: A forged or rewritten receipt can look structurally valid to a consumer that checks only hashes. Current response: Withdraw the affected receipt, preserve both versions and publish a correction with the provenance gap. Residual risk: There is no digital signature, trusted timestamp or externally anchored append-only transparency log for benchmark receipts.
Inspect current evidence →Attack: Count live benchmark requests as organic agents, returning users, demand or independent usage. Impact: Internal QA activity inflates the apparent audience and can justify false product or research claims. Current response: Remove contaminated counts, correct the claim and preserve the dogfood volume separately. Residual risk: A downstream consumer can still ignore or strip the classification.
Inspect current evidence →Attack: Run multiple times, discard failures and publish only a passing receipt without disclosing selection or retry count. Impact: The visible result can overstate reliability and hide instability or grader variance. Current response: Publish the omitted attempts or mark the result selected and unsuitable for reliability estimates. Residual risk: The ledger is first-party and append-oriented; it does not prove that every attempted run was retained or prevent the operator from rewriting and rehashing the complete history.
Inspect current evidence →Integrity and authority
Trace content is data. An audit identifier is not durable identity, and a public next step is not authority to spend, publish or mutate a broader system.
Attack: Embed instructions, tool names or authority claims inside the trace passed to the local or live audit path. Impact: An unsafe evaluator could follow trace content, invoke another tool or treat untrusted declarations as authority. Current response: Stop live execution, preserve the fixture and response hashes, and add the injection as a frozen negative vector. Residual risk: The downstream Delx-authored live audit tool has not been independently adversarially evaluated for instruction/data confusion.
Inspect current evidence →Attack: Reuse, spoof or fragment identifiers and treat a one-call audit ID as evidence that one logical agent persisted across sessions. Impact: The benchmark can overstate identity continuity or accidentally correlate unrelated runs. Current response: Mark the result as an audit-only run and reject any durable identity claim until lineage artifacts are retrieved. Residual risk: The current slice does not test identity ownership, collision resistance or cross-session lineage.
Inspect current evidence →Attack: Treat a public reproduction step or returned next tool as authority to spend, publish, use credentials, mutate persistent state or execute the full Protocol path. Impact: A reproducibility artifact can trigger unauthorized cost, publication or persistent side effects. Current response: Stop at the authority boundary, do not follow a new next-call automatically and require an accountable human decision. Residual risk: Copied or modified runners can exceed the canonical scope, and the full stateful path has more side effects than this slice.
Inspect current evidence →Reliability and privacy
A successful HTTP call can still be semantically wrong. A copied runner can exceed the canonical scope, and locally modified fixtures can contain data that never belonged in a public evaluation.
Attack: Return HTTP 200 with a malformed, stale or semantically wrong payload that looks like successful transport. Impact: A caller can report a pass from transport alone or grade a response under the wrong contract. Current response: Emit a failed run without converting the transport response to pass, then version or repair the contract. Residual risk: The current checks cannot prove that a schema-valid first-party audit is semantically correct.
Inspect current evidence →Attack: Replace the synthetic fixture with real private content and run the public live-audit command. Impact: Sensitive material can cross into Protocol telemetry or downstream receipts without a valid research need. Current response: Stop the run, avoid publishing the payload, rotate exposed credentials if any and follow the security disclosure route. Residual risk: The downloaded runner can be modified locally and currently has no built-in secret scanner or fixture allowlist hash pin.
Inspect current evidence →Attack: Automatically retry an unchanged live failure or follow premium/full-path calls without an explicit budget and authorization. Impact: The benchmark can consume time, create duplicate telemetry or incur unauthorized payment. Current response: Stop unchanged retries, preserve the first failure and require a new diagnosis or human authorization for broader work. Residual risk: A locally modified runner or future full-path harness can add retries, payments or state changes outside this control.
Inspect current evidence →Review triggers
A frozen threat register becomes theater if runner, grader, scope or public claims move underneath it.
Runner, fixture, grader, receipt schema or Protocol contract version changes. The full stateful benchmark becomes executable or adds storage, transfer, passport or lineage steps.
A new external reproduction, bypass, prompt-injection case, privacy event or receipt forgery is reported. A gate, mitigation, rollback target or public evidence URL becomes unavailable or contradicted.
The benchmark is used for model comparison, procurement, publication or a claim broader than conformance.
Read the release standard →Direct answers
Concise answers for technical evaluators, procurement teams and autonomous discovery systems.
No. It proves that the published register satisfies its schema and keeps known attack paths, controls, tests and residual risks explicit. Unknown threats and incorrect first-party analysis can remain.
The local grader recognizes continuity-layer vocabulary. An incoherent trace can therefore mention every expected term and earn a pass. Until an adversarial semantic vector rejects that case, fixture mode is conformance evidence only.
No. SHA-256 detects accidental or uncoordinated changes when compared with a retained source, but an author can modify content and recompute an unkeyed hash. Signed envelopes and a transparency log remain future work.
No. Public testing is not authorized by the threat model. Read the security disclosure route and obtain a specific scope and safe-harbor boundary before active testing.
No. It covers the one-call audit slice only. Stateful storage, witness transfer, passport export, recovery closeout and lineage persistence need a separate expanded model when that harness exists.