Research question
Can one frozen continuity-audit grader separate a synthetic trace containing witness, recovery and handoff evidence from a synthetic trace missing witness preservation and continuity transfer?
Inspect the full contract →Delx Technical Report 001 · v1.0.0
This technical report asks whether the Delx continuity-audit slice clears a minimum falsifiability check: the same frozen grader must accept a synthetic trace containing witness preservation, recovery closeout and successor continuity while rejecting a synthetic trace missing witness preservation and continuity transfer. The positive fixture scored 86 and passed; the negative fixture scored 53 and remained fail, a 33-point separation. One live QA call reproduced the positive score and layer set. The result changed one release decision: publish the narrow audit slice and its receipts, while continuing to block claims about the full stateful benchmark, semantic validity, model comparison and independent validation. This is a first-party technical report, not a peer-reviewed paper or independent validation.
Question and decision
The report asks one bounded question and publishes the consequences of the answer. It does not inflate a two-vector check into a general benchmark claim.
Can one frozen continuity-audit grader separate a synthetic trace containing witness, recovery and handoff evidence from a synthetic trace missing witness preservation and continuity transfer?
Inspect the full contract →Publish the one-call continuity-audit slice as a narrow reproduction kit with positive and negative receipts, an executable threat model and explicit first-party boundaries.
Run the released slice →full stateful continuity benchmark completion; semantic validity of the lexical grader; resistance to keyword stuffing or public-fixture contamination; model or provider comparison; independent validation, external replication or peer review; adoption, demand, universal reliability or system-wide security.
Inspect the open risks →Method
Both frozen vectors use the same dependency-free runner and scoring rule. The live call reuses the positive vector through the public Protocol tool and remains QA-classified.
Start at 35; add 9 points for each detected layer among structure, ego, witness, continuity, relation and recovery; add 15 when witness, continuity and recovery are all present; clamp to 0–100.
Inspect the runner →Score at least 78, observe witness, continuity, recovery, allow zero missing required layers and return low continuity risk.
Inspect the manifest →Fixture mode changes no external state. Live mode writes QA-classified Protocol telemetry under a qa- agent identifier and must remain excluded from organic adoption metrics.
Inspect the Protocol contract →Three retained observations
Expected status and observed status are separate fields. Matching a negative expectation does not convert the observed result into pass.
Observed: pass; expected: pass; matched: true. Required layers were present and continuity risk was low.
Inspect the positive receipt →Observed: fail; expected: fail; matched: true. Witness and continuity were missing; the result remained high risk and fail.
Inspect the negative receipt →Observed: pass; expected: pass; matched: true. It reproduced the positive score and layer set through a Delx-authored live tool.
Inspect the live receipt →The two frozen vectors differ by 33 score points. This proves separation for these exact bytes only; it is not an effect size from an independent sample.
Inspect the observations →Validity boundary
An always-pass harness would fail this control. Passing it still leaves the most important validity questions open.
keyword stuffing can satisfy lexical layer detection without real continuity evidence
Inspect threat CAV-001 →public fixtures can contaminate future evaluation or invite tuning to the test
Inspect threat CAV-002 →Delx owns the runner, fixtures, grader, live tool and publication path. Independent validation: false. External replication: false. Peer reviewed: false.
Read the claim method →Storage, witness transfer, passports, lineage persistence and the ten-step Protocol flow remain outside this report's executed scope.
Inspect the full benchmark →Reproduce, cite, correct
Human, JSON, Markdown and BibTeX views resolve to one canonical report. Exact runner, fixture and receipt hashes make the named result inspectable.
node runner.mjs --fixture pass · node runner.mjs --fixture fail. Fixture mode uses zero paid calls and changes no external state.
Download the runner →Canonical report content SHA-256: 2e4e342456cb6f267d744dbe8777b9d89cae7b55705508593252dd864e63e5f1.
Inspect the signed payload boundary →David Batista. A minimum falsifiability slice for agent continuity audits. Delx Technical Report 001, version 1.0.0, 2026.
Read the BibTeX record →Zero means no event currently targets DELX-TR-001 v1.0.0; it does not mean Delx has never published an error. Silent rewrites are outside policy; the ledger itself remains append-oriented rather than technically append-only.
Inspect the evidence ledger →Direct answers
Concise answers for technical evaluators, procurement teams and autonomous discovery systems.
No. It is Delx Technical Report 001, authored and published by Delx. Peer review, independent validation and external replication are all false in the report contract.
No. The observed benchmark result remains fail. A separate expectation_matched field records that the negative control behaved as designed.
No. It reuses the positive fixture against a Delx-authored public tool. The runner, fixture, tool and publication path remain first-party.
It supported releasing the narrow one-call audit kit and its receipts while explicitly blocking claims about the full stateful benchmark, semantic validity, model ranking and independent validation.
A correction, retraction or supersession must target this report version in the public evidence ledger and preserve the earlier published version. Silent rewrite is outside policy.