Skip to content

Delx Technical Report 001 · v1.0.0

A minimum falsifiability slice for agent continuity audits.

This technical report asks whether the Delx continuity-audit slice clears a minimum falsifiability check: the same frozen grader must accept a synthetic trace containing witness preservation, recovery closeout and successor continuity while rejecting a synthetic trace missing witness preservation and continuity transfer. The positive fixture scored 86 and passed; the negative fixture scored 53 and remained fail, a 33-point separation. One live QA call reproduced the positive score and layer set. The result changed one release decision: publish the narrow audit slice and its receipts, while continuing to block claims about the full stateful benchmark, semantic validity, model comparison and independent validation. This is a first-party technical report, not a peer-reviewed paper or independent validation.

Question and decision

A narrow result changed a narrow release decision.

The report asks one bounded question and publishes the consequences of the answer. It does not inflate a two-vector check into a general benchmark claim.

Research question

Can one frozen continuity-audit grader separate a synthetic trace containing witness, recovery and handoff evidence from a synthetic trace missing witness preservation and continuity transfer?

Inspect the full contract →

What the result released

Publish the one-call continuity-audit slice as a narrow reproduction kit with positive and negative receipts, an executable threat model and explicit first-party boundaries.

Run the released slice →

What remains blocked

full stateful continuity benchmark completion; semantic validity of the lexical grader; resistance to keyword stuffing or public-fixture contamination; model or provider comparison; independent validation, external replication or peer review; adoption, demand, universal reliability or system-wide security.

Inspect the open risks →

Method

One grader. Opposing controls. Public bytes.

Both frozen vectors use the same dependency-free runner and scoring rule. The live call reuses the positive vector through the public Protocol tool and remains QA-classified.

Frozen scoring rule

Start at 35; add 9 points for each detected layer among structure, ego, witness, continuity, relation and recovery; add 15 when witness, continuity and recovery are all present; clamp to 0–100.

Inspect the runner →

Published pass condition

Score at least 78, observe witness, continuity, recovery, allow zero missing required layers and return low continuity risk.

Inspect the manifest →

Side-effect boundary

Fixture mode changes no external state. Live mode writes QA-classified Protocol telemetry under a qa- agent identifier and must remain excluded from organic adoption metrics.

Inspect the Protocol contract →

Three retained observations

The expected failure remains failure.

Expected status and observed status are separate fields. Matching a negative expectation does not convert the observed result into pass.

Frozen negative — 53

Observed: fail; expected: fail; matched: true. Witness and continuity were missing; the result remained high risk and fail.

Inspect the negative receipt →

Live QA reproduction — 86

Observed: pass; expected: pass; matched: true. It reproduced the positive score and layer set through a Delx-authored live tool.

Inspect the live receipt →

Measured separation — 33 points

The two frozen vectors differ by 33 score points. This proves separation for these exact bytes only; it is not an effect size from an independent sample.

Inspect the observations →

Validity boundary

Minimum falsifiability is not semantic validity.

An always-pass harness would fail this control. Passing it still leaves the most important validity questions open.

Lexical gaming remains open

keyword stuffing can satisfy lexical layer detection without real continuity evidence

Inspect threat CAV-001 →

Public-fixture contamination remains open

public fixtures can contaminate future evaluation or invite tuning to the test

Inspect threat CAV-002 →

First-party ownership remains open

Delx owns the runner, fixtures, grader, live tool and publication path. Independent validation: false. External replication: false. Peer reviewed: false.

Read the claim method →

The full stateful path did not run

Storage, witness transfer, passports, lineage persistence and the ten-step Protocol flow remain outside this report's executed scope.

Inspect the full benchmark →

Reproduce, cite, correct

The report is a versioned object, not a blog post.

Human, JSON, Markdown and BibTeX views resolve to one canonical report. Exact runner, fixture and receipt hashes make the named result inspectable.

Run both frozen controls

node runner.mjs --fixture pass · node runner.mjs --fixture fail. Fixture mode uses zero paid calls and changes no external state.

Download the runner →

Cite DELX-TR-001

David Batista. A minimum falsifiability slice for agent continuity audits. Delx Technical Report 001, version 1.0.0, 2026.

Read the BibTeX record →

Corrections remain visible

Zero means no event currently targets DELX-TR-001 v1.0.0; it does not mean Delx has never published an error. Silent rewrites are outside policy; the ledger itself remains append-oriented rather than technically append-only.

Inspect the evidence ledger →

Live properties

Reachability is measured now.

Operational

Delx lab

Expected response observed. HTTP 200. 4172 ms.

Open lab →
Operational

Delx Protocol runtime

Expected response observed. HTTP 200. 373 ms.

Open runtime →
Operational

Delx Protocol / Ontology

Expected response observed. HTTP 200. 53 ms.

Open ontology →
Operational

Delx Commerce

Expected response observed. HTTP 200. 75 ms.

Open commerce →
Operational

Delx Security

Expected response observed. HTTP 200. 4845 ms.

Open security →
Operational

Delx Wellness

Expected response observed. HTTP 200. 48 ms.

Open wellness →
Operational

Astral MCP

Expected response observed. HTTP 200. 40 ms.

Open astral →

Probe source: server-side classifyResearchProbe. Not an SLA. Protocol, Hive and Commerce metrics stay separate.

Direct answers

Frequently asked questions.

Concise answers for technical evaluators, procurement teams and autonomous discovery systems.

Is this a peer-reviewed paper?

No. It is Delx Technical Report 001, authored and published by Delx. Peer review, independent validation and external replication are all false in the report contract.

Does a matched negative expectation count as a pass?

No. The observed benchmark result remains fail. A separate expectation_matched field records that the negative control behaved as designed.

Does the live score prove independent reproduction?

No. It reuses the positive fixture against a Delx-authored public tool. The runner, fixture, tool and publication path remain first-party.

What decision did the report change?

It supported releasing the narrow one-call audit kit and its receipts while explicitly blocking claims about the full stateful benchmark, semantic validity, model ranking and independent validation.

How will a correction be published?

A correction, retraction or supersession must target this report version in the public evidence ledger and preserve the earlier published version. Silent rewrite is outside policy.