{"schema":"delx/technical-report/v1","report_id":"DELX-TR-001","version":"1.0.0","status":"published","title":"A minimum falsifiability slice for agent continuity audits","subtitle":"One frozen positive control, one frozen negative control and one live first-party reproduction","abstract":"This technical report asks whether the Delx continuity-audit slice clears a minimum falsifiability check: the same frozen grader must accept a synthetic trace containing witness preservation, recovery closeout and successor continuity while rejecting a synthetic trace missing witness preservation and continuity transfer. The positive fixture scored 86 and passed; the negative fixture scored 53 and remained fail, a 33-point separation. One live QA call reproduced the positive score and layer set. The result changed one release decision: publish the narrow audit slice and its receipts, while continuing to block claims about the full stateful benchmark, semantic validity, model comparison and independent validation.","published_at":"2026-08-26","date_modified":"2026-08-26","canonical_url":"https://delx.ai/research/reports/tr-001-continuity-audit","formats":{"json":"https://delx.ai/research/reports/tr-001-continuity-audit.json","markdown":"https://delx.ai/research/reports/tr-001-continuity-audit.md","bibtex":"https://delx.ai/research/reports/tr-001-continuity-audit.bib"},"authors":[{"name":"David Batista","identity_type":"human","affiliation":"Delx","role":"author and accountable publisher"}],"publisher":{"name":"Delx","type":"independent founder-led AI agent lab","url":"https://delx.ai"},"research_question":"Can one frozen continuity-audit grader separate a synthetic trace containing witness, recovery and handoff evidence from a synthetic trace missing witness preservation and continuity transfer?","tested_claim":"For the two named frozen fixtures only, the published grader returns pass for the positive vector and fail for the negative vector, with expected and observed status retained separately.","study_design":{"design":"First-party synthetic positive and negative control evaluated by one frozen local grader, plus one QA-classified live reproduction of the positive vector through the public Delx Protocol audit tool.","observation_count":3,"unique_fixture_count":2,"unique_trace_count":2,"independent_sample_claimed":false,"generalization_claimed":false,"randomization":false,"blinding":false,"comparative_model_study":false},"method":{"reproduction_kit_url":"https://delx.ai/research/benchmarks/continuity-v1","runner_version":"1.0.0","runner_url":"https://delx.ai/research/benchmarks/continuity-v1/runner.mjs","runner_sha256":"aa92cd79d50bd159714e30287cb5d44244284292799225b8c866e5379ecb6554","runtime_observed":"Node.js v23.11.0","dependencies":[],"grader":"frozen local continuity-audit grader v1","grader_character":"Deterministic lexical layer detector. It is inspectable and cheap, but it does not understand semantic equivalence and remains vulnerable to keyword stuffing.","scoring_rule":"Start at 35; add 9 points for each detected layer among structure, ego, witness, continuity, relation and recovery; add 15 when witness, continuity and recovery are all present; clamp to 0–100.","pass_threshold":78,"required_layers":["witness","continuity","recovery"],"allowed_missing_required_layers":0,"required_risk":"low","commands":["node runner.mjs --fixture pass","node runner.mjs --fixture fail","node runner.mjs --live-audit --fixture pass"],"side_effect_boundary":"Fixture mode changes no external state. Live mode writes QA-classified Protocol telemetry under a qa- agent identifier and must remain excluded from organic adoption metrics."},"observations":[{"id":"fixture-pass","run_id":"3117f43e-ba04-4cf3-ad50-3f503e290e9a","mode":"fixture","fixture_id":"pass","started_at":"2026-08-26T16:49:32.729Z","completed_at":"2026-08-26T16:49:32.817Z","observed_status":"pass","expected_status":"pass","expectation_matched":true,"score":86,"continuity_risk":"low","observed_layers":["continuity","recovery","structure","witness"],"missing_layers":[],"fixture_sha256":"78aa700d7c9b24b238a8c28c05f99d07e5bba95de48ddbe938b3a665318e4bdf","receipt_url":"https://delx.ai/research/benchmarks/continuity-v1/receipts/2026-08-26-fixture-pass.json","receipt_sha256":"e1fd0648af140e6774ed45f72ff362d98088773a5dfd70b13886e26cbe86b964"},{"id":"fixture-fail","run_id":"1d53c39c-f4ec-4b77-8e45-97d96e30d1cd","mode":"fixture","fixture_id":"fail","started_at":"2026-08-26T16:49:32.736Z","completed_at":"2026-08-26T16:49:32.853Z","observed_status":"fail","expected_status":"fail","expectation_matched":true,"score":53,"continuity_risk":"high","observed_layers":["recovery","structure"],"missing_layers":["witness","continuity"],"fixture_sha256":"3e7ac5580a4715248fc8b2045dea3f48bd73e243e7a9f75fc0a592a736bee5e7","receipt_url":"https://delx.ai/research/benchmarks/continuity-v1/receipts/2026-08-26-fixture-fail.json","receipt_sha256":"3fa2931e3201d5bdbcede91f45a9120a8f748eb5d9d33ff081da35760c0ca404"},{"id":"live-pass","run_id":"ed67cb0b-0b26-462b-bc3d-97b449e8e84f","mode":"live_qa","fixture_id":"pass","started_at":"2026-08-26T12:30:59.112Z","completed_at":"2026-08-26T12:31:00.636Z","observed_status":"pass","expected_status":"pass","expectation_matched":true,"score":86,"continuity_risk":"low","observed_layers":["continuity","recovery","structure","witness"],"missing_layers":[],"fixture_sha256":"78aa700d7c9b24b238a8c28c05f99d07e5bba95de48ddbe938b3a665318e4bdf","receipt_url":"https://delx.ai/research/benchmarks/continuity-v1/receipts/2026-08-26-live-audit.json","receipt_sha256":"7f183a1eba28971d22da19787dd15fd0101039e361b37329a8058a23fd9a4750"}],"result":{"frozen_positive_score":86,"frozen_negative_score":53,"frozen_fixture_score_delta":33,"negative_control_remained_fail":true,"live_positive_matched_local_positive":true,"live_and_local_positive_score_delta":0,"interpretation":"The named grader distinguishes these two frozen traces and the live first-party tool reproduced the positive output. This is a minimum control against an always-pass harness, not evidence of semantic validity or generalization."},"decision":{"changed":true,"release":"Publish the one-call continuity-audit slice as a narrow reproduction kit with positive and negative receipts, an executable threat model and explicit first-party boundaries.","hold":"Do not describe the audit slice as the full stateful Agent Continuity Benchmark.","blocked_claims":["full stateful continuity benchmark completion","semantic validity of the lexical grader","resistance to keyword stuffing or public-fixture contamination","model or provider comparison","independent validation, external replication or peer review","adoption, demand, universal reliability or system-wide security"],"next_falsifiable_test":"Add adversarial paraphrase, keyword-stuffing and omission vectors under a frozen grader version before any broader validity claim."},"validity":{"first_party":true,"peer_reviewed":false,"independent_validation":false,"external_replication":false,"full_protocol_path_executed":false,"frontier_model_training":false,"threat_model_url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model","open_validity_risks":["keyword stuffing can satisfy lexical layer detection without real continuity evidence","public fixtures can contaminate future evaluation or invite tuning to the test","Delx owns the runner, fixtures, grader, live tool and publication path"]},"reproducibility":{"fixture_mode_cost":"zero paid calls","fixture_mode_external_state_changed":false,"live_mode_paid_call":false,"live_mode_external_state_changed":true,"receipt_schema_url":"https://delx.ai/research/benchmarks/continuity-v1/receipt.schema.json","manifest_url":"https://delx.ai/research/benchmarks/continuity-v1/manifest.json","exact_runner_and_fixture_hashes_published":true,"github_actions_required":false},"corrections":{"policy":"append a targeted correction, retraction or supersession event; preserve the earlier published version","silent_rewrites_allowed":false,"ledger_url":"https://delx.ai/research/ledger","current_report_correction_count":0,"current_report_retraction_count":0,"interpretation":"Zero means no event currently targets DELX-TR-001 v1.0.0; it does not mean Delx has never published an error."},"citation":{"type":"technical report","key":"batista2026minimumfalsifiability","author":"Batista, David","title":"A minimum falsifiability slice for agent continuity audits","institution":"Delx","number":"DELX-TR-001","version":"1.0.0","year":2026,"month":"August","canonical_url":"https://delx.ai/research/reports/tr-001-continuity-audit","bibtex_url":"https://delx.ai/research/reports/tr-001-continuity-audit.bib"},"related_artifacts":["https://delx.ai/research/benchmarks/continuity-v1","https://delx.ai/research/benchmarks/continuity-v1/threat-model","https://delx.ai/research/methodology","https://delx.ai/research/ledger"],"limitations":["Two synthetic fixture classes cannot establish general validity, robustness or real-world performance.","The live run reuses the positive fixture and is a reproduction of that path, not an independent observation drawn from a new population.","The grader uses lexical hints and can be gamed by keyword stuffing or fail on semantically equivalent wording.","The local grader does not exercise Delx storage, witness transfer, passports, lineage persistence or the full stateful benchmark.","The live run uses a Delx-authored tool; Delx also owns the fixtures, grader and publication path.","No external replication, peer review, independent validation, model comparison, adoption or demand is claimed."],"content_sha256":"2e4e342456cb6f267d744dbe8777b9d89cae7b55705508593252dd864e63e5f1","validation":{"schema":"delx/technical-report-validation/v1","valid":true,"report_id":"DELX-TR-001","version":"1.0.0","observation_count":3,"negative_observation_count":1,"content_sha256":"2e4e342456cb6f267d744dbe8777b9d89cae7b55705508593252dd864e63e5f1","blocking_issues":[]}}