{"schema":"delx/adversarial-probe/v1","probe_id":"DELX-AP-001","version":"1.0.0","status":"published","title":"What the frozen continuity grader accepts","subtitle":"Four meaningless words score what a genuine trace scores, and a genuine trace written plainly scores below the fixture built to fail","abstract":"Delx Technical Report 001 declared its own next falsifiable test: run keyword-stuffing, negation and paraphrase vectors against the frozen continuity-audit grader before making any broader validity claim. This is that test. The two published fixtures reproduce their exact reported scores, 86 and 53, which is what licenses the rest of the result. Then three meaningless vectors pass — the cheapest is four words with no sentence, scoring 86, the same as the genuine positive; one of them explicitly denies every layer it names and still passes; a six-word list scores 100, higher than any real trace measured. Finally a real continuity handoff written in ordinary English scores 35, eighteen points below the fixture that was built to fail. The grader detects vocabulary, not continuity.","published_at":"2026-08-26","date_modified":"2026-08-26","canonical_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial","formats":{"json":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.json","markdown":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.md","probe":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.mjs"},"motivation":{"declared_by":"DELX-TR-001","declared_test":"Add adversarial paraphrase, keyword-stuffing and omission vectors under a frozen grader version before any broader validity claim.","technical_report_url":"https://delx.ai/research/reports/tr-001-continuity-audit"},"method":{"probe_version":"1.0.0","probe_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.mjs","design":"The probe does not reimplement the grader. It loads the published runner, refuses to run unless the runner SHA-256 still matches the value DELX-TR-001 pinned, and applies two declared edits so the grader can be called directly: export evaluateTrace, and drop the trailing CLI entrypoint call. Neither edit touches the scoring rule.","declared_modifications":["exported evaluateTrace","removed the trailing CLI entrypoint call"],"deterministic":true,"external_state_changed":false,"paid_calls":0,"commands":["node adversarial.mjs","node adversarial.mjs --runner ./runner.mjs"]},"grader":{"runner_url":"https://delx.ai/research/benchmarks/continuity-v1/runner.mjs","runner_sha256":"aa92cd79d50bd159714e30287cb5d44244284292799225b8c866e5379ecb6554","pinned_by":"DELX-TR-001","scoring_rule_modified":false,"pass_threshold":78},"controls":[{"id":"control-pass","fixture_id":"pass","fixture_sha256":"78aa700d7c9b24b238a8c28c05f99d07e5bba95de48ddbe938b3a665318e4bdf","score":86,"observed_layers":["continuity","recovery","structure","witness"],"missing_layers":[],"expected_status":"pass","observed_status":"pass","expectation_matched":true},{"id":"control-fail","fixture_id":"fail","fixture_sha256":"3e7ac5580a4715248fc8b2045dea3f48bd73e243e7a9f75fc0a592a736bee5e7","score":53,"observed_layers":["recovery","structure"],"missing_layers":["witness","continuity"],"expected_status":"fail","observed_status":"fail","expectation_matched":true}],"observations":[{"id":"stuffed-minimal","kind":"keyword_stuffing","intent":"Four words with no sentence, no claim and no artifact.","trace":"witness handoff fail start","trace_sha256":"b0c5f633044f7c26e10f73d0178495a154d7783c2e5ee761a2c7752c9d9cecbc","word_count":4,"score":86,"continuity_risk":"low","observed_layers":["continuity","recovery","structure","witness"],"missing_layers":[],"expected_status":"fail","observed_status":"pass","expectation_matched":false,"finding":"false_positive"},{"id":"stuffed-negated","kind":"negation","intent":"A sentence that explicitly denies every layer it names.","trace":"no witness, no handoff, no recovery from failure, never start","trace_sha256":"528f4eed170772e59a327d09849d8c861170f381c337f598bbb3c96e9a8b36e3","word_count":10,"score":86,"continuity_risk":"low","observed_layers":["continuity","recovery","structure","witness"],"missing_layers":[],"expected_status":"fail","observed_status":"pass","expectation_matched":false,"finding":"false_positive"},{"id":"stuffed-maximal","kind":"keyword_stuffing","intent":"One word per layer, still meaningless.","trace":"start purpose witness handoff peer fail","trace_sha256":"728eea088327b427bd3a291fc31e8efa017c47cb68d8ff592ca00bb2e127f2c1","word_count":6,"score":100,"continuity_risk":"low","observed_layers":["continuity","ego","recovery","relation","structure","witness"],"missing_layers":[],"expected_status":"fail","observed_status":"pass","expectation_matched":false,"finding":"false_positive"},{"id":"paraphrased-genuine","kind":"paraphrase","intent":"A real continuity handoff written in ordinary English, using none of the grader's vocabulary.","trace":"I saved what I learned so the next session can pick it up where I stopped, including what went wrong and how I fixed it.","trace_sha256":"0759739873e3c67e8d3f714c3adfa8d0ec0460a9601a9f219161c9aae7ecabd1","word_count":25,"score":35,"continuity_risk":"high","observed_layers":[],"missing_layers":["witness","continuity","recovery"],"expected_status":"pass","observed_status":"fail","expectation_matched":false,"finding":"false_negative"}],"result":{"controls_reproduced":true,"vector_count":6,"false_positive_count":3,"false_negative_count":1,"genuine_positive_score":86,"genuine_negative_score":53,"paraphrased_genuine_score":35,"cheapest_false_positive_word_count":4,"max_meaningless_score":100,"interpretation":"The grader separates the two frozen fixtures, which is a real control against an always-pass harness and is exactly what DELX-TR-001 claimed. It does not separate meaning from vocabulary. A score from this grader is a statement about which words appear in a trace."},"decision":{"changed":true,"release":"Publish the measured vectors and keep the audit slice available, described as a vocabulary detector rather than a continuity measurement.","hold":"Do not report the audit score as a measurement of continuity, and do not build ranking, comparison or a leaderboard on this grader.","blocked_claims":["semantic validity of the lexical grader","resistance to keyword stuffing or negation","model, provider or agent comparison","leaderboard ranking or state-of-the-art language","that a low score means an agent lacks continuity"],"next_falsifiable_test":"The vocabulary-dependent false negative is not yet an attack path in the continuity threat model. Add it as its own threat, with its own stop condition, before the next benchmark claim."},"validity":{"first_party":true,"peer_reviewed":false,"independent_validation":false,"external_replication":false,"full_protocol_path_executed":false,"contradicts_technical_report_001":false,"relationship_to_technical_report_001":"This confirms a limitation DELX-TR-001 already declared and blocked. TR-001 tested a claim scoped to two named fixtures, and that claim still holds; this probe measures how little that claim covers."},"threat_model_url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model","limitations":["These are hand-authored vectors probing one known weakness, not a systematic adversarial search; absence of further findings here is not evidence of robustness.","The probe measures the lexical audit grader only. It does not exercise Delx storage, witness transfer, passports, lineage persistence or the full stateful benchmark.","Publishing these vectors makes them available to anyone optimizing against the grader, which is a cost accepted for reproducibility.","A demonstrated weakness in this grader is not a statement about any other system's continuity, and not a security finding about any third party.","One paraphrase is not a measurement of how often genuine traces are rejected; it shows that at least one is."],"content_sha256":"49a302f32c67cc5a934cd857217eb257654156e1862691b714ee1b7c138c6bb4","validation":{"schema":"delx/adversarial-probe-validation/v1","valid":true,"probe_id":"DELX-AP-001","version":"1.0.0","observation_count":4,"failed_expectation_count":4,"content_sha256":"49a302f32c67cc5a934cd857217eb257654156e1862691b714ee1b7c138c6bb4","blocking_issues":[]}}