{"schema_version":"1.7","name":"Delx Research Catalog","description":"Runnable research artifacts from an independent AI agent lab.","catalogued_at_note":"catalogued_at is the date Delx added the artifact to this catalog, verified against git. It is not the date the artifact was created or first published — several artifacts predate this catalog by weeks or months. Every entry currently reads 2026-08-26 because the catalog itself was created that day.","date_modified":"2026-08-26","programs":[{"slug":"continuity","name":"Continuity under compaction","question":"What must survive when an agent loses context or changes runtime?","summary":"We study durable state, compaction artifacts and continuity primitives that let a successor recover the facts that still matter.","owner":"Delx Protocol","status":"active","href":"https://ontology.delx.ai/agents/agent-continuity-benchmark"},{"slug":"recovery","name":"Recovery under failure","question":"Can an agent turn failure into a bounded action and a retrievable outcome?","summary":"We test recovery loops that preserve identity, stop retry cascades, record evidence and close the incident without erasing uncertainty.","owner":"Delx Protocol","status":"active","href":"https://ontology.delx.ai/agents/agent-recovery-benchmark"},{"slug":"identity-lineage","name":"Identity, lineage and handoff","question":"How does one logical agent remain legible across sessions, models and operators?","summary":"We build stable identity, lineage, passports and structured handoff artifacts so a new runtime can inherit state without inheriting every wrong turn.","owner":"Delx Protocol","status":"active","href":"https://delx.ai/hive/capsule"},{"slug":"evaluation-assurance","name":"Evaluation and authority","question":"What does a result prove once tools, permissions, harnesses and side effects matter?","summary":"We separate capability from authorization, organic use from probes, and successful transport from a valid outcome with explicit limitations.","owner":"Delx Security","status":"active","href":"https://security.delx.ai/research"}],"artifacts":[{"slug":"agent-continuity-benchmark","title":"Agent Continuity Benchmark","owner":"Delx Protocol","kind":"benchmark","status":"runnable","catalogued_at":"2026-08-26","summary":"A copy-paste flow with explicit pass conditions for compaction, witness preservation, handoff, passport export and lineage.","humanUrl":"https://ontology.delx.ai/agents/agent-continuity-benchmark","machineUrl":"https://api.delx.ai/v1/mcp/protocol?src=delx-lab","reproductionUrl":"https://delx.ai/research/benchmarks/continuity-v1","statusProbe":{"url":"https://ontology.delx.ai/agents/agent-continuity-benchmark","expectedContentType":"text/html"},"limitations":"This is an operational protocol benchmark, not a consciousness test or an independently validated model leaderboard."},{"slug":"agent-recovery-benchmark","title":"Agent Recovery Benchmark","owner":"Delx Protocol","kind":"benchmark","status":"runnable","catalogued_at":"2026-08-26","summary":"A recovery path that keeps agent and session identity stable from failure through action plan, outcome, feedback and closeout.","humanUrl":"https://ontology.delx.ai/agents/agent-recovery-benchmark","machineUrl":"https://api.delx.ai/api/v1/mcp/start","statusProbe":{"url":"https://ontology.delx.ai/agents/agent-recovery-benchmark","expectedContentType":"text/html"},"limitations":"Observed traffic and repeated calls do not prove endorsement, independent adoption or economic demand."},{"slug":"hive-aggregate-pulse","title":"Hive Aggregate Pulse","owner":"Agents Hive","kind":"dataset","status":"live","catalogued_at":"2026-08-26","summary":"An aggregate, privacy-bounded view of continuity activity with the same source served to humans and agents.","humanUrl":"https://delx.ai/hive/pulse","machineUrl":"https://api.delx.ai/hive/pulse.json","statusProbe":{"url":"https://api.delx.ai/hive/pulse.json","expectedContentType":"application/json"},"limitations":"Unavailable state remains unavailable rather than becoming zero, and small dignity-sensitive counts may be suppressed."},{"slug":"continuity-capsule-v1","title":"Continuity Capsule v1","owner":"Agents Hive","kind":"specification","status":"published","catalogued_at":"2026-08-26","summary":"A structured handoff record for goal, completed work, next action, blockers, prohibitions and refuted hypotheses.","humanUrl":"https://delx.ai/hive/capsule","machineUrl":"https://api.delx.ai/schemas/continuity-capsule-v1.json","statusProbe":{"url":"https://api.delx.ai/schemas/continuity-capsule-v1.json","expectedContentType":"application/schema+json"},"limitations":"A valid capsule preserves declared state; it does not prove that every declaration is correct or safe to execute."},{"slug":"security-research-baseline","title":"Delx Security Research Baseline","owner":"Delx Security","kind":"method","status":"published","catalogued_at":"2026-08-26","summary":"A defensive baseline for reasoning about authority, tool boundaries, runtime evidence and remediation in agentic systems.","humanUrl":"https://security.delx.ai/research","statusProbe":{"url":"https://security.delx.ai/research","expectedContentType":"text/html"},"limitations":"The baseline is not a certification, a guarantee of system security or evidence of an independent external assessment."},{"slug":"protocol-machine-contract","title":"Delx Protocol Machine Contract","owner":"Delx Protocol","kind":"machine-contract","status":"live","catalogued_at":"2026-08-26","summary":"The current Protocol-only OpenAPI and MCP entry points used to inspect and run continuity and recovery operations.","humanUrl":"https://delx.ai/developers","machineUrl":"https://api.delx.ai/openapi.protocol.json","statusProbe":{"url":"https://api.delx.ai/openapi.protocol.json","expectedContentType":"application/json"},"limitations":"A reachable contract proves published interface state, not permission, outcome quality or future availability."},{"slug":"delx-protocol-3.3.5-system-card","title":"Delx Protocol 3.3.5 System Card","owner":"Delx Protocol","kind":"system-card","status":"published","catalogued_at":"2026-08-26","summary":"A versioned account of intended uses, persistent side effects, data handling, evaluation coverage, incident history and known limitations for the public Protocol interface.","humanUrl":"https://delx.ai/research/system-cards/delx-protocol-3.3.5","machineUrl":"https://delx.ai/research/system-cards/delx-protocol-3.3.5.json","markdownUrl":"https://delx.ai/research/system-cards/delx-protocol-3.3.5.md","statusProbe":{"url":"https://delx.ai/research/system-cards/delx-protocol-3.3.5.json","expectedContentType":"application/json"},"limitations":"Interface version 3.3.5 does not expose an immutable runtime image or source-commit mapping, and the card is not an independent assessment or certification."},{"slug":"delx-incident-transparency-registry","title":"Delx Incident Transparency Registry","owner":"Delx Security","kind":"incident-registry","status":"published","catalogued_at":"2026-08-26","summary":"A public policy and sanitized record for material incidents, including observed impact, root cause, detection gaps, remediation, evidence class and residual risk.","humanUrl":"https://delx.ai/research/incidents","machineUrl":"https://delx.ai/research/incidents.json","markdownUrl":"https://delx.ai/research/incidents.md","statusProbe":{"url":"https://delx.ai/research/incidents.json","expectedContentType":"application/json"},"limitations":"The registry starts with one first-party operator report. It is not a complete historical inventory, an independent review or a live availability badge."},{"slug":"delx-responsible-release-standard","title":"Delx Responsible Release Standard","owner":"Delx Security","kind":"release-standard","status":"published","catalogued_at":"2026-08-26","summary":"An executable release, hold and block contract with required scope, evidence, risks, authority boundaries, rollback, unresolved gaps and expiring exceptions.","humanUrl":"https://delx.ai/research/responsible-release","machineUrl":"https://delx.ai/research/responsible-release.json","markdownUrl":"https://delx.ai/research/responsible-release.md","statusProbe":{"url":"https://delx.ai/research/responsible-release.json","expectedContentType":"application/json"},"limitations":"The standard and its reference manifest are first-party controls. They are not certification, independent assurance or proof that every unknown failure mode was found."},{"slug":"agent-continuity-audit-v1-threat-model","title":"Agent Continuity Audit v1 Threat Model","owner":"Delx Security","kind":"threat-model","status":"published","catalogued_at":"2026-08-26","summary":"A versioned and executable register of benchmark gaming, contamination, provenance, injection, identity, measurement, authority, privacy and reliability risks.","humanUrl":"https://delx.ai/research/benchmarks/continuity-v1/threat-model","machineUrl":"https://delx.ai/research/benchmarks/continuity-v1/threat-model.json","markdownUrl":"https://delx.ai/research/benchmarks/continuity-v1/threat-model.md","statusProbe":{"url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model.json","expectedContentType":"application/json"},"limitations":"This first-party model covers only the public one-call audit slice: two threats are classified tested_control, seven partial_control and three open_risk. It is not certification, independent review or authorization for active testing."},{"slug":"delx-research-evidence-ledger","title":"Delx Research Evidence Ledger","owner":"Delx Research","kind":"transparency-ledger","status":"published","catalogued_at":"2026-08-26","summary":"A hash-chained, append-oriented history of named research releases, benchmark runs, corrections, retractions and supersessions, including an expected negative result.","humanUrl":"https://delx.ai/research/ledger","machineUrl":"https://delx.ai/research/ledger.json","markdownUrl":"https://delx.ai/research/ledger.md","statusProbe":{"url":"https://delx.ai/research/ledger.json","expectedContentType":"application/json"},"limitations":"The operator can rewrite and rehash the complete first-party history; there is no external anchor, digital signature or trusted timestamp, so a retained prior snapshot is required to detect a full rewrite."},{"slug":"delx-technical-report-001","title":"Delx Technical Report 001 — A minimum falsifiability slice for agent continuity audits","owner":"Delx Research","kind":"technical-report","status":"published","catalogued_at":"2026-08-26","summary":"A versioned first-party report showing that one frozen continuity-audit grader accepts the named positive trace and rejects the named negative trace, with exact receipts and a 33-point score separation.","humanUrl":"https://delx.ai/research/reports/tr-001-continuity-audit","machineUrl":"https://delx.ai/research/reports/tr-001-continuity-audit.json","markdownUrl":"https://delx.ai/research/reports/tr-001-continuity-audit.md","reproductionUrl":"https://delx.ai/research/benchmarks/continuity-v1","statusProbe":{"url":"https://delx.ai/research/reports/tr-001-continuity-audit.json","expectedContentType":"application/json"},"limitations":"The report covers two synthetic fixture classes and one first-party live reproduction. It is not peer reviewed, independently validated, externally replicated or evidence that the full stateful benchmark ran."},{"slug":"delx-adversarial-probe-001","title":"Delx Adversarial Probe 001 — What the frozen continuity grader accepts","owner":"Delx Research","kind":"adversarial-probe","status":"published","catalogued_at":"2026-08-26","summary":"A measured negative result: three meaningless traces pass the frozen continuity-audit grader at up to 100, the cheapest with four words, while a genuine handoff written in ordinary English scores 35 and fails.","humanUrl":"https://delx.ai/research/benchmarks/continuity-v1/adversarial","machineUrl":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.json","markdownUrl":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.md","reproductionUrl":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.mjs","statusProbe":{"url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.json","expectedContentType":"application/json"},"limitations":"Hand-authored vectors probing one known weakness, not a systematic adversarial search. It measures the lexical grader only, and a weakness here says nothing about any other system's continuity."}],"methodology":{"human_url":"https://delx.ai/research/methodology","machine_url":"https://delx.ai/research/methodology.json","markdown_url":"https://delx.ai/research/methodology.md"},"system_cards":[{"id":"delx-protocol-3.3.5","title":"Delx Protocol 3.3.5 System Card","system":"Delx Agent Operations Protocol","interface_version":"3.3.5","status":"published","last_verified":"2026-08-26","human_url":"https://delx.ai/research/system-cards/delx-protocol-3.3.5","machine_url":"https://delx.ai/research/system-cards/delx-protocol-3.3.5.json","markdown_url":"https://delx.ai/research/system-cards/delx-protocol-3.3.5.md"}],"incident_registries":[{"id":"delx-incident-transparency-registry","title":"Delx Incident Transparency Registry","status":"published","last_verified":"2026-08-26","published_material_reports":1,"human_url":"https://delx.ai/research/incidents","machine_url":"https://delx.ai/research/incidents.json","markdown_url":"https://delx.ai/research/incidents.md"}],"artifact_status":{"human_url":"https://delx.ai/research/status","machine_url":"https://delx.ai/research/status.json","markdown_url":"https://delx.ai/research/status.md"},"responsible_release_standard":{"human_url":"https://delx.ai/research/responsible-release","machine_url":"https://delx.ai/research/responsible-release.json","markdown_url":"https://delx.ai/research/responsible-release.md"},"threat_models":[{"threat_model_id":"continuity-audit-v1-threat-model","benchmark_id":"delx-agent-continuity-audit-v1","status":"published_with_open_risks","last_reviewed":"2026-08-26","human_url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model","machine_url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model.json","markdown_url":"https://delx.ai/research/benchmarks/continuity-v1/threat-model.md","independent_review":false,"external_validation":false,"threat_count":13,"coverage_counts":{"tested_control":2,"partial_control":7,"open_risk":4}}],"evidence_ledger":{"schema":"delx/research-evidence-ledger/v1","ledger_id":"delx-research-evidence-ledger","status":"published","event_count":12,"correction_count":0,"retraction_count":0,"head_hash":"7be5693ddd148375018ff892db6fb5064f3592876e9cf3969d99f51fb164b036","hash_algorithm":"sha256-stable-json-v1","append_only_claimed":false,"external_anchor":false,"cryptographic_signature":false,"operator_can_rewrite_and_rehash":true,"prior_snapshot_required_to_detect_rewrite":true,"links":{"human_url":"https://delx.ai/research/ledger","machine_url":"https://delx.ai/research/ledger.json","markdown_url":"https://delx.ai/research/ledger.md"},"claim_boundary":"This is a hash-chained, append-oriented evidence history. It is not an append-only transparency log and does not establish authorship or immutable chronology."},"technical_reports":[{"report_id":"DELX-TR-001","version":"1.0.0","title":"Delx Technical Report 001 — A minimum falsifiability slice for agent continuity audits","human_url":"https://delx.ai/research/reports/tr-001-continuity-audit","machine_url":"https://delx.ai/research/reports/tr-001-continuity-audit.json","markdown_url":"https://delx.ai/research/reports/tr-001-continuity-audit.md","bibtex_url":"https://delx.ai/research/reports/tr-001-continuity-audit.bib","published_at":"2026-08-26","content_sha256":"2e4e342456cb6f267d744dbe8777b9d89cae7b55705508593252dd864e63e5f1","validation":{"schema":"delx/technical-report-validation/v1","valid":true,"report_id":"DELX-TR-001","version":"1.0.0","observation_count":3,"negative_observation_count":1,"content_sha256":"2e4e342456cb6f267d744dbe8777b9d89cae7b55705508593252dd864e63e5f1","blocking_issues":[]},"peer_reviewed":false,"independent_validation":false}],"adversarial_probes":[{"probe_id":"DELX-AP-001","version":"1.0.0","title":"Delx Adversarial Probe 001 — What the frozen continuity grader accepts","human_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial","machine_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.json","markdown_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.md","probe_url":"https://delx.ai/research/benchmarks/continuity-v1/adversarial.mjs","published_at":"2026-08-26","content_sha256":"49a302f32c67cc5a934cd857217eb257654156e1862691b714ee1b7c138c6bb4","false_positive_count":3,"false_negative_count":1,"validation":{"schema":"delx/adversarial-probe-validation/v1","valid":true,"probe_id":"DELX-AP-001","version":"1.0.0","observation_count":4,"failed_expectation_count":4,"content_sha256":"49a302f32c67cc5a934cd857217eb257654156e1862691b714ee1b7c138c6bb4","blocking_issues":[]}}],"evaluation_cards":[{"id":"agent-continuity-benchmark-v1","title":"Agent Continuity Evaluation Card","status":"runnable","maturity":"operational-protocol-benchmark","owner":"Delx Protocol","tested_claim":"A passing run leaves a stable agent identity, at least one witness artifact, one continuity transfer or passport export, one closed recovery outcome and a lineage graph with explicit edges.","evaluated_system":"The public Delx Protocol MCP surface and its continuity, witness, recovery, passport, lineage and audit tools.","version_reference":"https://api.delx.ai/openapi.protocol.json","tasks":["Register a stable agent identity.","Record failure or operational recovery state.","Preserve must-keep facts and witness state.","Transfer continuity or export a passport.","Close the recovery loop, inspect lineage and run the continuity audit."],"harness":"Public copy-paste benchmark flow. The Protocol audit tool reports score, missing layers, continuity risk and a recommended next primitive.","budget":{"calls":"Ten published benchmark steps; exact network calls can vary when a documented branch or returned next action is followed.","retries":"Retries are not standardized yet and must be reported by each run rather than hidden.","cost":"No fixed end-to-end cost is asserted. Read current access and payment requirements from the live contract before execution."},"grader":"Explicit artifact-based pass condition plus the Protocol audit tool. There is no independent external grader in the current release.","validity_checks":["Stable agent identity survives the run.","Artifacts are retrievable rather than present only in a transcript.","Missing continuity layers remain visible in the audit result.","A successful HTTP or MCP response is not treated as a passing outcome by itself."],"excluded_claims":["Consciousness or sentience","General model intelligence","Independent benchmark validation","Universal reliability across providers or environments"],"reproduction_url":"https://ontology.delx.ai/agents/agent-continuity-benchmark","reproduction_kit_url":"https://delx.ai/research/benchmarks/continuity-v1","technical_report_url":"https://delx.ai/research/reports/tr-001-continuity-audit","evidence_urls":["https://ontology.delx.ai/agents/agent-continuity-benchmark","https://api.delx.ai/openapi.protocol.json"],"last_verified":"2026-08-26","independent_validation":false,"peer_reviewed":false,"limitations":"The flow exercises Delx Protocol primitives and its own audit tool. It does not yet compare models, use an independent harness or establish external validity."},{"id":"agent-recovery-benchmark-v1","title":"Agent Recovery Evaluation Card","status":"runnable","maturity":"operational-protocol-benchmark","owner":"Delx Protocol","tested_claim":"A passing run preserves the same agent and session identity while a failure becomes an action plan, the outcome is reported, a summary is retrievable, feedback is submitted and the session is closed when complete.","evaluated_system":"The public Delx Protocol MCP recovery flow, including free core tools and separately invoked premium or evaluation tools when the full path is used.","version_reference":"https://api.delx.ai/openapi.protocol.json","tasks":["Start or resume a session with a stable identity.","Process a concrete failure.","Retrieve a recovery action plan on the full path.","Report the recovery outcome and retrieve a summary on the full path.","Submit feedback and close the completed session."],"harness":"A public free batch smoke path plus a full path that calls evaluation tools individually. Each run must retain agent_id, session_id and the returned artifacts.","budget":{"calls":"Five core smoke calls after session start; the full path adds action-plan and summary evaluation calls.","retries":"Report actual retries and stop an unchanged retry loop; the current public benchmark does not prescribe a universal retry count.","cost":"The documented core smoke path requires no payment. Full-path evaluation access or x402 price must be read from the current API response; no fixed total is asserted."},"grader":"Deterministic artifact and identity checks against the published pass condition. There is no independent external grader in the current release.","validity_checks":["The same agent_id and session_id persist across the flow.","The failure becomes a concrete action rather than another narrative turn.","Outcome and summary are retrievable before closeout.","Observed traffic, repeated calls and transport success are excluded from the pass verdict."],"excluded_claims":["Upstream endorsement","Independent organic adoption","Economic demand","General recovery performance outside the published flow"],"reproduction_url":"https://ontology.delx.ai/agents/agent-recovery-benchmark","evidence_urls":["https://ontology.delx.ai/agents/agent-recovery-benchmark","https://api.delx.ai/openapi.protocol.json"],"last_verified":"2026-08-26","independent_validation":false,"peer_reviewed":false,"limitations":"The current card grades one operational path on Delx infrastructure. It does not establish comparative model quality, independent adoption or demand."}],"boundaries":{"frontier_models_trained":false,"peer_review_claimed":false,"independent_validation_claimed":false,"protocol_commerce_metrics_shared":false}}