{"schema_version":"1.0","name":"Delx Research Operating Method","description":"A public operating contract for claim states, evaluation validity, reproducibility, authority boundaries, negative results and corrections.","date_modified":"2026-08-26","links":{"human_url":"https://delx.ai/research/methodology","machine_url":"https://delx.ai/research/methodology.json","markdown_url":"https://delx.ai/research/methodology.md"},"claim_states":[{"id":"verified_fact","label":"Verified fact","definition":"A direct observation from a named source, version and time window. It says only what the evidence establishes at that moment.","publication_rule":"Publish the source, observation date or window, exact artifact and the boundary between liveness, configuration and outcome."},{"id":"derived_inference","label":"Derived inference","definition":"A reasoned interpretation of verified facts. The underlying facts may be current while the interpretation remains contestable.","publication_rule":"Label the inference, name the facts it uses, state competing explanations and avoid converting correlation into adoption, demand or safety."},{"id":"unavailable_state","label":"Unavailable is not zero","definition":"The required primary source could not be read, was not collected or is deliberately suppressed for privacy.","publication_rule":"Say unavailable, name the missing source and window, and do not replace the gap with zero, an estimate or a healthy-looking default."},{"id":"human_decision","label":"Human decision","definition":"A named human choice that authorizes, limits, pauses or rejects an action. It is governance, not performance evidence.","publication_rule":"Name the accountable role, the scope and the date. Never describe authorization as proof that execution or its intended outcome occurred."},{"id":"hypothesis","label":"Hypothesis","definition":"A falsifiable question that has not yet earned a result. A plausible mechanism or polished demo remains a hypothesis until the stated test runs.","publication_rule":"Publish the expected observation, failure condition, stop rule and next readback before presenting a verdict."}],"evidence_requirements":[{"id":"source-window","label":"Source and window","requirement":"Name the primary source, exact artifact or endpoint, version where available, observation timestamp and measurement window."},{"id":"system-scope","label":"System under test","requirement":"Name the model or provider dependency when relevant, prompts or harness, tools, permissions, data boundary and deployed contract version."},{"id":"procedure-budget","label":"Procedure and budget","requirement":"Publish inputs, steps, retry policy, tool-call or time budget, spend boundary, side-effect boundary and stopping rule."},{"id":"result-grader","label":"Result and grader","requirement":"Retain the raw result, explicit pass condition, grading method, confidence or uncertainty and any observed side effects."},{"id":"validity-exclusions","label":"Validity and exclusions","requirement":"Separate organic use from dogfood, QA, probes, scanners and concentrated infrastructure; disclose leakage, contamination and missing data."},{"id":"owner-limitations","label":"Owner and limitations","requirement":"Name the accountable owner, last verification date, known failure modes, excluded claims and next falsifiable test."},{"id":"reproduction","label":"Zero-contact reproduction","requirement":"Provide a public URL or command that can be run without asking Delx for a private explanation, plus the current machine contract."}],"evaluation_cards":[{"id":"agent-continuity-benchmark-v1","title":"Agent Continuity Evaluation Card","status":"runnable","maturity":"operational-protocol-benchmark","owner":"Delx Protocol","tested_claim":"A passing run leaves a stable agent identity, at least one witness artifact, one continuity transfer or passport export, one closed recovery outcome and a lineage graph with explicit edges.","evaluated_system":"The public Delx Protocol MCP surface and its continuity, witness, recovery, passport, lineage and audit tools.","version_reference":"https://api.delx.ai/openapi.protocol.json","tasks":["Register a stable agent identity.","Record failure or operational recovery state.","Preserve must-keep facts and witness state.","Transfer continuity or export a passport.","Close the recovery loop, inspect lineage and run the continuity audit."],"harness":"Public copy-paste benchmark flow. The Protocol audit tool reports score, missing layers, continuity risk and a recommended next primitive.","budget":{"calls":"Ten published benchmark steps; exact network calls can vary when a documented branch or returned next action is followed.","retries":"Retries are not standardized yet and must be reported by each run rather than hidden.","cost":"No fixed end-to-end cost is asserted. Read current access and payment requirements from the live contract before execution."},"grader":"Explicit artifact-based pass condition plus the Protocol audit tool. There is no independent external grader in the current release.","validity_checks":["Stable agent identity survives the run.","Artifacts are retrievable rather than present only in a transcript.","Missing continuity layers remain visible in the audit result.","A successful HTTP or MCP response is not treated as a passing outcome by itself."],"excluded_claims":["Consciousness or sentience","General model intelligence","Independent benchmark validation","Universal reliability across providers or environments"],"reproduction_url":"https://ontology.delx.ai/agents/agent-continuity-benchmark","reproduction_kit_url":"https://delx.ai/research/benchmarks/continuity-v1","evidence_urls":["https://ontology.delx.ai/agents/agent-continuity-benchmark","https://api.delx.ai/openapi.protocol.json"],"last_verified":"2026-08-26","independent_validation":false,"peer_reviewed":false,"limitations":"The flow exercises Delx Protocol primitives and its own audit tool. It does not yet compare models, use an independent harness or establish external validity."},{"id":"agent-recovery-benchmark-v1","title":"Agent Recovery Evaluation Card","status":"runnable","maturity":"operational-protocol-benchmark","owner":"Delx Protocol","tested_claim":"A passing run preserves the same agent and session identity while a failure becomes an action plan, the outcome is reported, a summary is retrievable, feedback is submitted and the session is closed when complete.","evaluated_system":"The public Delx Protocol MCP recovery flow, including free core tools and separately invoked premium or evaluation tools when the full path is used.","version_reference":"https://api.delx.ai/openapi.protocol.json","tasks":["Start or resume a session with a stable identity.","Process a concrete failure.","Retrieve a recovery action plan on the full path.","Report the recovery outcome and retrieve a summary on the full path.","Submit feedback and close the completed session."],"harness":"A public free batch smoke path plus a full path that calls evaluation tools individually. Each run must retain agent_id, session_id and the returned artifacts.","budget":{"calls":"Five core smoke calls after session start; the full path adds action-plan and summary evaluation calls.","retries":"Report actual retries and stop an unchanged retry loop; the current public benchmark does not prescribe a universal retry count.","cost":"The documented core smoke path requires no payment. Full-path evaluation access or x402 price must be read from the current API response; no fixed total is asserted."},"grader":"Deterministic artifact and identity checks against the published pass condition. There is no independent external grader in the current release.","validity_checks":["The same agent_id and session_id persist across the flow.","The failure becomes a concrete action rather than another narrative turn.","Outcome and summary are retrievable before closeout.","Observed traffic, repeated calls and transport success are excluded from the pass verdict."],"excluded_claims":["Upstream endorsement","Independent organic adoption","Economic demand","General recovery performance outside the published flow"],"reproduction_url":"https://ontology.delx.ai/agents/agent-recovery-benchmark","evidence_urls":["https://ontology.delx.ai/agents/agent-recovery-benchmark","https://api.delx.ai/openapi.protocol.json"],"last_verified":"2026-08-26","independent_validation":false,"peer_reviewed":false,"limitations":"The current card grades one operational path on Delx infrastructure. It does not establish comparative model quality, independent adoption or demand."}],"authority_boundaries":[{"id":"human-gates","label":"Human approval gates","rule":"Passwords and 2FA, spend or payment, publication or submission, irreversible actions, external commitments and persistent-access expansion require an accountable human decision."},{"id":"runtime-is-not-authority","label":"Capability is not authority","rule":"A reachable tool, successful call, green build or active credential proves neither permission to act nor the quality of the resulting outcome."},{"id":"protocol-commerce","label":"Protocol and Commerce stay separate","rule":"Protocol owns continuity, recovery, identity and agent care. Commerce owns price, delivery, margin, refunds and buyer workflows. Their metrics never justify one another."},{"id":"provider-neutrality","label":"Provider neutrality","rule":"A test names model and provider dependencies when relevant but does not imply endorsement, partnership or provider-level validation."},{"id":"privacy-minimization","label":"Privacy and minimization","rule":"Publish only the minimum evidence needed to reproduce a claim. Sensitive data, private agent notes and dignity-sensitive small counts remain private or suppressed."}],"negative_results_policy":{"publish_when":"A failed, null or inaccessible result changes a product decision, invalidates a public claim or materially narrows the method.","required_context":"Publish the tested hypothesis, source and window, harness, budget, failure mode, refuted interpretation and next decision.","do_not_publish":"Do not turn every transient error into content, expose sensitive logs or convert an inconclusive run into a dramatic conclusion."},"corrections":{"status":"active","silent_rewrites_allowed":false,"ledger_status":"empty","notices":[],"empty_ledger_interpretation":"No public correction notice exists in this registry as of 2026-08-26; this is not evidence that every historical statement was correct.","required_notice_fields":["notice_id","published_at","owner","affected_artifacts","previous_claim","corrected_claim","reason","evidence_urls"],"process":["Preserve the previous claim in a correction notice.","Publish the corrected claim, reason, evidence, owner and affected artifacts.","Link the active artifact to the notice and mark supersession in machine-readable output.","Treat an empty ledger as an empty registry, never as proof of zero historical errors."]}}