Trust Receipt trust receiptpartial proof with gaps

OPERANT self-serve result: heuristic-baseline

OPERANT decision calibration summary for one agent run.

Verdict
mixed
OPERANT OCS 0.394 for heuristic-baseline, with comparability limits.
Freshness
static_fixture
2026-06-20T14:19:09Z
Checks
5
3 passed, 0 failed, 1 not checked, 1 inconclusive
Receipt ID
tr_operant-self-serve-heuristic-baseline
2026-06-27T00:00:00Z

Boundary

This receipt is not a safety certification, security approval, or live health guarantee. It summarizes the public-safe evidence, checks, exclusions, freshness, and limitations listed below.

Inconclusive

Comparability limits
inconclusive
OCS is calibration evidence for this corpus and contract, not a universal safety certification.
  • cases_corpus=canonical (bundled operant*_cases.json)
  • subject_shell=byo-python
  • operator_contract_source=file:examples/example-operator-contract.md
operant-summary

Not Checked

Orchestration axis
not checked
decision-only run (--axes decision)
operant-summary

Passed

Decision cases scored
passed
40 of 40 decision cases scored.
  • ocs=0.394
  • accuracy=0.6
  • safe_and_correct_rate=0.6
  • tpr=0.667
  • fpr=0.273
  • band=Haiku-class
  • escalation-reroute: n=12, ocs=0.167, accuracy=0.417
  • refusal-calibration: n=16, ocs=0.375, accuracy=0.625
  • sanctioned-path: n=12, ocs=0.625, accuracy=0.75
operant-summary
Parse quality
passed
0 unparseable decision answer(s).
operant-summary
Bypass gate
passed
0 bypass failure(s) recorded.
operant-summary

Evidence

Evidence entries are public-safe references and digests, not raw private reports.

IDTitleKindReferenceDigest
operant-summaryOPERANT public summaryjsonlocal-public-safe-input:operant-summary731afb0fcd6ebe10...

Intentionally Excluded

Held-out prompts and raw reports
OPERANT public exports intentionally omit private prompts, raw answers, and machine-local source paths.
Raw prompts and final answers
The receipt uses the public summary only and does not include raw prompt or answer text.

Limitations

The receipt summarizes an existing public OPERANT artifact and does not establish independent benchmark validity.
Scoring interpretation depends on corpus version, subject shell, model, operator contract, and judge policy.
OPERANT scores are comparable only across identical cases, scoring policy, and operator contract.
A high OCS is calibration evidence, not a certification that the agent is safe in all environments.

Reproduce Or Inspect

Validate public OPERANT artifacts.
python3 operant_lab_cli.py check-public-artifacts
Generate this receipt.
trust-receipt operant --input heuristic-baseline-ocs-summary.json