hotato.assert_ checks a call's transcript, ingested trace, and timing
against a fixed set of typed assertions, each kind scored on its own lane by
construction. Every kind is a regex, checksum, or span/dict lookup --
deterministic end to end, no model call in the loop.
from hotato import assert_ as A
ctx = A.build_context(
transcript_path="transcript.json", # or transcript=[...] turns already in hand
trace_path="voice_trace.jsonl", # or spans=[...] already in hand
timing=envelope, # optional: a run's envelope.v1 events
)
env = A.run_assertions_from_file("assertions.yaml", ctx)
print(env["exit_code"]) # 0 pass, 1 failPython API: docs/API.md covers the scoring envelope this
complements; the schema is src/hotato/schema/assert.v1.json.
Runnable ground truth:
examples/reference-agent exercises these
kinds end to end: a 375-run offline suite (25 scenarios, 5 caller
behaviours, 3 audio environments) whose tool_result, state, sequence,
handoff, and tool_error assertions surface four seeded agent defects,
deterministically. Method: the "Say-do verification methodology" section of
METHODOLOGY.md.
Every kind is deterministic, so every result carries deterministic: true,
including INCONCLUSIVE. The five below are the original core; the full
vocabulary is under the whole kind vocabulary.
| Kind | Checks | Reads |
|---|---|---|
phrase |
a regex is present (or, in absent mode, never present), with an optional role filter and position (first/last/any) |
transcript text |
pii |
deterministic detectors (ssn, card_luhn with full Luhn validation, email, phone) find nothing, mode: must_not_leak |
transcript text |
policy |
a named, versioned, offline rule pack's banned-language and required-disclosure rules | transcript text |
tool_call |
a tool was (or was not) called, with an optional argument subset, a count bound, a required order across tools, or a "never before" ordering constraint | ingested voice_trace.v1 spans only |
outcome |
task success as all_of/any_of a list of the sub-predicates above (tool_called, phrase, field_present), reported as a met/of fraction |
whichever context each sub-predicate needs |
tool_call checks only the ingested trace (hotato trace ingest,
docs/TRACE.md) -- that's the evidence a tool ran; an agent's own
words claiming it ran don't count. pii surfaces only a [REDACTED]
transcript artifact plus hit metadata (detector name, turn index, role) --
never the matched text.
The full assert.v1 kind vocabulary is 20 deterministic kinds, all
deterministic: true, all model-free. Beyond the five core kinds:
tool_result and tool_error (a tool's returned value or raised error, from
the trace), http_result (a recorded HTTP exchange span's method, URL,
status, and response subset, from the trace), state and state_change (a
state adapter's snapshot or
transition -- see STATE-ADAPTERS.md), handoff, dtmf,
termination, latency, timing_contract, entity_accuracy, sequence,
count, formula (a boolean composite over other assertions' results in
the same run, below), and order (a transcript ordering check, below). Two
more, human_rubric and judge_rubric, belong to the SEPARATE
model-judged rubric lane (RUBRIC.md); inside a raw assert.v1
document they resolve to a deterministic INCONCLUSIVE, so no model runs here
and the guarantee holds.
http_result reads http_exchange spans from the ingested
hotato.voice_trace.v1 trace -- a recorded request/response report carrying
method, url, status_code, and response. Evaluation is a pure span
lookup: hotato never performs the request, so a rerun is deterministic and
offline, and it shares the Authority-1 wall with tool_result /
tool_error -- an agent's spoken claim about a request can never satisfy it.
version: 1
assertions:
- id: refund-posted
kind: http_result
method: POST # matched case-insensitively
url_matches: "/v1/refunds$" # a regex searched against the span's url
status: 201 # one status, or a list: [200, 201]
response_subset: {status: ok} # optional: fields the response must carryThe span's method and url select the exchange; status and
response_subset then judge it. PASS carries the grounding span_ids;
FAIL distinguishes "no exchange matched the method and URL" from "the
exchange matched but its status/response did not"; a context with no trace at
all reports INCONCLUSIVE, never a guess. A malformed assertion (missing
method, an invalid url_matches regex, a status outside 100-599) is a
usage error caught up front, before any assertion runs.
formula combines OTHER named assertions' results from the same run into one
verdict. Its expr is a boolean expression over assertion ids -- and, or,
not, parentheses -- plus a weighted-sum comparison
(0.6*a + 0.4*b >= 0.5, where a bare id weighs 1). A referenced PASS is
true and a FAIL is false; the weighted form sums the weights of the
referenced assertions that passed and compares the total with the threshold.
version: 1
assertions:
- id: refund-issued
kind: tool_call
name: issue_refund
- id: confirmation-said
kind: phrase
regex: "confirmation number"
role: agent
- id: identity-verified
kind: tool_call
name: verify_identity
- id: happy-path
kind: formula
expr: "refund-issued and (confirmation-said or identity-verified)"
- id: mostly-healthy
kind: formula
expr: "0.5*refund-issued + 0.3*confirmation-said + 0.2*identity-verified >= 0.7"The expression is parsed by a small recursive-descent parser -- never
eval() -- and the whole reference graph is checked up front, before any
assertion runs: an unknown id, a self-reference, or a reference cycle
(including one through another formula or a when:) is a usage error
(ValueError, exit 2). References evaluate first whatever the document
order (a formula may reference another formula), and results are always
emitted in document order, byte-stable.
A referenced INCONCLUSIVE makes the formula INCONCLUSIVE, with a reason
naming the reference -- absent input propagates; a composite never guesses.
A determinate result carries refs (the referenced ids), met/of (how many
referenced assertions passed), and, on FAIL, a reason listing which
references passed and which did not.
order checks the sequence of two phrases in the transcript: the FIRST turn
matching the before regex must precede (be strictly earlier than) the FIRST
turn matching after. It measures where two phrases first appear, not why they
appear. It reads only transcript text, so it needs a --transcript.
version: 1
assertions:
- id: verify-before-account-details
kind: order
before: "verify your identity|date of birth|last four"
after: "your account number is|current balance"
role: agent # optional: match only this speaker's turns
case_sensitive: false # optional, default falsePASS carries before_turn and after_turn (the zero-based turn indices the
two phrases first matched at). If either phrase never matches there is no
ordering constraint to violate, so the result is a vacuous PASS carrying
vacuous: true. A context with no transcript reports INCONCLUSIVE, never a
guess. The phrases present but in the wrong order is a FAIL. A malformed
assertion (a missing or invalid before/after regex, a non-string role) is
a usage error caught up front, before any assertion runs.
Any assertion -- any kind -- may carry an optional when: precondition: one
assertion id, or a list of ids. Unless every referenced assertion PASSed,
the assertion is SKIPPED: an INCONCLUSIVE result with skipped: true and a
reason naming the unmet reference(s), because the check itself never ran --
there is no verdict to report.
version: 1
assertions:
- id: refund-issued
kind: tool_call
name: issue_refund
- id: refund-amount-correct
kind: tool_result
name: issue_refund
result_subset: {status: ok}
when: refund-issued # or a list: [refund-issued, other-id]A referenced FAIL and a referenced INCONCLUSIVE both skip -- the
precondition demands a PASS. when: references obey the same up-front
unknown-name and cycle refusal as formula, and the referenced assertions
always evaluate first. Under inconclusive_policy: fail/refuse a skip gates
like any other INCONCLUSIVE, so a suite that must not stay green on skipped
checks can opt in.
build_context assembles the three inputs an assertion run needs, each built
from hotato's existing primitives:
- transcript:
hotato.transcribe(the opt-in[transcribe]extra, faster-whisper) produces one, or pass--transcript FILE/transcript_path=with a JSON file -- a plain array of{role, text, start, end}turns, or the{"segments": [...]}shapehotato.transcribeand the MCP surface write. - trace:
hotato trace ingest --otel FILE --out voice_trace.jsonl(docs/TRACE.md,docs/OTEL.md) produces thehotato.voice_trace.v1spanstool_callreads (name,arguments, and -- when the source trace carries them --result/error) andhttp_resultreads (method,url,status_code,response). - timing: a scoring run's own envelope (
hotato run --format json,docs/API.md) passed straight through as read-only context foroutcome'sfield_presentsub-predicate. Nothing here recomputes it.
Context you never supply stays None, distinct from a supplied [] or {}
that happens to be empty. An assertion whose required input is absent reports
INCONCLUSIVE. tool_call with spans=[] (a trace was ingested with zero
spans) is a FAIL, distinct from tool_call with no trace at all, which
reports INCONCLUSIVE.
A small, dependency-free YAML subset (block mappings/sequences, flow
[...]/{...}, quoted or bare scalars, # comments) -- or valid JSON,
accepted directly. Hotato parses this subset itself, so the core stays zero
third-party dependency either way.
version: 1
assertions:
- id: refund-confirmed
kind: outcome
all_of: [{tool_called: issue_refund}, {phrase: "confirmation number", role: agent}]
- id: tool-order
kind: tool_call
require_order: [verify_identity, lookup_account, issue_refund]
never_before: {tool: issue_refund, until: verify_identity}
- id: disclosure
kind: phrase
regex: "recorded for quality"
role: agent
position: first
- id: no-ssn-leak
kind: pii
detectors: [ssn, card_luhn]
mode: must_not_leakEvery assertion needs a unique id and a recognized kind. Kind-specific
fields are validated (bad regex, unknown detector, missing required field,
unsupported version) up front -- a malformed file is caught whole, before
partial results exist.
This is the entire point of the module: structural, not a convention someone can quietly break.
- Every result carries
kindanddeterministic: true-- true on every deterministic kind and every status, includingINCONCLUSIVE, itself a deterministic read of missing required input. - The envelope's
summarysplitsdeterministic({pass, fail, inconclusive}) fromjudge({pass, fail}), each in its own count -- the schema (src/hotato/schema/assert.v1.json) enforces this with"overall_score": falseand anot: {required: [overall_score]}on the summary object. judge-- an LLM-scored rubric kind -- stays structurally quarantined from the deterministic count, so a model-scored result can never blend in.summary.judgereports{"pass": 0, "fail": 0};summary.notestates how many judge-scored assertions ran.- Same inputs, same file, same result, every time:
run_assertionsis byte-stable across repeated calls on identical input -- no wall-clock timestamp or random id in the mix.
The report (below) renders this as two visually separate shelves, not one number -- visible on the page, not just in the JSON.
hotato.report.build_report_html / build_report_md accept an optional
assertions= parameter: an already-evaluated assert.v1 envelope (build one
with run_assertions / run_assertions_from_file / run_assertions_from_yaml
above). Like base (a previous run envelope) and transcript (an
already-produced ASR artifact), the report purely renders whatever result
it's handed.
from hotato import assert_ as A, report
env = A.run_assertions_from_file("assertions.yaml", ctx)
html, _ = report.build_report_html(suite="barge-in", assertions=env)When present, it adds one "Assertions" section:
- The headline. Two counts side by side, each scored on its own:
N deterministic pass / M fail K judge-scored (advisory). - Deterministic (audio / timing / transcript / trace derived): one
PER-DIMENSION TYPED card per result -- a kind tag, the
deterministicflag, the PASS/FAIL/INCONCLUSIVE chip, and that result's kind-specific fields (apiicard's hit detail and redacted transcript, apolicycard's matched rules and pack name, atool_callcard's grounding span ids, anoutcomecard's met/of fraction). - Model-assisted (advisory, quarantined): stays empty here by design, with a note pointing at the model-judged rubric lane (RUBRIC.md) where that scoring runs.
assertions=None (the default) is byte-identical to a report built before this
parameter existed.
By default an INCONCLUSIVE result -- a check whose required input was
absent -- leaves the exit code unaffected, the right default for an
exploratory run. inconclusive_policy lets a suite gate on that instead, so a
transcript or trace that never arrived fails loudly instead of leaving the
suite silently green:
| value | how INCONCLUSIVE gates |
exit_code |
|---|---|---|
report (default) |
reports missing input and leaves the gate unchanged | 1 if any FAIL, else 0 |
fail |
gates exactly like a FAIL |
1 if any FAIL or INCONCLUSIVE, else 0 |
refuse |
refuses to return a verdict at all | 2 if any INCONCLUSIVE, else 1 if any FAIL, else 0 |
refuse precedence. Under refuse, an INCONCLUSIVE result exits 2
even if another assertion also FAILed -- the exit-2 refusal takes precedence
over the FAIL. A run that cannot fully see its inputs withholds its verdict.
The default is report, so a suite that sets nothing gates exactly as
before this field existed -- fully backward-compatible. CI and compliance
suites should set fail or refuse, so a missing transcript or trace fails
loudly.
Set it as an optional top-level key:
version: 1
inconclusive_policy: fail # or: refuse | report (default)
assertions:
- id: disclosure
kind: phrase
regex: "recorded for quality"
role: agentor override the document's key from the CLI (the flag wins):
hotato assert run --assertions assertions.yaml --transcript call.json \
--inconclusive-policy refuseor from Python (an explicit argument overrides the document's key; absent both,
report applies):
env = A.run_assertions_from_file("assertions.yaml", ctx, inconclusive_policy="fail")A bad value (anything but report/fail/refuse) -- in the document or
passed explicitly -- is a usage error (ValueError, exit 2), raised during
validation before any assertion runs. The envelope always carries the applied
inconclusive_policy, stated with the counts in summary.note.
Same exit-code convention as every hotato command:
| Exit | Meaning |
|---|---|
0 |
every assertion passed (under report, an INCONCLUSIVE reports missing input and leaves the exit at 0; under fail it gates like FAIL; under refuse it exits 2) |
1 |
at least one deterministic status is FAIL (or, under fail, INCONCLUSIVE) |
2 |
a refusal under refuse (an INCONCLUSIVE, taking precedence over a FAIL), or a malformed file / bad input, raised before any assertion runs |
python3 - <<'PY'
from hotato import assert_ as A
import sys
ctx = A.build_context(transcript_path="transcript.json", trace_path="voice_trace.jsonl")
env = A.run_assertions_from_file("assertions.yaml", ctx)
print(env["summary"]["note"])
sys.exit(env["exit_code"])
PYA gate on assert.v1 and a gate on the timing scorer's exit_code
(hotato run / hotato verify, docs/CI.md) are two different,
composable guarantees: one gates turn-taking timing, the other gates
transcript/trace content. Run both; neither exit code substitutes for the
other's.