Skip to content

[deferred] BENCH-001: Direct-frontier comparison benchmarks (children of #262) #275

Description

@mrnicholasbcarter-code

Scope (minimum v1)

Deterministic benchmark harness runs the same published task via DIRECT API model vs Verdict route, reports cost / latency / completion / regression / failover. Reproducible when seeded.

Do not duplicate

This issue surfaces the existing benchmarking.py + reproducible_benchmarks.py modules (already shipped) and the PR #22 final V1-007 ticket.

Deferred (per audit, not for V1)

  • Multi-model tournament scoring
  • Provider-rotation benchmarks
  • GNN evaluation harness

Refs: evidence/V1_READINESS_AUDIT_2026-08-03.md #239

Metadata

Metadata

Assignees

No one assigned

    Labels

    area-verificationVerification, evidence, and release qualitypythonPython implementation

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions