Skip to content

Phase 2 — evals + observability (traces, OTel, replay, HITL) - #3

Merged
hutusi merged 7 commits into
mainfrom
phase-2-evals-observability
Jun 18, 2026
Merged

hutusi merged 7 commits into
mainfrom
phase-2-evals-observability

Conversation

@hutusi

@hutusi hutusi commented Jun 18, 2026 •

Copy link
Copy Markdown
Contributor

Phase 2 — evals + observability

Builds on Phase 1 (on main). Six focused commits. (Reopened; supersedes #2.)

What's new

  • Trace model (@auriga/core/trace) — every run records an ordered Trace of model_response (with usage), tool_call, skill_loaded (exact name@version+hash), compaction, and verify events. model_response events make a run replayable. JobSpec.require_approval added.
  • Trace emission (@auriga/currus) — runLoop + runJob emit the full event stream via an onTrace hook.
  • Observability (@auriga/capella) — Recorder seals events into a Trace; traceCost rolls up tokens + USD; emitSpans produces an OpenTelemetry span tree (root job span + child span per step, GenAI-style attributes — no-op until an exporter is registered); formatTrace for the CLI.
  • Evals (@auriga/evals) — ReplayProvider replays a trace's recorded responses deterministically (no model calls, throws on divergence); runEval/runEvals replay a batch through the real harness against fresh sandboxes and score (matches recorded state, verify passed, steps, cost); loadEvalCases reads a disk suite.
  • Trace persistence + HITL (@auriga/habenae) — JobStore.saveTrace/loadTrace (in-memory + file + Postgres, with a traces table); JobRecord.approved; the Worker records + persists each run's trace and consults an approval gate — a job with require_approval returns state paused until approved, then resumes to done.
  • CLI — auriga trace <id> · approve <id> · run <id> · eval <dir>.

Verification

  • bun run check — 133 pass / 6 skip, typecheck clean.
  • Highlights: record→replay reproduces done on the failing-test fixture (deterministic), batch eval scoring, divergence detection, OTel span tree via an in-memory exporter, HITL pause→approve→resume, trace persisted with model+verify events.

Notes

  • OTel export is deployment config (register an OTLP/Jaeger/Tempo exporter); the instrumentation + an in-memory-exporter test are included.
  • Postgres trace table + approved column are in the migration; verified live once Docker/Postgres is up (tested here via the in-memory/file stores).

🤖 Generated with Claude Code

Summary by CodeRabbit

Release Notes

New Features

  • Structured trace recording: Job execution now records detailed event traces (model responses, tool calls, skill loads, verification attempts).
  • HITL approval gate: Jobs can require human approval via require_approval flag and new approve command.
  • Trace replay & evaluation: Deterministically replay recorded traces to verify consistency; new eval command scores test cases.
  • OpenTelemetry integration: Traces expose structured spans for external observability platforms.
  • New CLI commands: trace (view job traces and costs), approve (grant approvals), eval (run evaluation suites).

hutusi added 6 commits June 18, 2026 19:36
- core/trace: Trace + TraceEvent (model_response, tool_call, skill_loaded,
  compaction, verify) + TraceResult; recordedResponses() for replay
- model_response events carry the full ModelResponse so a run is replayable
- JobSpec.require_approval (optional) for the HITL gate; schema regenerated
- runLoop emits model_response (with usage), tool_call, and compaction events
- runJob forwards loop events and emits skill_loaded (exact name@version+hash for
  required and model-invoked skills) and verify (per-criterion pass/evidence)
- onTrace threaded through RunLoopOptions + RunJobOptions
- tests assert event order, recorded model usage, and skill version capture
- Recorder collects trace events into a Trace (record = onTrace hook; finish seals)
- traceCost rolls up tokens + USD from model_response events
- emitSpans: OTel span tree (root job span + child span per step) with GenAI-style
  attributes; no-op without an SDK, ships when an exporter is registered
- formatTrace: human-readable trace + cost footer for the CLI/console
- tests incl. an in-memory OTel exporter asserting the emitted span tree
- @auriga/evals: ReplayProvider replays a trace's recorded model responses
  deterministically (no model calls); throws on divergence (more calls than recorded)
- runEval/runEvals: replay a batch of {spec, trace} cases through the real harness
  against fresh sandboxes and score (matches recorded state, verify passed, steps, cost)
- summarize() rolls up a batch; loadEvalCases() reads a disk suite
- tests: record→replay reproduces "done" on the failing-test fixture, batch
  summary, and divergence detection (replay exhausted)
- JobStore gains saveTrace/loadTrace (in-memory + file + postgres, with orphan
  guards + traces table/migration); FileJobStore.list now excludes .trace.json
- JobRecord.approved field across all stores
- runJob: ApprovalGate consulted before execution — when spec.require_approval
  and not approved, the job returns state "paused" (no work done)
- Worker records the run via a Recorder and persists the trace; builds the
  approval gate from the store; pause→approve(store.update)→resume→done
- tests: trace persisted (model+verify events), HITL pause/approve/resume,
  store trace round-trip + orphan rejection
- auriga trace <id>     print the recorded trace + cost rollup
- auriga approve <id>   grant HITL approval to a paused job
- auriga run <id>       run/resume an existing job (e.g. after approval)
- auriga eval <dir>     replay a suite of recorded traces and score them
- submit refactored to create + runWorker; usage + README updated for Phase 2
- smoke tests for the new commands
@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

@hutusi, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 41 minutes and 40 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more credits in the billing tab to continue.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits.

🚦 How do rate limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan refill rate.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, the refill rate gradually slows as usage increases. The highest same-day bursts are limited more strictly.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 7dde24e8-d71b-46bc-a991-1499b4783b5b

📥 Commits

Reviewing files that changed from the base of the PR and between a6194a1 and c392d2b.

📒 Files selected for processing (8)
  • migrations/0002_add_approved_and_traces.sql
  • packages/currus/src/job-runner.test.ts
  • packages/currus/src/job-runner.ts
  • packages/evals/src/load.ts
  • packages/evals/src/runner.test.ts
  • packages/evals/src/runner.ts
  • packages/habenae/src/postgres-store.ts
  • packages/habenae/src/worker.ts
📝 Walkthrough

Walkthrough

Implements Phase 2 of the Auriga platform: core Trace/TraceEvent types, trace event emission in the run loop and job-runner, a HITL approval gate (require_approval), observability utilities (OTel spans, cost rollup, trace formatter) in @capella, trace and approval persistence across all store backends, a new @auriga/evals package with a deterministic replay provider and batch eval runner, and four new CLI subcommands (run, approve, trace, eval).

Changes

Phase 2: Evals + Observability + HITL

Layer / File(s) Summary
Core trace type definitions and HITL job spec field
packages/core/src/trace/types.ts, packages/core/src/trace/index.ts, packages/core/src/index.ts, packages/core/src/job/spec.ts, packages/core/schema/job.schema.json, packages/core/src/trace/trace.test.ts
Introduces TraceEvent discriminated union, Trace, TraceResult, VerifiedCriterion, and recordedResponses helper; re-exports them through the core index; adds optional require_approval: boolean to JobSpecSchema, JobSpec, and the JSON schema.
Trace emission in run loop and job-runner with HITL gate
packages/currus/src/loop.ts, packages/currus/src/job-runner.ts, packages/currus/src/index.ts, packages/currus/src/trace-emit.test.ts
Extends RunLoopOptions with onTrace to emit model_response, tool_call, and compaction events; adds ApprovalGate interface and "paused" state to RunJobResult; wires onTrace into runLoop and emits skill_loaded and verify events after each phase; tests verify event sequence and skill trace content.
Capella observability utilities: Recorder, cost rollup, OTel spans, formatter
packages/capella/src/recorder.ts, packages/capella/src/rollup.ts, packages/capella/src/tracing.ts, packages/capella/src/format.ts, packages/capella/src/index.ts, packages/capella/package.json, packages/capella/src/observability.test.ts, packages/capella/src/tracing.test.ts
Adds Recorder to collect events and seal a Trace; TraceCost/traceCost for token/cost rollup; emitSpans to emit an OTel span tree; formatTrace for human-readable output; re-exports all from capella index; adds @opentelemetry/api and @opentelemetry/sdk-trace-base dependencies; tests cover all utilities.
Trace and approval persistence across all store backends
migrations/0001_init.sql, packages/habenae/src/types.ts, packages/habenae/src/memory-store.ts, packages/habenae/src/file-store.ts, packages/habenae/src/postgres-store.ts, packages/habenae/src/worker.ts, packages/habenae/package.json, packages/habenae/src/memory-store.test.ts, packages/habenae/src/worker.test.ts
Adds approved: boolean to JobRecord and saveTrace/loadTrace to JobStore; implements across all three stores; adds approved column and traces table in DB migration; wires Recorder into Worker.run() to persist the sealed trace before updating job state; tests cover trace round-trip and HITL two-phase run.
New @auriga/evals package: replay provider, runner, and loader
packages/evals/package.json, packages/evals/tsconfig.json, packages/evals/src/replay.ts, packages/evals/src/runner.ts, packages/evals/src/load.ts, packages/evals/src/index.ts, packages/evals/src/runner.test.ts
Creates ReplayProvider for deterministic FIFO replay from recordedResponses; implements runEval/runEvals/summarize for scoring replay outcomes; loadEvalCases to load .json eval fixtures; tests cover single replay, batch summary, and queue-exhaustion divergence.
CLI new subcommands: run, approve, trace, eval
packages/cli/src/main.ts, packages/cli/package.json, packages/cli/src/cli.test.ts
Adds run, approve, trace, and eval subcommands; consolidates job execution into runWorker with selectCliDriver; updates printUsage and requireJob error text; adds @auriga/evals dependency; tests verify new command presence in usage and error exits.
README Phase 2 documentation
README.md
Adds Phase 2 bullet to status section and extends CLI example with trace, approve, and eval commands.

Sequence Diagram(s)

sequenceDiagram
  participant CLI
  participant Worker
  participant runJob
  participant Recorder
  participant JobStore

  rect rgba(70, 130, 180, 0.5)
    note over CLI, JobStore: Job submission and HITL flow
    CLI->>Worker: run(jobId)
    Worker->>Recorder: new Recorder(jobId, model)
    Worker->>runJob: {onTrace: recorder.record, approvalGate}
    runJob->>JobStore: approvalGate.isApproved()
    JobStore-->>runJob: approved=false
    runJob-->>Worker: {state: "paused"}
    Worker->>JobStore: saveTrace(recorder.finish({state:"paused"}))
    Worker->>JobStore: update(jobId, {state:"paused"})
  end

  rect rgba(34, 139, 34, 0.5)
    note over CLI, JobStore: After approve command
    CLI->>JobStore: update(jobId, {approved: true})
    CLI->>Worker: run(jobId)
    Worker->>Recorder: new Recorder(jobId, model)
    Worker->>runJob: {onTrace: recorder.record, approvalGate}
    runJob->>JobStore: approvalGate.isApproved()
    JobStore-->>runJob: approved=true
    runJob->>runJob: execute loop, emit trace events
    runJob-->>Worker: {state: "done"}
    Worker->>JobStore: saveTrace(recorder.finish({state:"done"}))
    Worker->>JobStore: update(jobId, {state:"done"})
  end
Loading
sequenceDiagram
  participant CLI
  participant runEvals
  participant runEval
  participant ReplayProvider
  participant runJob

  CLI->>runEvals: loadEvalCases(dir) → cases
  loop for each EvalCase
    runEvals->>runEval: {spec, trace}, driver
    runEval->>ReplayProvider: new ReplayProvider(trace)
    runEval->>runJob: {provider: ReplayProvider, spec}
    loop per model call
      runJob->>ReplayProvider: complete(req)
      ReplayProvider-->>runJob: next recorded ModelResponse
    end
    runJob-->>runEval: RunJobResult
    runEval-->>runEvals: EvalScore{matches, verify_passed, cost_usd}
  end
  runEvals->>runEvals: summarize(scores)
  runEvals-->>CLI: {scores, summary}
  CLI->>CLI: print results, set exitCode=2 if not all matched
Loading

Estimated code review effort

🎯 5 (Critical) | ⏱️ ~120 minutes

Poem

🐇 Hoppity-hop through the trace event queue,
Each span gets a name, each token its due.
The HITL gate opens when approved turns true,
And evals replay what the model once knew.
Cost rolled up neatly, OTel shining bright—
Phase 2 has landed, the rabbit's delight! ✨

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 29.73% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title clearly summarizes the main change: introducing Phase 2 with evals and observability features including traces, OpenTelemetry, replay, and HITL approval workflows.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch phase-2-evals-observability

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
migrations/0001_init.sql (1)

4-17: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Do not rely on Line 10 in 0001_init.sql for upgrades; add a forward migration.

On existing Phase 1 databases, jobs already exists, so this create table if not exists jobs (...) block will not add approved. That leaves runtime code expecting jobs.approved with a schema mismatch after deploy. Keep 0001 as baseline for fresh installs, and add a new migration that explicitly alters existing schemas.

Proposed migration (new file)
-- migrations/0002_add_approved_and_traces.sql
alter table jobs
  add column if not exists approved boolean not null default false;

create table if not exists traces (
  job_id     text primary key references jobs(id) on delete cascade,
  data       jsonb not null,
  updated_at timestamptz not null default now()
);
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@migrations/0001_init.sql` around lines 4 - 17, The create table if not exists
statement in the jobs table definition will not add the approved column to
existing Phase 1 databases where the jobs table already exists, causing a schema
mismatch. Create a new migration file (0002_add_approved_and_traces.sql) that
explicitly alters the jobs table to add the approved column using ALTER TABLE
with "add column if not exists" clause, ensuring both fresh installs and
existing databases have the correct schema after deployment. Keep the current
0001_init.sql as the baseline for fresh installs.
packages/habenae/src/worker.ts (1)

36-48: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Check approval before sandbox initialization.

Line 37 creates the sandbox before the approval gate at Line 48 is evaluated. For require_approval jobs, this can fail before pausing (for example, unsupported workspace seeding), which violates the HITL “pause first” behavior. Short-circuit unapproved jobs before seedFor/sandbox creation.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/habenae/src/worker.ts` around lines 36 - 48, The approval gate check
for `require_approval` jobs needs to happen before sandbox initialization to
comply with "pause first" behavior. Before calling seedFor and
this.opts.sandboxDriver.create, first retrieve the job approval status from the
store using store.get(jobId) and check if it's approved. Only proceed with
creating the sandbox using this.opts.sandboxDriver.create(seedFor(record,
checkpoint)) if the job is either approved or approval is not required. This
ensures that unapproved jobs pause for approval before attempting any
resource-intensive operations like sandbox creation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/currus/src/job-runner.ts`:
- Around line 219-223: The deduplication logic in the loop that processes
resolver.loadedSkills() uses only skill.name as the dedupe key in
recordedSkills, which causes different versions of the same skill to be dropped
from traces. Modify the deduplication to use a composite key that combines skill
name with its version and hash (like name@version+hash) instead of just
skill.name when checking recordedSkills.has() and recordedSkills.add() to ensure
all distinct skill versions are captured in the trace.
- Around line 152-155: The approval gate validation in the conditional check at
line 153 uses AND logic that allows execution to proceed when require_approval
is true but opts.approvalGate is not provided. Fix this by restructuring the
condition so that if spec.require_approval is true, the job pauses when either
opts.approvalGate is missing OR when the gate exists but isApproved() returns
false. This ensures that approval-required jobs cannot bypass the approval
requirement simply by not providing an approval gate.

In `@packages/evals/src/load.ts`:
- Around line 14-15: The trace object loaded from the file on line 14 is cast to
the Trace type without validation, causing invalid traces to be accepted and
fail later with unclear errors. After parsing the JSON and before pushing to
cases, add validation logic that checks the required trace fields exist and have
the correct shape. If validation fails, throw an error that includes the
filename (from the file variable) to make debugging easier. This validation
should happen right after the JSON.parse call and before the cases.push on line
15.

In `@packages/evals/src/runner.test.ts`:
- Around line 112-114: The test for loadEvalCases is currently importing from
the internal module ./load instead of from the public barrel export at ./index
or package root. This means the test would still pass even if loadEvalCases is
accidentally removed from the public API. Change the import statement in the
test to import from ./index instead of ./load, so that the test validates
loadEvalCases is properly exported through the intended public API surface.

In `@packages/evals/src/runner.ts`:
- Around line 39-42: The seedFor function silently converts any workspace kind
that is not "dir" to "empty", which can cause test replays to use the wrong
initial state. Replace the ternary operator with explicit error handling: check
if the workspace kind is "dir" and return the appropriate seed, or throw an
error for any unsupported or unexpected workspace kinds instead of silently
defaulting to empty.
- Around line 84-85: The sandbox.destroy() call in the finally block of the
runEval function can throw an error, which will reject the entire runEval
promise and stop batch execution in runEvals() even if a score was already
computed. Wrap the await sandbox.destroy() statement in a try-catch block to
handle any cleanup failures locally, logging or suppressing the error so that
sandbox teardown failures do not abort the entire batch evaluation run.

---

Outside diff comments:
In `@migrations/0001_init.sql`:
- Around line 4-17: The create table if not exists statement in the jobs table
definition will not add the approved column to existing Phase 1 databases where
the jobs table already exists, causing a schema mismatch. Create a new migration
file (0002_add_approved_and_traces.sql) that explicitly alters the jobs table to
add the approved column using ALTER TABLE with "add column if not exists"
clause, ensuring both fresh installs and existing databases have the correct
schema after deployment. Keep the current 0001_init.sql as the baseline for
fresh installs.

In `@packages/habenae/src/worker.ts`:
- Around line 36-48: The approval gate check for `require_approval` jobs needs
to happen before sandbox initialization to comply with "pause first" behavior.
Before calling seedFor and this.opts.sandboxDriver.create, first retrieve the
job approval status from the store using store.get(jobId) and check if it's
approved. Only proceed with creating the sandbox using
this.opts.sandboxDriver.create(seedFor(record, checkpoint)) if the job is either
approved or approval is not required. This ensures that unapproved jobs pause
for approval before attempting any resource-intensive operations like sandbox
creation.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e19e8cfb-6a1c-4a59-a1ce-8bf9c07d6d54

📥 Commits

Reviewing files that changed from the base of the PR and between 64bc2d6 and a6194a1.

⛔ Files ignored due to path filters (1)
  • bun.lock is excluded by !**/*.lock
📒 Files selected for processing (38)
  • README.md
  • migrations/0001_init.sql
  • packages/capella/package.json
  • packages/capella/src/format.ts
  • packages/capella/src/index.ts
  • packages/capella/src/observability.test.ts
  • packages/capella/src/recorder.ts
  • packages/capella/src/rollup.ts
  • packages/capella/src/tracing.test.ts
  • packages/capella/src/tracing.ts
  • packages/cli/package.json
  • packages/cli/src/cli.test.ts
  • packages/cli/src/main.ts
  • packages/core/schema/job.schema.json
  • packages/core/src/index.ts
  • packages/core/src/job/spec.ts
  • packages/core/src/trace/index.ts
  • packages/core/src/trace/trace.test.ts
  • packages/core/src/trace/types.ts
  • packages/currus/src/index.ts
  • packages/currus/src/job-runner.ts
  • packages/currus/src/loop.ts
  • packages/currus/src/trace-emit.test.ts
  • packages/evals/package.json
  • packages/evals/src/index.ts
  • packages/evals/src/load.ts
  • packages/evals/src/replay.ts
  • packages/evals/src/runner.test.ts
  • packages/evals/src/runner.ts
  • packages/evals/tsconfig.json
  • packages/habenae/package.json
  • packages/habenae/src/file-store.ts
  • packages/habenae/src/memory-store.test.ts
  • packages/habenae/src/memory-store.ts
  • packages/habenae/src/postgres-store.ts
  • packages/habenae/src/types.ts
  • packages/habenae/src/worker.test.ts
  • packages/habenae/src/worker.ts

Comment thread packages/currus/src/job-runner.ts Outdated
Comment thread packages/currus/src/job-runner.ts
Comment thread packages/evals/src/load.ts Outdated
Comment thread packages/evals/src/runner.test.ts Outdated
Comment thread packages/evals/src/runner.ts
Comment thread packages/evals/src/runner.ts Outdated
- runJob approval gate fails closed: require_approval with no/!approved gate pauses
  (previously a missing gate let execution proceed)
- worker short-circuits unapproved require_approval jobs BEFORE creating a sandbox
  (pause-first; no resource spend / seedFor failure before approval)
- skill_loaded dedupe keyed on name@version+content_hash (distinct versions recorded)
- evals: validate trace shape at load with the filename in the error; seedFor throws
  on unsupported workspace kind; sandbox teardown failure no longer aborts a batch
- evals test imports loadEvalCases from the public barrel
- migrations: keep 0001 as the Phase-1 baseline, add 0002 (ALTER add approved +
  traces table); migrate() is additive (add column if not exists) for existing DBs
- test: require_approval with no gate pauses (fail-closed)
@hutusi

hutusi commented Jun 18, 2026

Copy link
Copy Markdown
Contributor Author

Addressed the review in c392d2b — all 8 findings (6 inline + 2 in the body):

Major

  • Approval bypass (job-runner.ts): require_approval now fails closed — a missing gate or !approved pauses; execution can't proceed without explicit approval.
  • Worker pause-first (worker.ts): unapproved require_approval jobs short-circuit to paused before sandbox creation, so no resources are spent (and workspace seeding can't fail) pre-approval.
  • skill_loaded dedupe (job-runner.ts): keyed on name@version+content_hash so distinct versions are each recorded.
  • Eval trace validation (evals/load.ts): validates trace shape at load and includes the filename in the error.
  • Eval seedFor (evals/runner.ts): throws on an unsupported workspace kind instead of silently seeding empty (which could fake divergence).
  • Eval teardown (evals/runner.ts): sandbox.destroy() failures are swallowed so one cleanup error can't abort a batch or drop a score.
  • Migration (migrations/): reverted 0001 to the Phase-1 baseline and added 0002_add_approved_and_traces.sql (ALTER add approved + traces table); migrate() is now additive (add column if not exists) so existing DBs upgrade cleanly.

Minor

  • Eval test imports loadEvalCases from the public barrel (./index) to guard the exported API.

Added a regression test (require_approval with no gate → paused). Suite: 134 pass / 6 skip, typecheck clean.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant