Phase 2 — evals + observability (traces, OTel, replay, HITL) - #3
Conversation
- core/trace: Trace + TraceEvent (model_response, tool_call, skill_loaded, compaction, verify) + TraceResult; recordedResponses() for replay - model_response events carry the full ModelResponse so a run is replayable - JobSpec.require_approval (optional) for the HITL gate; schema regenerated
- runLoop emits model_response (with usage), tool_call, and compaction events - runJob forwards loop events and emits skill_loaded (exact name@version+hash for required and model-invoked skills) and verify (per-criterion pass/evidence) - onTrace threaded through RunLoopOptions + RunJobOptions - tests assert event order, recorded model usage, and skill version capture
- Recorder collects trace events into a Trace (record = onTrace hook; finish seals) - traceCost rolls up tokens + USD from model_response events - emitSpans: OTel span tree (root job span + child span per step) with GenAI-style attributes; no-op without an SDK, ships when an exporter is registered - formatTrace: human-readable trace + cost footer for the CLI/console - tests incl. an in-memory OTel exporter asserting the emitted span tree
- @auriga/evals: ReplayProvider replays a trace's recorded model responses
deterministically (no model calls); throws on divergence (more calls than recorded)
- runEval/runEvals: replay a batch of {spec, trace} cases through the real harness
against fresh sandboxes and score (matches recorded state, verify passed, steps, cost)
- summarize() rolls up a batch; loadEvalCases() reads a disk suite
- tests: record→replay reproduces "done" on the failing-test fixture, batch
summary, and divergence detection (replay exhausted)
- JobStore gains saveTrace/loadTrace (in-memory + file + postgres, with orphan guards + traces table/migration); FileJobStore.list now excludes .trace.json - JobRecord.approved field across all stores - runJob: ApprovalGate consulted before execution — when spec.require_approval and not approved, the job returns state "paused" (no work done) - Worker records the run via a Recorder and persists the trace; builds the approval gate from the store; pause→approve(store.update)→resume→done - tests: trace persisted (model+verify events), HITL pause/approve/resume, store trace round-trip + orphan rejection
- auriga trace <id> print the recorded trace + cost rollup - auriga approve <id> grant HITL approval to a paused job - auriga run <id> run/resume an existing job (e.g. after approval) - auriga eval <dir> replay a suite of recorded traces and score them - submit refactored to create + runWorker; usage + README updated for Phase 2 - smoke tests for the new commands
|
Warning Review limit reached
More reviews will be available in 41 minutes and 40 seconds. Learn how PR review limits work. Your organization has run out of usage credits. Purchase more credits in the billing tab to continue. ⌛ How to resolve this issue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based credits. 🚦 How do rate limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan refill rate. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, the refill rate gradually slows as usage increases. The highest same-day bursts are limited more strictly. Please see our Fair Usage Limits Policy for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (8)
📝 WalkthroughWalkthroughImplements Phase 2 of the Auriga platform: core ChangesPhase 2: Evals + Observability + HITL
Sequence Diagram(s)sequenceDiagram
participant CLI
participant Worker
participant runJob
participant Recorder
participant JobStore
rect rgba(70, 130, 180, 0.5)
note over CLI, JobStore: Job submission and HITL flow
CLI->>Worker: run(jobId)
Worker->>Recorder: new Recorder(jobId, model)
Worker->>runJob: {onTrace: recorder.record, approvalGate}
runJob->>JobStore: approvalGate.isApproved()
JobStore-->>runJob: approved=false
runJob-->>Worker: {state: "paused"}
Worker->>JobStore: saveTrace(recorder.finish({state:"paused"}))
Worker->>JobStore: update(jobId, {state:"paused"})
end
rect rgba(34, 139, 34, 0.5)
note over CLI, JobStore: After approve command
CLI->>JobStore: update(jobId, {approved: true})
CLI->>Worker: run(jobId)
Worker->>Recorder: new Recorder(jobId, model)
Worker->>runJob: {onTrace: recorder.record, approvalGate}
runJob->>JobStore: approvalGate.isApproved()
JobStore-->>runJob: approved=true
runJob->>runJob: execute loop, emit trace events
runJob-->>Worker: {state: "done"}
Worker->>JobStore: saveTrace(recorder.finish({state:"done"}))
Worker->>JobStore: update(jobId, {state:"done"})
end
sequenceDiagram
participant CLI
participant runEvals
participant runEval
participant ReplayProvider
participant runJob
CLI->>runEvals: loadEvalCases(dir) → cases
loop for each EvalCase
runEvals->>runEval: {spec, trace}, driver
runEval->>ReplayProvider: new ReplayProvider(trace)
runEval->>runJob: {provider: ReplayProvider, spec}
loop per model call
runJob->>ReplayProvider: complete(req)
ReplayProvider-->>runJob: next recorded ModelResponse
end
runJob-->>runEval: RunJobResult
runEval-->>runEvals: EvalScore{matches, verify_passed, cost_usd}
end
runEvals->>runEvals: summarize(scores)
runEvals-->>CLI: {scores, summary}
CLI->>CLI: print results, set exitCode=2 if not all matched
Estimated code review effort🎯 5 (Critical) | ⏱️ ~120 minutes Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 6
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
migrations/0001_init.sql (1)
4-17:⚠️ Potential issue | 🟠 Major | ⚡ Quick winDo not rely on Line 10 in
0001_init.sqlfor upgrades; add a forward migration.On existing Phase 1 databases,
jobsalready exists, so thiscreate table if not exists jobs (...)block will not addapproved. That leaves runtime code expectingjobs.approvedwith a schema mismatch after deploy. Keep0001as baseline for fresh installs, and add a new migration that explicitly alters existing schemas.Proposed migration (new file)
-- migrations/0002_add_approved_and_traces.sql alter table jobs add column if not exists approved boolean not null default false; create table if not exists traces ( job_id text primary key references jobs(id) on delete cascade, data jsonb not null, updated_at timestamptz not null default now() );🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@migrations/0001_init.sql` around lines 4 - 17, The create table if not exists statement in the jobs table definition will not add the approved column to existing Phase 1 databases where the jobs table already exists, causing a schema mismatch. Create a new migration file (0002_add_approved_and_traces.sql) that explicitly alters the jobs table to add the approved column using ALTER TABLE with "add column if not exists" clause, ensuring both fresh installs and existing databases have the correct schema after deployment. Keep the current 0001_init.sql as the baseline for fresh installs.packages/habenae/src/worker.ts (1)
36-48:⚠️ Potential issue | 🟠 Major | ⚡ Quick winCheck approval before sandbox initialization.
Line 37 creates the sandbox before the approval gate at Line 48 is evaluated. For
require_approvaljobs, this can fail before pausing (for example, unsupported workspace seeding), which violates the HITL “pause first” behavior. Short-circuit unapproved jobs beforeseedFor/sandbox creation.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/habenae/src/worker.ts` around lines 36 - 48, The approval gate check for `require_approval` jobs needs to happen before sandbox initialization to comply with "pause first" behavior. Before calling seedFor and this.opts.sandboxDriver.create, first retrieve the job approval status from the store using store.get(jobId) and check if it's approved. Only proceed with creating the sandbox using this.opts.sandboxDriver.create(seedFor(record, checkpoint)) if the job is either approved or approval is not required. This ensures that unapproved jobs pause for approval before attempting any resource-intensive operations like sandbox creation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@packages/currus/src/job-runner.ts`:
- Around line 219-223: The deduplication logic in the loop that processes
resolver.loadedSkills() uses only skill.name as the dedupe key in
recordedSkills, which causes different versions of the same skill to be dropped
from traces. Modify the deduplication to use a composite key that combines skill
name with its version and hash (like name@version+hash) instead of just
skill.name when checking recordedSkills.has() and recordedSkills.add() to ensure
all distinct skill versions are captured in the trace.
- Around line 152-155: The approval gate validation in the conditional check at
line 153 uses AND logic that allows execution to proceed when require_approval
is true but opts.approvalGate is not provided. Fix this by restructuring the
condition so that if spec.require_approval is true, the job pauses when either
opts.approvalGate is missing OR when the gate exists but isApproved() returns
false. This ensures that approval-required jobs cannot bypass the approval
requirement simply by not providing an approval gate.
In `@packages/evals/src/load.ts`:
- Around line 14-15: The trace object loaded from the file on line 14 is cast to
the Trace type without validation, causing invalid traces to be accepted and
fail later with unclear errors. After parsing the JSON and before pushing to
cases, add validation logic that checks the required trace fields exist and have
the correct shape. If validation fails, throw an error that includes the
filename (from the file variable) to make debugging easier. This validation
should happen right after the JSON.parse call and before the cases.push on line
15.
In `@packages/evals/src/runner.test.ts`:
- Around line 112-114: The test for loadEvalCases is currently importing from
the internal module ./load instead of from the public barrel export at ./index
or package root. This means the test would still pass even if loadEvalCases is
accidentally removed from the public API. Change the import statement in the
test to import from ./index instead of ./load, so that the test validates
loadEvalCases is properly exported through the intended public API surface.
In `@packages/evals/src/runner.ts`:
- Around line 39-42: The seedFor function silently converts any workspace kind
that is not "dir" to "empty", which can cause test replays to use the wrong
initial state. Replace the ternary operator with explicit error handling: check
if the workspace kind is "dir" and return the appropriate seed, or throw an
error for any unsupported or unexpected workspace kinds instead of silently
defaulting to empty.
- Around line 84-85: The sandbox.destroy() call in the finally block of the
runEval function can throw an error, which will reject the entire runEval
promise and stop batch execution in runEvals() even if a score was already
computed. Wrap the await sandbox.destroy() statement in a try-catch block to
handle any cleanup failures locally, logging or suppressing the error so that
sandbox teardown failures do not abort the entire batch evaluation run.
---
Outside diff comments:
In `@migrations/0001_init.sql`:
- Around line 4-17: The create table if not exists statement in the jobs table
definition will not add the approved column to existing Phase 1 databases where
the jobs table already exists, causing a schema mismatch. Create a new migration
file (0002_add_approved_and_traces.sql) that explicitly alters the jobs table to
add the approved column using ALTER TABLE with "add column if not exists"
clause, ensuring both fresh installs and existing databases have the correct
schema after deployment. Keep the current 0001_init.sql as the baseline for
fresh installs.
In `@packages/habenae/src/worker.ts`:
- Around line 36-48: The approval gate check for `require_approval` jobs needs
to happen before sandbox initialization to comply with "pause first" behavior.
Before calling seedFor and this.opts.sandboxDriver.create, first retrieve the
job approval status from the store using store.get(jobId) and check if it's
approved. Only proceed with creating the sandbox using
this.opts.sandboxDriver.create(seedFor(record, checkpoint)) if the job is either
approved or approval is not required. This ensures that unapproved jobs pause
for approval before attempting any resource-intensive operations like sandbox
creation.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: e19e8cfb-6a1c-4a59-a1ce-8bf9c07d6d54
⛔ Files ignored due to path filters (1)
bun.lockis excluded by!**/*.lock
📒 Files selected for processing (38)
README.mdmigrations/0001_init.sqlpackages/capella/package.jsonpackages/capella/src/format.tspackages/capella/src/index.tspackages/capella/src/observability.test.tspackages/capella/src/recorder.tspackages/capella/src/rollup.tspackages/capella/src/tracing.test.tspackages/capella/src/tracing.tspackages/cli/package.jsonpackages/cli/src/cli.test.tspackages/cli/src/main.tspackages/core/schema/job.schema.jsonpackages/core/src/index.tspackages/core/src/job/spec.tspackages/core/src/trace/index.tspackages/core/src/trace/trace.test.tspackages/core/src/trace/types.tspackages/currus/src/index.tspackages/currus/src/job-runner.tspackages/currus/src/loop.tspackages/currus/src/trace-emit.test.tspackages/evals/package.jsonpackages/evals/src/index.tspackages/evals/src/load.tspackages/evals/src/replay.tspackages/evals/src/runner.test.tspackages/evals/src/runner.tspackages/evals/tsconfig.jsonpackages/habenae/package.jsonpackages/habenae/src/file-store.tspackages/habenae/src/memory-store.test.tspackages/habenae/src/memory-store.tspackages/habenae/src/postgres-store.tspackages/habenae/src/types.tspackages/habenae/src/worker.test.tspackages/habenae/src/worker.ts
- runJob approval gate fails closed: require_approval with no/!approved gate pauses (previously a missing gate let execution proceed) - worker short-circuits unapproved require_approval jobs BEFORE creating a sandbox (pause-first; no resource spend / seedFor failure before approval) - skill_loaded dedupe keyed on name@version+content_hash (distinct versions recorded) - evals: validate trace shape at load with the filename in the error; seedFor throws on unsupported workspace kind; sandbox teardown failure no longer aborts a batch - evals test imports loadEvalCases from the public barrel - migrations: keep 0001 as the Phase-1 baseline, add 0002 (ALTER add approved + traces table); migrate() is additive (add column if not exists) for existing DBs - test: require_approval with no gate pauses (fail-closed)
|
Addressed the review in c392d2b — all 8 findings (6 inline + 2 in the body): Major
Minor
Added a regression test (require_approval with no gate → paused). Suite: 134 pass / 6 skip, typecheck clean. |
Phase 2 — evals + observability
Builds on Phase 1 (on
main). Six focused commits. (Reopened; supersedes #2.)What's new
@auriga/core/trace) — every run records an orderedTraceofmodel_response(with usage),tool_call,skill_loaded(exactname@version+hash),compaction, andverifyevents.model_responseevents make a run replayable.JobSpec.require_approvaladded.@auriga/currus) —runLoop+runJobemit the full event stream via anonTracehook.@auriga/capella) —Recorderseals events into aTrace;traceCostrolls up tokens + USD;emitSpansproduces an OpenTelemetry span tree (root job span + child span per step, GenAI-style attributes — no-op until an exporter is registered);formatTracefor the CLI.@auriga/evals) —ReplayProviderreplays a trace's recorded responses deterministically (no model calls, throws on divergence);runEval/runEvalsreplay a batch through the real harness against fresh sandboxes and score (matches recorded state, verify passed, steps, cost);loadEvalCasesreads a disk suite.@auriga/habenae) —JobStore.saveTrace/loadTrace(in-memory + file + Postgres, with atracestable);JobRecord.approved; the Worker records + persists each run's trace and consults an approval gate — a job withrequire_approvalreturns statepauseduntil approved, then resumes todone.auriga trace <id>·approve <id>·run <id>·eval <dir>.Verification
bun run check— 133 pass / 6 skip, typecheck clean.doneon the failing-test fixture (deterministic), batch eval scoring, divergence detection, OTel span tree via an in-memory exporter, HITL pause→approve→resume, trace persisted with model+verify events.Notes
approvedcolumn are in the migration; verified live once Docker/Postgres is up (tested here via the in-memory/file stores).🤖 Generated with Claude Code
Summary by CodeRabbit
Release Notes
New Features
require_approvalflag and newapprovecommand.evalcommand scores test cases.trace(view job traces and costs),approve(grant approvals),eval(run evaluation suites).