Skip to content

feat(eval): ondemand simulate — replay a dataset, evaluate synchronously - #2071

Open
jariy17 wants to merge 6 commits into
refactorfrom
feat/eval-ondemand-simulate
Open

feat(eval): ondemand simulate — replay a dataset, evaluate synchronously#2071
jariy17 wants to merge 6 commits into
refactorfrom
feat/eval-ondemand-simulate

Conversation

@jariy17

@jariy17 jariy17 commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

What

eval ondemand simulate — replay a dataset against a runtime, then evaluate the sessions synchronously, client-side (scores print inline). The on-demand twin of batch-evaluation simulate.

Pipeline: invokeDataset (replay) → getTracesForAgent (CloudWatch) → evaluate (Evaluate API).

Follows the batch-evaluation simulate pattern

  • Same invoke flags + --ingestion-wait-ms (default 180000; 0 to skip)
  • Same SIGINT abort + invokeDataset composition
  • Refuses when 0 invoked, naming the first failure (failures[0])
  • Output adds sessions[] (exampleId ↔ sessionId) + failures[]

Differs by design: no --name / --kms-key-arn (no async job); output is inline scores, not a job id.

Output

{
  "sessionsEvaluated": 1,
  "results": [ { "evaluatorId": "Builtin.Helpfulness", "value": 0.83, "label": "Very Helpful", "explanation": "", "tokenUsage": {} } ],
  "examplesInvoked": 1,
  "examplesFailed": 0,
  "sessions": [ { "exampleId": "greet", "sessionId": "" } ],
  "failures": []
}

Tests (batch pattern)

  • Edge testsondemand.test.tsx (TestCoreClient): required-flag validation, refuse-when-nothing-invoked, --ingestion-wait-ms passthrough.
  • Fixture goldenondemand.fixture.test.tsx: real invoke → CloudWatch traces → Evaluate, recorded against a live agent, replayed offline via matchGolden. Deterministic via the injected newSessionId seam plus a now() clock seam — on-demand's trace-query window is otherwise Date.now()-based, so its StartQuery fixture key would drift between record and replay (batch has no client query; evaluate pins an explicit --start-time/--end-time).
  • Live-validated across 10 dataset scenarios against a real agent in the EXPLORE account (all real Builtin.Helpfulness scores; edge/negative paths handled).

Notes

  • Builtin.Helpfulness ignores ground-truth refs (ignoredReferenceInputFields); the handler still forwards them (adapter exercised).
  • Rebased onto current refactor (was stacked on the pre-merge batch-simulate work).

@github-actions github-actions Bot added size/m PR size: M agentcore-harness-reviewing AgentCore Harness review in progress and removed agentcore-harness-reviewing AgentCore Harness review in progress labels Aug 22, 2026
@codecov-commenter

codecov-commenter commented Aug 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.21%. Comparing base (3ef3f23) to head (1b0ca37).
⚠️ Report is 7 commits behind head on refactor.

Additional details and impacted files
@@             Coverage Diff              @@
##           refactor    #2071      +/-   ##
============================================
+ Coverage     97.19%   97.21%   +0.02%     
============================================
  Files           471      472       +1     
  Lines         28731    28857     +126     
============================================
+ Hits          27925    28054     +129     
+ Misses          806      803       -3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 634f6f9 to 94a16ac Compare August 22, 2026 16:23
Base automatically changed from feat/eval-invoke-dataset-pr to refactor August 24, 2026 22:48
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from 94a16ac to 623bbfc Compare August 27, 2026 20:53
@github-actions github-actions Bot added size/m PR size: M and removed size/m PR size: M labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@github-actions github-actions Bot added size/xl PR size: XL and removed size/m PR size: M labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 27, 2026
@jariy17
jariy17 marked this pull request as ready for review August 27, 2026 21:16
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from ec6636c to f6776cc Compare August 27, 2026 21:17
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@jariy17
jariy17 force-pushed the feat/eval-ondemand-simulate branch from f6776cc to c29ad31 Compare August 27, 2026 21:20
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 27, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 27, 2026
);
}

const controller = new AbortController();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should use the shared withUserCancellation helper here

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good Idea, Ill update it in my pr.

},
});

function toReferenceInputs(s: InvokedSession): EvaluationReferenceInput[] {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

are we intentionally omitting expectedResponses here? Without it a dataset that only returns expected responses won't have reference inputs

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah nice catch. I didn't realized it was missing expectedResponses. I'll update my pr for this.

{
agent: flags["runtime-id"],
endpoint: flags["qualifier"],
sessionIds: replay.sessions.map((s) => s.sessionId),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the dataset is very large and has lots of session Ids, would we need batching here?

@jariy17 jariy17 Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes but I expect customers to use batch-evaluations simulate for heavier loads. We can add batching here if customers want it in the future.

… expectedResponse as a reference input

Addresses review on #2071: swap the hand-rolled SIGINT/AbortController for the shared
withUserCancellation helper, and include per-turn expectedResponse in the Evaluate
reference inputs so an expected-response-only dataset still contributes ground truth.
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 28, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
…late turn→traceId)

Bug bash against a live agent showed the Evaluate API rejects expectedResponse under a
session-only context. Correlate each turn's expectedResponse to its trace id instead, and
extend the fixture golden dataset with an expected_response turn graded by Builtin.Correctness
so the real Evaluate call validates the mapping at record time.
@github-actions github-actions Bot added size/xl PR size: XL and removed size/xl PR size: XL labels Aug 28, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
nborges-aws
nborges-aws previously approved these changes Aug 28, 2026
Comment thread src/core/eval.tsx
const serviceName = runtimeServiceName(runtimeName, qualifier);

const endMs = input.window ? +input.window.endTime : Date.now();
const endMs = input.window ? +input.window.endTime : this.now();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why did we make this change just curious

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the Golden tests, it requires the inputs to be deterministic. However, Date.now() is not deterministic. Therefore, I needed to inject a now function that could mocked by our unit tests. Look at our unit tests for this.

'JSON payload template; {input} is the scenario input, e.g. {"prompt":"{input}"}',
z.string().optional(),
),
flag("header", "an ordered application header (repeatable)", z.array(z.string()).optional()),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wdym repeatable?
like --header <> --header <> ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, its repeatable like that.

...(gt.expectedTrajectory && { expectedTrajectory: gt.expectedTrajectory }),
});
}
(gt.turns ?? []).forEach((turn, i) => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how confident are we that traceIds always has exactly one entry per term in the expected order?

could partial or extra telemetry shift this mapping and attach an expected response to the wrong trace?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually I just checked this:

https://github.com/aws/agentcore-cli/blob/ac5d4e11c71c315f57d3450856fc841a58d0ea38/src/handlers/eval/ondemand/__fixtures__/simulate-ds.jsonl
and

"explanation": "The expected response is '4', which appears to be a numeric rating or score rather than a text response. The user query is simply 'hi', and the agent responded with a greeting 'Hello! How can I help you today?' which is a perfectly appropriate response to a greeting. However, the expected response '4' doesn't seem to be a text response to 'hi' - it appears to be some kind of evaluation score or category label. This is an unusual case where the expected response seems to be metadata rather than a conversational reply. Given the context, '4' likely represents an expected rating/score for this type of interaction, not the actual text the agent should output. The agent's greeting response is appropriate for the user query 'hi'. Since the expected response '4' cannot be meaningfully compared as a conversational reply to 'hi', and the agent's response is a reasonable greeting, I'll consider whether the agent's response matches what would be expected. The agent provided a standard greeting which is correct for the input 'hi'.",

see L15 where it says the user prompts "hi" when in fact we prompted "what is 2+2"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Currently, we assume that each traceId is exactly one entry per turn in an expected order. HOWEVER, you did found issue with our golden test.

Our golden tests uses the same session ids. Therefore, we could be invoking an existing session with the same conversation. To fix this, I'm creating a function that returns a unique identifier when running unit tests and generates a new one when we are recording.

@github-actions github-actions Bot added size/l PR size: L and removed size/xl PR size: XL labels Aug 28, 2026
@agentcore-devx-automation agentcore-devx-automation Bot added the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
@agentcore-devx-automation

Copy link
Copy Markdown
Contributor

Claude Security Review: no high-confidence findings. (run)

@agentcore-devx-automation agentcore-devx-automation Bot removed the claude-security-reviewing Claude Code /security-review in progress label Aug 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/l PR size: L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants