eval: skill comparison script - #3070
Conversation
update sample output from comparison run
collect artifacts of multiple stimuli
Details# 🔍 Token Analysis Report
fatal: path '.github/skills/analyze-comparison-tests/SKILL.md' exists on disk, but not in 'origin/main' 📊 Token Change ReportComparing Summary
Changed Files
📊 Token Limit Check ReportChecked: 850 files
|
| File | Tokens | Limit | Over By |
|---|---|---|---|
.github/skills/analyze-comparison-tests/SKILL.md |
579 | 500 | +79 |
.github/skills/analyze-skill-issues/SKILL.md |
2109 | 500 | +1609 |
.github/skills/analyze-test-run/SKILL.md |
2471 | 500 | +1971 |
.github/skills/file-test-bug/SKILL.md |
628 | 500 | +128 |
.github/skills/sensei/SKILL.md |
861 | 500 | +361 |
.github/skills/skill-authoring/SKILL.md |
839 | 500 | +339 |
.github/skills/vally-eval/SKILL.md |
1737 | 500 | +1237 |
plugins/azure-kusto-graph-skills/skills/azure-kusto-graph/SKILL.md |
4810 | 500 | +4310 |
plugins/azure-kusto-graph-skills/skills/azure-kusto-irql/SKILL.md |
2625 | 500 | +2125 |
plugins/azure-kusto-graph-skills/skills/azure-kusto-irql-graph/SKILL.md |
4781 | 500 | +4281 |
plugins/azure-kusto-graph-skills/skills/azure-kusto-irql-graph/references/DEPLOY_IRQL_FUNCTIONS.md |
2854 | 2000 | +854 |
plugins/azure-skills/changelog-2026-07-24.md |
4226 | 2000 | +2226 |
plugins/azure-skills/skills/airunway-aks-setup/SKILL.md |
1025 | 500 | +525 |
plugins/azure-skills/skills/appinsights-instrumentation/SKILL.md |
937 | 500 | +437 |
plugins/azure-skills/skills/azure-ai/SKILL.md |
820 | 500 | +320 |
plugins/azure-skills/skills/azure-aigateway/SKILL.md |
1261 | 500 | +761 |
plugins/azure-skills/skills/azure-aigateway/references/policies.md |
2342 | 2000 | +342 |
plugins/azure-skills/skills/azure-app-onboard/SKILL.md |
1623 | 500 | +1123 |
plugins/azure-skills/skills/azure-app-onboard/deploy/SKILL.md |
1816 | 500 | +1316 |
plugins/azure-skills/skills/azure-app-onboard/deploy/references/code-deployment-container-apps.md |
2140 | 2000 | +140 |
plugins/azure-skills/skills/azure-app-onboard/deploy/references/deploy-checklist-template.md |
2564 | 2000 | +564 |
plugins/azure-skills/skills/azure-app-onboard/prepare/SKILL.md |
1566 | 500 | +1066 |
plugins/azure-skills/skills/azure-app-onboard/prepare/references/pricing-guide-services.md |
2142 | 2000 | +142 |
plugins/azure-skills/skills/azure-app-onboard/prepare/references/service-mapping.md |
2022 | 2000 | +22 |
plugins/azure-skills/skills/azure-app-onboard/prepare/references/sku-quota-validation.md |
2278 | 2000 | +278 |
plugins/azure-skills/skills/azure-app-onboard/references/session-protocol.md |
2026 | 2000 | +26 |
plugins/azure-skills/skills/azure-app-onboard/scaffold/SKILL.md |
3592 | 500 | +3092 |
plugins/azure-skills/skills/azure-app-onboard/scaffold/references/bicep-container-apps.md |
2354 | 2000 | +354 |
plugins/azure-skills/skills/azure-app-onboard/scaffold/references/self-review-checklist.md |
2269 | 2000 | +269 |
plugins/azure-skills/skills/azure-app-onboard/scaffold/references/subagent-iac-gen.md |
2406 | 2000 | +406 |
plugins/azure-skills/skills/azure-app-onboard-prereq/SKILL.md |
2517 | 500 | +2017 |
plugins/azure-skills/skills/azure-app-onboard-prereq/references/dependency-compatibility.md |
2239 | 2000 | +239 |
plugins/azure-skills/skills/azure-cloud-migrate/SKILL.md |
1085 | 500 | +585 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/container-apps/cloudrun-deployment-guide.md |
2029 | 2000 | +29 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/container-apps/deployment-guide.md |
2458 | 2000 | +458 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/container-apps/fargate-deployment-guide.md |
2587 | 2000 | +587 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/container-apps/spring-deployment-guide.md |
3871 | 2000 | +1871 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/functions/lambda-to-functions.md |
2600 | 2000 | +600 |
plugins/azure-skills/skills/azure-cloud-migrate/references/services/functions/runtimes/javascript.md |
2181 | 2000 | +181 |
plugins/azure-skills/skills/azure-compliance/SKILL.md |
1188 | 500 | +688 |
plugins/azure-skills/skills/azure-compute/SKILL.md |
657 | 500 | +157 |
plugins/azure-skills/skills/azure-compute/workflows/essential-machine-management/references/emm-enable-flow.md |
2344 | 2000 | +344 |
plugins/azure-skills/skills/azure-deploy/SKILL.md |
1645 | 500 | +1145 |
plugins/azure-skills/skills/azure-deploy/references/pre-deploy-checklist.md |
4692 | 2000 | +2692 |
plugins/azure-skills/skills/azure-deploy/references/recipes/azd/errors.md |
4004 | 2000 | +2004 |
plugins/azure-skills/skills/azure-deploy/references/troubleshooting.md |
2038 | 2000 | +38 |
plugins/azure-skills/skills/azure-diagnostics/SKILL.md |
1596 | 500 | +1096 |
plugins/azure-skills/skills/azure-enterprise-infra-planner/SKILL.md |
911 | 500 | +411 |
plugins/azure-skills/skills/azure-enterprise-infra-planner/references/constraints/compute-apps.md |
2022 | 2000 | +22 |
plugins/azure-skills/skills/azure-kubernetes/SKILL.md |
2723 | 500 | +2223 |
plugins/azure-skills/skills/azure-kubernetes/azure-kubernetes-automatic-readiness/SKILL.md |
3690 | 500 | +3190 |
plugins/azure-skills/skills/azure-kusto/SKILL.md |
2152 | 500 | +1652 |
plugins/azure-skills/skills/azure-messaging/SKILL.md |
821 | 500 | +321 |
plugins/azure-skills/skills/azure-prepare/SKILL.md |
3145 | 500 | +2645 |
plugins/azure-skills/skills/azure-prepare/references/aspire.md |
4617 | 2000 | +2617 |
plugins/azure-skills/skills/azure-prepare/references/plan-template.md |
2560 | 2000 | +560 |
plugins/azure-skills/skills/azure-prepare/references/recipes/azd/aspire.md |
2275 | 2000 | +275 |
plugins/azure-skills/skills/azure-prepare/references/recipes/azd/terraform.md |
3555 | 2000 | +1555 |
plugins/azure-skills/skills/azure-prepare/references/research.md |
2196 | 2000 | +196 |
plugins/azure-skills/skills/azure-prepare/references/resources-limits-quotas.md |
3322 | 2000 | +1322 |
plugins/azure-skills/skills/azure-prepare/references/security.md |
2147 | 2000 | +147 |
plugins/azure-skills/skills/azure-prepare/references/services/functions/bicep.md |
3043 | 2000 | +1043 |
plugins/azure-skills/skills/azure-prepare/references/services/functions/templates/recipes/composition.md |
2813 | 2000 | +813 |
plugins/azure-skills/skills/azure-prepare/references/services/functions/terraform.md |
3404 | 2000 | +1404 |
plugins/azure-skills/skills/azure-quotas/SKILL.md |
3006 | 500 | +2506 |
plugins/azure-skills/skills/azure-quotas/references/commands.md |
2644 | 2000 | +644 |
plugins/azure-skills/skills/azure-reliability/SKILL.md |
5922 | 500 | +5422 |
plugins/azure-skills/skills/azure-reliability/references/configure-multi-region.md |
4729 | 2000 | +2729 |
plugins/azure-skills/skills/azure-reliability/references/services/app-service/reliability.md |
2591 | 2000 | +591 |
plugins/azure-skills/skills/azure-resource-lookup/SKILL.md |
1367 | 500 | +867 |
plugins/azure-skills/skills/azure-resource-visualizer/SKILL.md |
2122 | 500 | +1622 |
plugins/azure-skills/skills/azure-storage/SKILL.md |
1228 | 500 | +728 |
plugins/azure-skills/skills/azure-upgrade/SKILL.md |
1542 | 500 | +1042 |
plugins/azure-skills/skills/azure-upgrade/references/languages/java/INSTRUCTION.md |
2893 | 2000 | +893 |
plugins/azure-skills/skills/azure-upgrade/references/languages/java/package-specific/com.microsoft.azure.management.md |
2428 | 2000 | +428 |
plugins/azure-skills/skills/azure-upgrade/references/languages/java/templates/PLAN_TEMPLATE.md |
2411 | 2000 | +411 |
plugins/azure-skills/skills/azure-upgrade/references/languages/java/templates/PROGRESS_TEMPLATE.md |
2315 | 2000 | +315 |
plugins/azure-skills/skills/azure-upgrade/references/languages/java/templates/SUMMARY_TEMPLATE.md |
2190 | 2000 | +190 |
plugins/azure-skills/skills/azure-upgrade/references/services/functions/automation.md |
3463 | 2000 | +1463 |
plugins/azure-skills/skills/azure-upgrade/references/services/functions/consumption-to-flex.md |
2773 | 2000 | +773 |
plugins/azure-skills/skills/azure-validate/SKILL.md |
897 | 500 | +397 |
plugins/azure-skills/skills/entra-agent-id/SKILL.md |
3994 | 500 | +3494 |
plugins/azure-skills/skills/entra-app-registration/SKILL.md |
2058 | 500 | +1558 |
plugins/azure-skills/skills/entra-app-registration/references/api-permissions.md |
2545 | 2000 | +545 |
plugins/azure-skills/skills/entra-app-registration/references/cli-commands.md |
2211 | 2000 | +211 |
plugins/azure-skills/skills/entra-app-registration/references/console-app-example.md |
2752 | 2000 | +752 |
plugins/azure-skills/skills/entra-app-registration/references/oauth-flows.md |
2375 | 2000 | +375 |
plugins/azure-skills/skills/microsoft-foundry/SKILL.md |
6501 | 500 | +6001 |
plugins/azure-skills/skills/microsoft-foundry/finetuning/SKILL.md |
1376 | 500 | +876 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/azd-guidance/references/azd-ai-cli.md |
2230 | 2000 | +230 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/create/create-hosted.md |
7598 | 2000 | +5598 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/create/quick-start-hosted.md |
5635 | 2000 | +3635 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/create/references/foundry-tool-catalog.md |
10922 | 2000 | +8922 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/create/references/local-run.md |
2443 | 2000 | +443 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/create/references/use-toolbox-in-hosted-agent.md |
2643 | 2000 | +643 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/deploy/deploy.md |
5033 | 2000 | +3033 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/eval-datasets/eval-datasets.md |
2863 | 2000 | +863 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/eval-datasets/references/generate-seed-dataset.md |
2212 | 2000 | +212 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/eval-datasets/references/trace-to-dataset.md |
4325 | 2000 | +2325 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/invocations-ws/invocations-ws.md |
2652 | 2000 | +652 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/observe/observe.md |
3856 | 2000 | +1856 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/observe/references/continuous-eval.md |
3855 | 2000 | +1855 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/observe/references/evaluate-step.md |
2175 | 2000 | +175 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/observe/references/evaluation-suite-generation.md |
3134 | 2000 | +1134 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/routine/routine.md |
2032 | 2000 | +32 |
plugins/azure-skills/skills/microsoft-foundry/foundry-agent/trace/references/kql-templates.md |
2701 | 2000 | +701 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/SKILL.md |
1798 | 500 | +1298 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/capacity/SKILL.md |
1739 | 500 | +1239 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/customize/SKILL.md |
2236 | 500 | +1736 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/customize/references/customize-workflow.md |
3336 | 2000 | +1336 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/preset/SKILL.md |
1227 | 500 | +727 |
plugins/azure-skills/skills/microsoft-foundry/models/deploy-model/preset/references/preset-workflow.md |
5539 | 2000 | +3539 |
plugins/azure-skills/skills/microsoft-foundry/project/create/create-foundry-project.md |
2287 | 2000 | +287 |
plugins/azure-skills/skills/microsoft-foundry/quota/quota.md |
2289 | 2000 | +289 |
plugins/azure-skills/skills/microsoft-foundry/quota/references/capacity-planning.md |
2081 | 2000 | +81 |
plugins/azure-skills/skills/microsoft-foundry/references/agent-metadata-contract.md |
2217 | 2000 | +217 |
plugins/azure-skills/skills/python-appservice-deploy/SKILL.md |
688 | 500 | +188 |
Consider moving content to
references/subdirectories.
Automated token analysis. See skill authoring guidelines for best practices.
There was a problem hiding this comment.
Pull request overview
Adds a comparison-test development workflow to run integration tests across model/skill configurations, collect trajectories from Azure Storage, and document the process via a new meta-skill.
Changes:
- Add
compare:runto queue GitHub Actions integration runs across a model/with-skill matrix. - Add
compare:collectto download and organize run trajectories from themanual-integration-reportscontainer. - Extend integration workflows and test executor support for
MODEL_OVERRIDEandNO_SKILLS, plus add a new.github/skills/meta-skill for analysis.
Show a summary per file
| File | Description |
|---|---|
| tests/vally/vally-executor.ts | Allows model selection via MODEL_OVERRIDE for Vally-based runs. |
| tests/package.json | Adds compare:run and compare:collect scripts for comparison tooling. |
| tests/comparison/run-compare.ts | New script to enqueue a matrix of integration workflow runs and write a run map JSON. |
| tests/comparison/collect-artifacts.ts | New script to discover and download trajectory artifacts from Azure Storage for queued runs. |
| tests/.gitignore | Ignores downloaded comparison-artifacts/ output directory. |
| .github/workflows/test-azure-deploy.yml | Adds no-skills input and passes NO_SKILLS env through to tests. |
| .github/workflows/test-all-integration.yml | Makes model-override flexible (string) and adds no-skills support. |
| .github/skills/analyze-comparison-tests/SKILL.md | New meta-skill guiding artifact collection + analysis workflow. |
| .github/skills/analyze-comparison-tests/references/report-template.md | New report template for comparison findings. |
Review details
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Suppressed comments (1)
tests/comparison/collect-artifacts.ts:194
- Similarly, when no trajectory blobs are found for a stimuli, the script logs a warning and continues without marking the overall run as failed. This conflicts with the stated success criteria (exit 0 only when everything is collected).
if (blobList.length === 0) {
console.error(
`Warning: no trajectory blobs found for run ${runId}, stimuli ${stimuliPart}.`
);
continue;
- Files reviewed: 9/9 changed files
- Comments generated: 10
- Review effort level: Lite
There was a problem hiding this comment.
Review details
Suppressed comments (8)
Previously missed (4) — in code that hasn't changed since the last review.
tests/comparison/run-compare.ts:87
queueComparisonRun()resolves withstdoutfromgh workflow run, butgh workflow rundoesn't reliably print a run URL/ID. Downstream,collect-artifacts.tsassumesrun.runis a GitHub Actions URL and extracts the run ID viasplit('/').pop(), which will break if the output is the default human-readable confirmation text.
Consider resolving the run URL explicitly after dispatch (e.g., query gh run list --workflow <id> --branch <branch> --json url,createdAt,event and pick the newest workflow_dispatch run) and store that URL in the output JSON.
async function queueComparisonRun(branch: string, skill: SkillRef, option: CompareOption): Promise<string> {
const skillsInput = `${skill.pluginDirname}/${skill.name}`;
const args = ["workflow", "run", integrationTestWorkflowId, "--repo", repo, "--ref", branch, "--json"];
const inputs = JSON.stringify({
skills: skillsInput,
"model-override": option.model,
// Note: gh cli use string values for boolean input
"no-skills": !option.withSkill ? "true" : "false"
});
tests/comparison/run-compare.ts:149
resultsis initialized as an untyped empty array, so TypeScript will inferany[]here and you lose type-safety forBranchOutput.runsentries.
const results = [];
.github/skills/analyze-comparison-tests/references/report-template.md:3
- Placeholders in repo skill docs are typically represented with angle brackets (e.g.,
<skill-name>) to make them unambiguous in Markdown. Using{...}here can be mistaken for literal curly-brace syntax.
Skill: {plugin dirname/skill name}
.github/skills/analyze-comparison-tests/references/report-template.md:13
- For consistency with other templates in this repo, prefer angle-bracket placeholders for the headings and answer blocks as well (so users don’t copy
{User question 1}literally).
### {User question 1}
{Answer to question 1}
### {User question 2}
{Answer to question 2}
tests/comparison/collect-artifacts.ts:143
- The header comment says exit code 1 covers "no blobs found", but when no stimuli directories are discovered the script only warns and continues. This can produce an exit code 0 even though artifacts are missing for a run.
if (stimuliSet.size === 0) {
console.warn(
`Warning: no stimuli directories discovered for run ${runId}`
);
continue;
}
tests/comparison/collect-artifacts.ts:183
- When no trajectory blobs are found, the message is labeled "Warning" but is logged via
console.error, and the script does not mark the run as failed. This again conflicts with the documented exit code behavior ("no blobs found" -> exit 1).
if (blobList.length === 0) {
console.error(
`Warning: no trajectory blobs found for run ${runId}, stimuli ${stimuliPart}.`
);
continue;
}
.github/skills/analyze-comparison-tests/SKILL.md:10
- Meta-skills under
.github/skills/are expected to include a dedicated "When to Use" section (not only aWHEN TO USE:phrase in frontmatter). Without it, the activation scenarios and routing intent are unclear.
# Steps
.github/skills/analyze-comparison-tests/SKILL.md:50
- Meta-skills under
.github/skills/should include an Error Handling table (Error / Message / Remediation). Right now the skill doesn't describe what to do when the input JSON is missing/invalid or when Azure blob discovery/download fails.
Each `<branch-name>/<stimulus-name>/<model>-with-skill` or `<branch-name>/<stimulus-name>/<model>-without-skill` directory contains the test run trajectories for that stimulus and model on that branch, with or without skills. Each trajectory is a markdown file that records user prompts, tool call requests, tool execution results, assistant responses that happened during the run. It also contains statistics such as token usage and turns. Based on the trajectories, answer the user's questions for each test run. Generate a report following the [report-template](./references/report-template.md) to show your answers.
- Files reviewed: 9/9 changed files
- Comments generated: 1
- Review effort level: Lite
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Review details
Suppressed comments (7)
Previously missed (5) — in code that hasn't changed since the last review.
tests/comparison/run-compare.ts:165
run-compare.tswrites its output JSON into the source directory (tests/comparison/). This makes it easy to accidentally commit generated files and is surprising when running fromtests/. Consider writing the file to the current working directory (or a dedicated output dir) and printing the output path.
const outputFilename = `comparison-runs-${new Date().toISOString().replace(/[:.]/g, "-")}.json`;
writeFileSync(path.resolve(__dirname, outputFilename), JSON.stringify(output, null, 2));
tests/comparison/run-compare.ts:142
- The
daterecorded in the output is computed when queuing runs, but the artifact upload prefix uses the UTC date at publish time in the GitHub Actions job. If a run starts after midnight UTC,collect-artifacts.tswill look under the wrong date prefix and miss artifacts.
const date = new Date().toISOString().slice(0, 10); // Get yyyy-mm-dd date string
const output: CompareRunOutput = {
skill: input.skill,
date: date,
results: []
tests/.gitignore:4
run-compare.tsproducescomparison-runs-*.jsonfiles that are generated artifacts. They should be gitignored to avoid accidental commits (especially if the script is run multiple times).
comparison-artifacts/
.github/skills/analyze-comparison-tests/SKILL.md:16
- Grammar: use “a JSON file” (not “an JSON file”). Also “should be” reads more naturally than “is supposed to be” here.
The user must provide an JSON file to correlate each comparison test run with the GitHub Actions run. The script expects one input argument as the path to this JSON file. The JSON input is supposed to be the JSON output when queuing the comparison test runs using the `npm run compare:run` command.
.github/skills/analyze-comparison-tests/SKILL.md:23
- Grammar: use “a directory” (not “an directory”).
The collect-artifacts script will download the test run artifacts to a directory named `comparison-artifacts` in the current working directory. Before executing the script, check if there is already such an directory. If so, skip executing the script and proceed to step 2.
tests/comparison/collect-artifacts.ts:150
- When no stimuli directories are discovered for a run, the script logs a warning and continues without marking the overall run as failed. This can exit 0 even though artifacts are missing, which contradicts the header comment (“0 = success (all runs collected)”).
if (stimuliSet.size === 0) {
console.warn(
`Warning: no stimuli directories discovered for run ${runId}`
);
continue;
}
tests/comparison/collect-artifacts.ts:197
- If no trajectory blobs are found for a stimuli directory, this is currently logged as a “Warning” but it neither sets
failednor usesconsole.warn. That can lead to a successful exit even though some trajectories were not downloaded.
const blobList = blobs.trim().split("\n").filter((line: string) => line);
if (blobList.length === 0) {
console.error(
`Warning: no trajectory blobs found for run ${runId}, stimuli ${stimuliPart}.`
);
continue;
}
- Files reviewed: 9/9 changed files
- Comments generated: 0 new
- Review effort level: Lite
Description
Add development tools for running skill comparison tests. The tools include two scripts compare:run and compare:collect, and a skill for orchestrating these two scripts using an LLM agent.
compare:run script queues one run for the target skill in the integration test workflow for each comparison options, which configures the model and whether to include skills. The user can specify branches other than "main" to test modified versions of the skill or eval suites.
compare:collect script consumes the output from compare:run script to download and organize the trajectories of the test runs. The users can use an LLM agent or review the trajectories themselves to extract information.
The analyze-comparison-tests skill provides instructions on how to use compare:collect to download artifacts and extract information.
Users can use the tools to do the following comparisons:
Checklist
cd tests && npm test)fix:,feat:,feature:,chore:,misc:,test:,eval:tests/,npm run test:integration -- <skill>ornpm run test:vally -- --skill <skill>)Related Issues