Skip to content

Add bundle size & tree-shaking benchmark with CI regression guard - #281

Merged
DZakh merged 11 commits into
mainfrom
claude/bundle-tree-shaking-benchmarks-uq8hcl
Jul 4, 2026
Merged

DZakh merged 11 commits into
mainfrom
claude/bundle-tree-shaking-benchmarks-uq8hcl

Conversation

@DZakh

@DZakh DZakh commented Jul 1, 2026 •

Copy link
Copy Markdown
Owner

Adds a comprehensive bundle size benchmark to track Sury's tree-shaking effectiveness and prevent regressions, aligning with the project's third core goal (after DX and runtime performance).

Summary

Introduces bundle.bench.mjs, a new benchmark that measures minified + gzipped bundle sizes for six representative usage scenarios. Each scenario bundles only the APIs it imports, allowing the output size to serve as a tree-shaking measurement. The benchmark compares against a committed baseline snapshot and fails CI if any scenario drifts beyond ±1% tolerance, with results posted as a sticky PR comment.

Key Changes

  • New benchmark script (packages/sury/tests/bundle.bench.mjs):

    • Bundles six scenarios with esbuild (minified, tree-shaking enabled)
    • Measures min, gzip, and brotli sizes
    • Supports --update to re-baseline after intentional changes
    • Supports --tolerance=N to adjust the drift threshold
    • Generates GitHub-flavored markdown reports for CI
  • Baseline snapshot (packages/sury/tests/bundle-size.snapshot.json):

    • Captures esbuild version and baseline sizes for all scenarios
    • Demonstrates strong tree-shaking: string + parser is ~18% of total gzip size
  • CI integration (.github/workflows/ci.yml):

    • Runs benchmark after build (measures fresh Sury.res.mjs)
    • Posts sticky PR comment with results table via marocchino/sticky-pull-request-comment
    • Appends report to GitHub job summary
    • Fails if any scenario exceeds tolerance
  • Documentation (CONTRIBUTING.md):

    • Explains how to run the benchmark locally
    • Describes the tree-shaking measurement approach
    • Links to the benchmark implementation
  • Package scripts:

    • Added bench:bundle to packages/sury/package.json
    • Added benchmark:bundle to root package.json for consistency
  • Cleanup:

    • Removed unnest and compile from public exports (not yet implemented; were blocking tree-shaking measurements)
    • Updated .gitignore to exclude generated bundle-size.md

Implementation Details

The benchmark guards against regressions on the gzip metric (the most relevant for shipped code). Scenarios range from minimal (string + parser at 3.7 kB gzip) to comprehensive (total at 21.1 kB gzip), with intermediate cases covering objects, unions, recursion, and JSON Schema generation. The /*#__PURE__*/ annotations in src/S.js are critical to achieving the tree-shaking wins measured here.

https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c

Summary by CodeRabbit

  • New Features

    • Added bundle size and tree-shaking benchmarking with a report shown in CI and PR comments.
    • Added a new command to run the benchmark locally.
  • Bug Fixes

    • Removed a couple of runtime exports from the public bundle surface.
  • Documentation

    • Added guidance for checking bundle size changes and updating the baseline snapshot.

claude added 3 commits June 30, 2026 18:31
Sury's third core goal (after DX and runtime perf) is bundle size, but it
had no automated guard — only the manual bundlejs.com recipes in
CONTRIBUTING.md and the hand-keyed README table. This adds the size
analogue of the type-instantiation bench (types.bench.ts) and the runtime
bench (sury.bench.ts).

tests/bundle.bench.mjs bundles tiny per-feature entry points with esbuild
(already in the tree via vitest; now a direct devDep), minifies, and gzips
them. Because each entry imports only part of the API, the output size
measures how well unused code tree-shakes away — string+parser (3.66 kB
gzip) vs the full surface (20.64 kB gzip) is the headline win, and it
guards the /*#__PURE__*/ annotations in src/S.js against regression.

Sizes are pinned in tests/bundle-size.snapshot.json; check mode fails if
any scenario's gzip drifts beyond ±1%, with a --update flag to re-baseline
(mirrors the bench:types philosophy). Wired into the sury CI job after
build via `pnpm bench:bundle`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
Both were leftover entries in the Pack filesMapping that generate
`export var unnest = S.unnest` / `export var compile = S.compile` in the
facade, but neither symbol is exported from Sury.res.mjs — so the names
resolved to `undefined` and esbuild warned on every bundle. Neither is in
Sury.resi or S.d.ts. The `unnest` concept shipped as `compactColumns`
(still present); `S.compile` was replaced (the docs already reference it in
the past tense), so the docs need no change.

Removed from the source mapping (Pack.res), its committed compiled output
(Pack.res.mjs), and the generated facade (src/S.js); re-baselined the
bundle-size snapshot (total drops ~18 B gzip).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
The bundle-size guard previously only failed CI on drift; this surfaces the
numbers. The bench now renders a GitHub-flavoured table (via --markdown) and
appends it to $GITHUB_STEP_SUMMARY, and a sticky PR comment
(marocchino/sticky-pull-request-comment) keeps it updated on each push. The
comment step runs on `if: always()` so a regression stays visible inline, and
is best-effort (continue-on-error) since fork PRs get a read-only token; the
job gains `pull-requests: write` for it.

Each row shows min / gzip / brotli and the Δ vs the snapshot baseline
(✅ within tolerance, ⚠️ regression, 🔵 improvement), plus a tree-shaking
summary line. The generated bundle-size.md is gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
@coderabbitai

coderabbitai Bot commented Jul 1, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds a bundle-size/tree-shaking benchmark script with an esbuild-based scenario suite and JSON snapshot baseline, wires it into CI to post a sticky PR comment and job summary, adds supporting package/vitest/docs changes, and removes unnest and compile re-exports from the packaged runtime.

Changes

Bundle size benchmark and runtime export cleanup

Layer / File(s) Summary
Benchmark script core
packages/sury/tests/bundle.bench.ts
New benchmark script with file purpose notes, CLI arg parsing (--update, --tolerance, markdown path), scenario definitions exercising partial API surfaces, and an esbuild-based measure function computing minified/gzip/brotli sizes plus formatting/markdown helpers.
Benchmark execution and reporting
packages/sury/tests/bundle.bench.ts
Computes measurements for all scenarios, implements update mode (writes snapshot) and check mode (compares against baseline with tolerance), writes markdown/GITHUB_STEP_SUMMARY output, and exits with failure code on drift.
Baseline and package wiring
packages/sury/tests/bundle-size.snapshot.json, packages/sury/package.json, package.json, .gitignore, packages/sury/vitest.config.mjs
Adds the JSON snapshot baseline (esbuild version, scenario sizes), bench:bundle scripts at both root and package level, pins esbuild devDependency, excludes the new bench file from Vitest, and ignores the generated bundle-size.md report.
CI comment and docs
.github/workflows/ci.yml, CONTRIBUTING.md
Grants pull-requests: write permission, adds a bundle-size benchmark step producing a markdown report and a sticky PR comment step (best-effort, pull_request-only), guards coverage/Codecov on non-cancellation and bumps Codecov to v7, and documents the bundle-size workflow in CONTRIBUTING.
Remove unnest and compile exports
packages/sury/scripts/pack/Pack.res, packages/sury/scripts/pack/Pack.res.mjs, packages/sury/src/S.js
Removes unnest and compile symbol mappings from the pack script's filesMapping and removes their corresponding re-exports from S.js.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CI as CI Workflow
  participant BenchScript as bundle.bench.ts
  participant Esbuild as esbuild
  participant Snapshot as bundle-size.snapshot.json
  participant PRComment as Sticky PR Comment

  CI->>BenchScript: run pnpm bench:bundle --markdown=bundle-size.md
  BenchScript->>Esbuild: bundle each scenario
  Esbuild-->>BenchScript: minified output
  BenchScript->>BenchScript: compute gzip/brotli sizes
  BenchScript->>Snapshot: load baseline sizes
  Snapshot-->>BenchScript: baseline gzip values
  BenchScript->>BenchScript: compute diffs, mark failures beyond tolerance
  BenchScript-->>CI: write bundle-size.md, append GITHUB_STEP_SUMMARY, exit code
  CI->>PRComment: post/update sticky comment (always, best-effort)
Loading

Compact metadata

  • Related issues: None referenced
  • Related PRs: None referenced
  • Suggested labels: ci, benchmarking, breaking-change
  • Suggested reviewers: DZakh

🐰 A rabbit hopped through bundles small,
Trimmed two exports, watched sizes fall,
Esbuild hums, gzip counts true,
A sticky comment lands on cue,
Tree-shaken code, tidy for all.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: adding a bundle size and tree-shaking benchmark with CI regression protection.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/bundle-tree-shaking-benchmarks-uq8hcl

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jul 1, 2026 •

Copy link
Copy Markdown

📦 Bundle size & tree-shaking

Scenario min min+gzip brotli Tree-shaking Δ gzip vs baseline
total (export *) 56.39 kB 20.62 kB 18.40 kB — ✅ +5 B (+0.02%)
string + parser 8.03 kB 3.68 kB 3.35 kB 82% smaller ✅ +16 B (+0.43%)
object schema + parser 23.62 kB 9.73 kB 8.81 kB 53% smaller ✅ +51 B (+0.51%)
union + transform 23.79 kB 9.84 kB 8.92 kB 52% smaller ✅ +50 B (+0.50%)
recursive + array 23.96 kB 9.93 kB 9.03 kB 52% smaller ✅ +50 B (+0.49%)
toJSONSchema 30.49 kB 12.09 kB 10.92 kB 41% smaller ✅ +40 B (+0.32%)

Baseline: esbuild 0.28.1, ±1% gzip tolerance. Re-baseline with pnpm bench:bundle --update.

@codspeed

codspeed Bot commented Jul 1, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 28 untouched benchmarks


Comparing claude/bundle-tree-shaking-benchmarks-uq8hcl (fcdd0fd) with main (a1b033c)

Open in CodSpeed

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (2)
.github/workflows/ci.yml (1)

35-53: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win

Bundle-size regression skips coverage/codecov/artifact-upload steps.

The "Bundle size benchmarks" step has no continue-on-error/always() guard, so on a size regression the job stops before pnpm coverage, codecov-action, and upload-artifact run (only the sticky-comment step is guarded to still run). This means any bundle-size drift also hides code-coverage results and blocks artifact publishing for that run, which can make debugging a PR harder than necessary.

♻️ Let downstream steps run regardless of the bundle-size gate
       - run: pnpm coverage
+        if: ${{ !cancelled() }}
         working-directory: packages/sury

       - uses: codecov/codecov-action@v3
+        if: ${{ !cancelled() }}
         with:
           files: packages/sury/coverage/lcov.info
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/ci.yml around lines 35 - 53, The bundle-size guard in the
CI workflow is stopping the rest of the job when `pnpm bench:bundle` fails,
which prevents later coverage and artifact steps from running. Update the
`Bundle size benchmarks` job step so it is non-blocking for the workflow control
flow (for example by allowing the step to fail without cancelling the job),
while keeping the bundle-size regression visible in the PR. Make sure the
downstream `pnpm coverage`, `codecov-action`, and `upload-artifact` steps still
execute after a bundle-size mismatch, and keep the existing sticky PR comment
step behavior intact.
packages/sury/package.json (1)

71-71: 🔒 Security & Privacy | 🔵 Trivial | ⚡ Quick win

Bump esbuild to 0.28.1. It includes fixes for GHSA-g7r4-m6w7-qqqr and GHSA-gv7w-rqvm-qjhr; since this is only a devDependency, the risk is limited to the build environment, but the upgrade is low-cost hardening.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/sury/package.json` at line 71, Update the esbuild devDependency in
the package manifest from 0.28.0 to 0.28.1. Keep the change limited to the
package.json entry for esbuild so the build tooling picks up the patched version
without altering unrelated dependencies.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@packages/sury/tests/bundle.bench.mjs`:
- Around line 122-126: Remove the obsolete esbuild-warning comment in the bundle
benchmark setup and keep the `logLevel` configuration minimal in that
`bundle.bench.mjs` test. The stale note about `src/S.js` intentionally
re-exporting missing `unnest`/`compile` symbols no longer applies, so delete it
from the `build` options block and verify `logLevel: "silent"` is still only
being used to suppress the remaining intentional warnings.

---

Nitpick comments:
In @.github/workflows/ci.yml:
- Around line 35-53: The bundle-size guard in the CI workflow is stopping the
rest of the job when `pnpm bench:bundle` fails, which prevents later coverage
and artifact steps from running. Update the `Bundle size benchmarks` job step so
it is non-blocking for the workflow control flow (for example by allowing the
step to fail without cancelling the job), while keeping the bundle-size
regression visible in the PR. Make sure the downstream `pnpm coverage`,
`codecov-action`, and `upload-artifact` steps still execute after a bundle-size
mismatch, and keep the existing sticky PR comment step behavior intact.

In `@packages/sury/package.json`:
- Line 71: Update the esbuild devDependency in the package manifest from 0.28.0
to 0.28.1. Keep the change limited to the package.json entry for esbuild so the
build tooling picks up the patched version without altering unrelated
dependencies.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 62f332ac-8868-46f0-baa0-19fa1501dc60

📥 Commits

Reviewing files that changed from the base of the PR and between 92e15e3 and a09e1de.

⛔ Files ignored due to path filters (1)
  • pnpm-lock.yaml is excluded by !**/pnpm-lock.yaml
📒 Files selected for processing (10)
  • .github/workflows/ci.yml
  • .gitignore
  • CONTRIBUTING.md
  • package.json
  • packages/sury/package.json
  • packages/sury/scripts/pack/Pack.res
  • packages/sury/scripts/pack/Pack.res.mjs
  • packages/sury/src/S.js
  • packages/sury/tests/bundle-size.snapshot.json
  • packages/sury/tests/bundle.bench.mjs
💤 Files with no reviewable changes (3)
  • packages/sury/src/S.js
  • packages/sury/scripts/pack/Pack.res
  • packages/sury/scripts/pack/Pack.res.mjs

Comment thread packages/sury/tests/bundle.bench.mjs Outdated
claude added 2 commits July 1, 2026 19:23
- bundle.bench.mjs: refresh the stale esbuild `logLevel: "silent"` comment.
  It referenced the now-removed unnest/compile re-exports; the setting now
  only silences the pre-existing package.json "types" export-condition warning.
- ci.yml: guard `pnpm coverage`, codecov, and upload-artifact with
  `if: ${{ !cancelled() }}` so a bundle-size regression still fails the job
  without hiding coverage results or blocking the artifact upload.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
Addresses the review's security note (GHSA-gv7w-rqvm-qjhr, GHSA-g7r4-m6w7-qqqr,
patched in 0.28.1). esbuild is a build-only devDependency here — no dev-server
use and pnpm installs from a lockfile with integrity hashes — so real exposure
is negligible, but 0.28.1 is the standard patched release and keeps the range
clean. Output is byte-identical to 0.28.0, so the size snapshot only needs its
recorded esbuildVersion bumped; the whole tree (incl. vite/vitest) dedupes to
0.28.1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/ci.yml:
- Around line 60-63: The Codecov upload step in the CI workflow is using the
obsolete codecov/codecov-action@v3, so update that step to the current v7 major
and keep the existing coverage file path. If you want to remove the static
token, configure the job with id-token: write and set use_oidc: true on the
Codecov action; otherwise, ensure the existing upload auth still works after the
version bump.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: c5cd9087-a6de-464f-bbd3-4d0331301431

📥 Commits

Reviewing files that changed from the base of the PR and between a09e1de and 50f6ce7.

📒 Files selected for processing (2)
  • .github/workflows/ci.yml
  • packages/sury/tests/bundle.bench.mjs
🚧 Files skipped from review as they are similar to previous changes (1)
  • packages/sury/tests/bundle.bench.mjs

Comment thread .github/workflows/ci.yml Outdated
claude and others added 6 commits July 1, 2026 19:35
actionlint flags codecov/codecov-action@v3 as too old to run on GitHub
Actions (deprecated runner). v7.0.0 is the current major. The step stays
tokenless — v7 supports tokenless uploads for public repos, which also keeps
fork PRs working (OIDC would require an id-token forks can't be granted).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
Rewrites the bundle-size benchmark as bundle.bench.ts with full type
annotations (Sizes/Scenario/Snapshot/Row), matching the .ts convention used
by types.bench.ts and sury.bench.ts. Runs via tsx (bench:bundle now
`tsx tests/bundle.bench.ts`, like bench:types).

Since the file now matches Vitest's `tests/**/*.bench.ts` benchmark glob, it's
added to the benchmark `exclude` list alongside types.bench.ts — it's a
standalone tsx script (calls process.exit), not a Vitest benchmark. Verified:
tsc --strict passes, output is byte-identical, and `pnpm bench` no longer
collects it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
CI's `pnpm coverage` failed (Sury job) because vitest's typecheck pass
compiles every .ts file in the package against tsconfig.json (no `include`
restricts it) — including bundle.bench.ts — and that config sets
noUncheckedIndexedAccess + module: ES2020. Neither was covered by the ad-hoc
`tsc --strict` sanity check used when the file was first converted.

- Record/array index reads (`measured[s.name]`, `split("=")[1]`,
  `outputFiles[0]`) return `T | undefined` under noUncheckedIndexedAccess;
  asserted the two `measured` reads (guaranteed populated by the preceding
  loop) and gave the markdown-path split a fallback.
- Top-level `await measure(s)` isn't allowed under module: ES2020 (needs
  es2022+); wrapped the executable body in an async `main()` instead of
  changing the shared tsconfig.

Verified against the real config this time: `tsc --noEmit -p tsconfig.json`
and `pnpm coverage` both pass clean, and bench output/--update/--markdown are
unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
Replaces the single "🌳 Tree-shaking: ..." sub-line (which only compared the
smallest scenario against the full surface) with a per-row "Tree-shaking"
column: each scenario's % reduction vs the `total (export *)` ceiling. Now
every scenario's tree-shaking effectiveness is visible at a glance, not just
the smallest one.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MTDvBdMTEzVVm4pg4Rc31c
Picks up main's sync merge (S.pattern Input-type fix, #283) alongside the
tree-shaking column change.
@DZakh
DZakh merged commit 92b2c97 into main Jul 4, 2026
10 checks passed
@DZakh
DZakh deleted the claude/bundle-tree-shaking-benchmarks-uq8hcl branch July 4, 2026 09:31
DZakh pushed a commit that referenced this pull request Jul 4, 2026
…ark+CI

Merges origin/main (bundle-size benchmark PR #281, S.pattern Input-type fix,
fixed-property-object encoding), then reverts the parts superseded by the
spec harness while keeping the parts that aren't:

Kept:
- Dead S.unnest/S.compile export removal (Pack.res/Pack.res.mjs/S.js) — a
  real, unrelated correctness fix bundled into the same PR.
- codecov-action v3 -> v7 (actionlint-driven, independent of bundle-size).
- The two unrelated library fixes from main.

Reverted (superseded by this commit's harness integration):
- .github/workflows/ci.yml: the bundle-size CI step, sticky PR-comment
  posting, the permissions it needed, and the `if: !cancelled()` guards added
  specifically to accommodate it.
- packages/sury/tests/bundle.bench.ts and bundle-size.snapshot.json (fixed
  hand-picked scenarios against a committed baseline, its own CI gate).
- The bench:bundle/benchmark:bundle scripts, esbuild devDependency on
  packages/sury, vitest.config.mjs's exclusion, CONTRIBUTING's new section.

Integrated instead:
- packages/spec/bundleSize.ts derives `ts.bundleBytes` per spec: bundles a
  tiny `S.parser(schema)` entry with esbuild (aliased to the dev source),
  minifies, gzips — same technique, now scoped to the schema under test
  rather than a handful of fixed global scenarios.
- harness.ts's recomputeGoldens now derives ts.bundleBytes the same way it
  already derives jsonSchema/operations/ts.input/ts.output/ts.instantiations
  — one unified `spec update`/`spec new`, no separate command.

Async + parallel, per follow-up request:
- deriveBundleBytes uses esbuild's async `build()` (not `buildSync`) so
  concurrent specs' bundle measurements run as genuinely parallel child-process
  builds via Promise.all — unlike the TS-introspection environment (introspect.ts),
  which is inherently synchronous/single-instance and stays that way, just wrapped
  in an async signature for uniform composition.
- recomputeGoldens is now async; it kicks off the bundle-size build *before*
  the synchronous TS-introspection/operation work, so a single spec's async
  and sync derivations genuinely overlap rather than running back-to-back.
- cli.ts's cmdUpdate/cmdCheck parallelize their per-spec work via Promise.all
  (collecting results and printing in original order afterward, avoiding
  interleaved output); cmdNew parallelizes its two derivations the same way.
  Top-level dispatch uses an async main() (matching the existing
  bundle.bench.ts convention) rather than top-level await, since the shared
  tsconfig targets module: ES2020.

Verified: full pnpm test (99 files, 1515 tests, no type errors); spec check/
update/new all green with real derived ts.bundleBytes (3765 for S.string,
matching the standalone measurement); drift detection confirmed on the new
dimension; a deliberately malformed spec still fails gracefully without
aborting the batch. Also confirmed @ark/attest's peer-resolved TypeScript
version in this repo is 5.8.3 (matching Sury's own pin), separate from its
internal devDependency (5.9.3) which is irrelevant here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f
DZakh added a commit that referenced this pull request Jul 11, 2026
* Add AI-first declarative spec test suite (string spec)

Introduce a declarative, machine-authored test-spec harness under
packages/sury/specs. One YAML file captures a schema's full contract —
type, input/output JSON Schema, and per-operation (parse/decode/encode)
generated code plus named input→output|error examples — and is expanded
into committed Vitest files that run with plain `pnpm test`.

Design:
- Format defined *as a Sury schema* (scripts/spec/format.mjs), which emits
  specs/spec.schema.json and validates specs via Sury's own parser.
- Golden-master: authors write inputs; `spec update` runs the real schema
  and fills every derivable golden (expression, jsonSchema, example results).
- Exhaustive by construction: every dimension/operation is required; gaps
  must be an explicit `_skip: <reason>` (enum or todo(#…)), keeping coverage
  greppable and preventing silent omissions by machine authors.
- Closed world + byte-deterministic canonical form (`spec fmt`); `spec:check`
  gates format validity, skip lint, canonical form, live-golden match, and
  generated-file freshness without needing a ReScript build.

CLI: `spec new|update|gen|fmt|check|schema`. Documented in CONTRIBUTING.
Property-based testing is scoped out and tracked as a planned dimension
(properties: _skip: todo(#pbt-dimension)).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec harness: TS CLI, publish specs, snapshot generated tests

Address three refinements to the spec suite:

- Publish specs with the package: `specs/*.yaml` added to `files` (yaml only;
  the emitted spec.schema.json and scripts stay out of the tarball).
- Stop committing generated tests. `tests/generated/` is now gitignored and
  regenerated before `pnpm test` via a `pretest` hook, so behavior and types
  are still asserted in CI without generated code landing in git. The committed,
  reviewable record of generation is a file snapshot in tests/spec_test.ts
  (tests/__snapshots__/string.gen_test.ts.snap).
- Rewrite the CLI in TypeScript (scripts/spec/*.ts, run via tsx), typed spec
  format and harness. `check` drops generated-file freshness (nothing committed
  to compare) and instead gates spec.schema.json freshness plus, per spec,
  format validity, skip-lint, canonical form, and goldens matching live behavior.

tests/spec_test.ts is the harness's own suite: validates every spec, asserts
goldens are fresh, checks the closed-world guarantee, and snapshots the
generated output. CONTRIBUTING updated to match.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: add Claude skill, trim package scripts, pin yaml

- Replace the CONTRIBUTING "Specs" section with a concise, AI-first Claude
  skill at .claude/skills/spec/SKILL.md (golden-master workflow, rules, format,
  dimensions, layout).
- Drop the redundant `spec:check` script — `pnpm spec check` covers it.
- Remove the `res` watch script from the sury package; point the root `res`
  script directly at `rescript watch` via pnpm exec.
- Pin `yaml` to the exact version 2.9.0 (matches the repo's exact-version
  devDependency convention).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Move spec CLI to packages/spec: published infra + dev goldens

Split the spec harness's two uses of sury so the CLI stays stable while
Sury's internals are iterated on:

- New `packages/spec` workspace package holds the CLI (format.ts, harness.ts,
  cli.ts, run via tsx).
- `format.ts` (spec format as a Sury schema, spec validation, spec.schema.json)
  imports a *published* sury via the `sury-published` alias
  (npm:sury@11.0.0-alpha.9). This infra no longer breaks when the working
  tree's core is mid-refactor.
- `harness.ts` (golden recomputation, generation) imports the *dev* source
  (../sury/src/S.js), so expression/jsonSchema/example goldens still track the
  code under test — the whole point of the suite.

sury delegates `pnpm spec` / `pretest` to `tsx ../spec/cli.ts`; specs and
generated tests still live in packages/sury (specs ship with the package).
tests/spec_test.ts imports the harness across the package boundary and validates
via published sury. `yaml` moved from sury to the spec package.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: root command, drop properties dimension, enforce identity invariant

- Add a root `spec` script so `pnpm spec …` runs from the repo root (paths are
  cwd-independent). `[id]` is optional for update/check/fmt/gen (all specs when
  omitted); only `new` requires one.
- Remove the `properties` (PBT) dimension from the format entirely for now —
  from the Sury format schema, Spec type, scaffold, spec.schema.json, and the
  string spec. It'll be re-added when PBT is built.
- Enforce the identity invariant both ways: an operation compiles to Sury's
  pass-through (`noopOperation`) iff it is declared `_skip: identity`. `update`
  refuses and `check` fails if a full op block compiles to identity (use
  `_skip: identity`) or if `_skip: identity` is claimed for a non-identity op.
  string.yaml's decode/encode are now `_skip: identity`.
- Skill: root-run commands, id optionality, identity rule, drop properties, and
  stop documenting harness internals (just "don't touch packages/spec").

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Skill: record missing harness features under Spec Harness Suggestions

Add an "Improving the harness" note to the spec skill: when a strictness or
author-guidance gap surfaces while writing specs, add it under a new
"Spec Harness Suggestions" section in CONTRIBUTING.md instead of working around
it. Seed that section with the identity-detection hardening idea.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: infer TS types from schema, drop schema.res, auto-derive `spec new`

- format.ts: every exported type (Skip, Example, Operation, Spec, OpName) is
  now INFERRED from its Sury schema via `S.Output<typeof x>` instead of
  hand-duplicated — the schema is the single source of truth for both runtime
  validation and the TS shape. `orSkip` is generic so the wrapped schema's
  Output type isn't widened away.
- Remove `schema.res` (the ReScript surface) from the format for now — only
  `schema.ts` is authored/executed. Tracked as a Spec Harness Suggestion to
  re-add once the harness can compile+run `.res` source.
- `spec new` is now `spec new --id <id> --ts <schema>` (both required) and
  immediately derives `jsonSchema` and `operations` by executing the given
  schema (new `scaffoldJsonSchema`/`scaffoldOperations` in harness.ts) —
  identity ops collapse to `_skip: identity` automatically. Only `types` and
  example inputs remain manual; that gap is tracked as a Suggestion too
  (auto-deriving the type string would need TS Compiler API integration).

Fallout from properly inferring types (each a real gap, not a workaround):
- `isNoop`'s parameter type didn't structurally overlap with a function type
  (TS's weak-type check) — typed as `Function` instead of `{name?: string}`.
- Per-example `_skip` was never actually schema-backed (`examples:
  S.record(example)`, no `orSkip`) — removed the dead Skip-handling code that
  the old hand-written types incorrectly implied was live.
- `S.toJSONSchema`'s `JSONSchema7` return type doesn't structurally satisfy
  Sury's own `JSON` type (no index signature) — bridged with one explicit,
  documented cast (`asJson`) at that boundary.
- Hardened `spec check`'s per-file loop so a malformed spec (shape assumed by
  identityViolations/recomputeGoldens but not matching the format schema)
  reports its own failure and lets the batch continue, instead of throwing
  and aborting the whole run.

Verified: standalone tsc on cli.ts clean, full vitest typecheck graph clean,
`spec new --id --ts` end-to-end (auto-derives correctly incl. identity
collapse), and a deliberately malformed spec fails gracefully without
aborting the batch. Skill and CONTRIBUTING updated to match.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: restructure per-surface (ts/*), bare `identity` literal, named noop const

- Regroup everything surface-specific under `ts`: `schema`, `input`/`output`
  (replacing the combined `types.ts` generic-string assertion with two
  independently-skippable type strings, generated as separate
  `expectTypeOf<S.Output<schema>>`/`S.Input<schema>` assertions), plus
  `instantiations`/`bundleBytes` (moved from top-level). `jsonSchema` and
  `operations` stay top-level (language-independent). Paves the way for a
  future `res` (ReScript) surface with its own, smaller shape — no `input`
  (S.t<'value> has no separate Input type param) and no `instantiations`
  (TS/attest-only). string.yaml's existing `string`/`string` type coverage is
  preserved as real `ts.input`/`ts.output` values, not regressed to a skip.
- An operation that compiles to Sury's pass-through is now the bare literal
  `identity` (e.g. `decode: identity`), not `_skip: identity` — skip stays
  reserved for dimensions that are genuinely optional; identity is a real,
  verified fact about the operation, not an absence.
- Extract the identity-detection magic string into
  `NOOP_OPERATION_WHICH_WILL_NEVER_CHANGE`, replacing the earlier suggestion to
  additionally compare the compiled `.toString()` body — the loud, named
  constant is considered sufficient (any rename breaks every identity-marked
  op across every spec, immediately and visibly).
- CONTRIBUTING Spec Harness Suggestions: drop the now-addressed identity and
  schema.res/cross-surface-conformance bullets, reword the type-string one for
  `ts.input`/`ts.output`, add a blank placeholder bullet.

Verified: pnpm spec check green, full vitest typecheck graph (28 tests, no
type errors) with the new `S.Output`/`S.Input` per-field assertions, standalone
tsc on cli.ts clean, spec new produces the correct nested scaffold with
identity auto-collapse, and both identity-invariant violation directions still
fire correctly with the bare-literal spelling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: derive ts.input/output/instantiations via vendored TS introspection

Unify `spec update`/`spec new` into one step that derives everything the
harness knows how to derive, instead of leaving ts.input/ts.output/
ts.instantiations for manual fill.

- New packages/spec/introspect.ts vendors just the mechanism @ark/attest
  itself uses internally for instantiation counting (bench/type.js +
  cache/utils.js): an isolated @typescript/vfs virtual environment (the tech
  behind the TS Playground), diffed against a bare-import baseline via the
  real `program.getInstantiationCount()`, plus `checker.typeToString()` for
  the Output/Input type strings. Deliberately does NOT depend on @ark/attest —
  what makes attest's own `setup()` slow (~15s, measured) is a separate,
  unrelated whole-project scan for pre-written hardcoded-expected-value
  assertions, which this harness has no use for. The vendored path measures
  ~1s cold, ~50-200ms warm per additional schema in the same process (also
  measured) — fast enough to run on every update, no separate slow command.
- harness.ts's recomputeGoldens now derives ts.input/ts.output/
  ts.instantiations the same way it already derives jsonSchema/operations/
  expression — respecting an existing `_skip` (an author's deliberate skip,
  e.g. `not-applicable`, is never overwritten). This means `spec check`'s
  existing goldens-match-live diff automatically covers the new dimension too,
  with no new checking logic needed.
- cli.ts's `spec new` now populates ts.input/output/instantiations for real
  instead of scaffolding them as `_skip: todo(#fill)` — only `ts.bundleBytes`
  (a genuinely separate, unbuilt dimension) and example inputs remain manual.
- string.yaml: instantiations filled in (226, replacing the stale
  `todo(#instantiations-dimension)` skip from before the feature existed).

Verified: full `pnpm test` (99 files, 1509 tests) passes with no meaningful
duration regression (~15s, same as baseline) since the environment is
memoized once per process; drift detection confirmed on the new
ts.instantiations field; spec new end-to-end with a fresh schema (transform
and refinement cases) produces fully-derived, correct input/output/
instantiations. Skill and CONTRIBUTING updated; the now-resolved manual-fill
suggestion removed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Integrate bundle-size into the spec harness, revert standalone benchmark+CI

Merges origin/main (bundle-size benchmark PR #281, S.pattern Input-type fix,
fixed-property-object encoding), then reverts the parts superseded by the
spec harness while keeping the parts that aren't:

Kept:
- Dead S.unnest/S.compile export removal (Pack.res/Pack.res.mjs/S.js) — a
  real, unrelated correctness fix bundled into the same PR.
- codecov-action v3 -> v7 (actionlint-driven, independent of bundle-size).
- The two unrelated library fixes from main.

Reverted (superseded by this commit's harness integration):
- .github/workflows/ci.yml: the bundle-size CI step, sticky PR-comment
  posting, the permissions it needed, and the `if: !cancelled()` guards added
  specifically to accommodate it.
- packages/sury/tests/bundle.bench.ts and bundle-size.snapshot.json (fixed
  hand-picked scenarios against a committed baseline, its own CI gate).
- The bench:bundle/benchmark:bundle scripts, esbuild devDependency on
  packages/sury, vitest.config.mjs's exclusion, CONTRIBUTING's new section.

Integrated instead:
- packages/spec/bundleSize.ts derives `ts.bundleBytes` per spec: bundles a
  tiny `S.parser(schema)` entry with esbuild (aliased to the dev source),
  minifies, gzips — same technique, now scoped to the schema under test
  rather than a handful of fixed global scenarios.
- harness.ts's recomputeGoldens now derives ts.bundleBytes the same way it
  already derives jsonSchema/operations/ts.input/ts.output/ts.instantiations
  — one unified `spec update`/`spec new`, no separate command.

Async + parallel, per follow-up request:
- deriveBundleBytes uses esbuild's async `build()` (not `buildSync`) so
  concurrent specs' bundle measurements run as genuinely parallel child-process
  builds via Promise.all — unlike the TS-introspection environment (introspect.ts),
  which is inherently synchronous/single-instance and stays that way, just wrapped
  in an async signature for uniform composition.
- recomputeGoldens is now async; it kicks off the bundle-size build *before*
  the synchronous TS-introspection/operation work, so a single spec's async
  and sync derivations genuinely overlap rather than running back-to-back.
- cli.ts's cmdUpdate/cmdCheck parallelize their per-spec work via Promise.all
  (collecting results and printing in original order afterward, avoiding
  interleaved output); cmdNew parallelizes its two derivations the same way.
  Top-level dispatch uses an async main() (matching the existing
  bundle.bench.ts convention) rather than top-level await, since the shared
  tsconfig targets module: ES2020.

Verified: full pnpm test (99 files, 1515 tests, no type errors); spec check/
update/new all green with real derived ts.bundleBytes (3765 for S.string,
matching the standalone measurement); drift detection confirmed on the new
dimension; a deliberately malformed spec still fails gracefully without
aborting the batch. Also confirmed @ark/attest's peer-resolved TypeScript
version in this repo is 5.8.3 (matching Sury's own pin), separate from its
internal devDependency (5.9.3) which is irrelevant here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: drop generated test files, add schema descriptions, help, error snapshots

Four changes, largely enabling each other:

1. Remove code-generation entirely (generateTest/genPath/GEN_DIR, the `gen`
   CLI command, the pretest hook, tests/generated/). packages/sury/tests/
   spec_test.ts is a single, committed, hand-written Vitest file that already
   dynamically loops over every spec at run time and calls straight into
   recomputeGoldens — so example execution and every dimension's freshness
   were already exercised (and coverage-attributed) by a real Vitest run,
   with no per-spec generated file needed. The type-level checks the
   generated files' `expectTypeOf` provided are redundantly covered by
   introspect.ts's own type-string drift detection. Also fixed a real bug
   while touching valueToCode: it used JSON.stringify/String() to serialize
   example outputs, which throws on bigint (a real S.bigint output type) and
   silently mishandles Date/Map/Set/NaN/-0 — special-cased bigint (`${v}n`),
   documented the rest as a Suggestion.

2. format.ts: every schema/field gets a `.with(S.meta, {description})`, so
   spec.schema.json (consumed by yaml-language-server hover/autocomplete and
   by AI authors) is self-documenting. Found and fixed a real type-inference
   regression along the way: a generic `desc<T extends Schema<unknown,
   unknown>>(schema, description)` helper reliably collapsed `Spec` to
   `unknown` at specSchema's nesting depth (reproduced in isolation; direct
   `.with(S.meta,...)` chaining at the identical depth was unaffected) —
   dropped the helper for direct chains everywhere, per the CLAUDE.md note
   about this schema shape being costly to instantiate.

3. cli.ts: a real `help`/`--help`/`-h` command with full per-command
   descriptions, shown (with an "Unknown command" header when applicable)
   as the default output on any invalid/missing command instead of a
   one-line usage string.

4. New packages/sury/tests/spec_errors_test.ts: snapshots the exact guiding
   messages checkSpec produces for ten deliberately-broken mutations of a
   real spec (stale expression, stale example output, invalid _skip reason,
   non-canonical form, both identity-mismatch directions, two format-
   validation failures, a schema syntax error, and multiple simultaneous
   problems) — proving they name the problem and say what to run, not just
   pass/fail. Required extracting cmdCheck's per-spec logic into a new
   harness.checkSpec (also moving lintSkips there), since cli.ts's
   side-effecting main() call makes it unsafe to import directly into a
   test; cli.ts's cmdCheck is now a thin wrapper so the CLI and the tests
   exercise the exact same code. Found (and documented as a Suggestion,
   not fixed) a minor cascading double-error when a format-validation
   failure leaves ts.schema evaluating to a non-schema value.

Verified: full pnpm test (99 files, 1517 tests, no type errors), standalone
tsc on cli.ts clean, spec check/update/new all functionally verified
end-to-end, descriptions confirmed present in the regenerated
spec.schema.json, help/--help/-h/unknown-command/no-command all produce
correct output and exit codes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec: trim verbose descriptions and help text

Shorten the new-command help block, drop the now-superfluous
.type-probe.ts gitignore entry, and tighten every S.meta description in
format.ts down to the guiding essentials so spec.schema.json stays
scannable for AI authors and editor hover alike.

* Spec: value-based golden comparison via @vitest/expect, drop update command

Generalize valueToCode to any JS value (bigint/Date/Map/Set/RegExp/NaN/-0 at
any nesting depth, not just top-level bigint), and use @vitest/expect's
equals() — the exact function toStrictEqual calls internally — to decide
whether a recorded example output still matches live behavior. An
equivalent-but-differently-formatted recorded value now survives a recompute
instead of being churned into a new (also correct) rendering.

Remove the standalone `update` command; `check` gains a --write flag that
persists whatever's safely fixable (canonical form, stale goldens), mirroring
prettier --write / eslint --fix. --write is a no-op for a format-invalid spec
or a live identity mismatch, both of which need a human decision rather than
a rewrite. Updates cli.ts, the SKILL.md workflow, spec.schema.json's
descriptions, and the CLI's own guiding messages accordingly.

* Spec: non-circular --write message, Symbol support in valueToCode

checkSpec now says "resolve the identity mismatch above first, then
--write can fix it" instead of just pointing back at --write when an
identity violation is what's actually blocking the write — the old
message was circular for someone who'd just run --write and had it
no-op. Resolves the CONTRIBUTING.md suggestion.

Also extend valueToCode to registry symbols (Symbol.for(key), via
Symbol.keyFor to round-trip) — a bare Symbol() can't be represented in
source since every call produces a unique value, so that case throws a
guiding error instead of silently producing wrong output.

* Spec: add object/tuple/union/merge specs, fix 3 harness gaps found writing them

13 new specs covering the schema shapes types.bench.ts used to benchmark:
object (1/5/10 fields, required and mixed-optional, 3-level nesting), tuple
(2/5/10 items), union (2/5 primitive members, 5 discriminated-object
members), and merge (all-required, and optional-preserving per issue #157).

Writing them surfaced three real harness gaps, all fixed here:

- introspect.ts's typeToString() printed a union (or any type structurally
  matching a builtin alias like Record<K,V>) as the useless literal string
  "__Output"/"__Input" instead of expanding it, because the resolved type
  keeps an alias symbol back to the type alias being inspected. Fixed by
  passing TypeFormatFlags.InTypeAlias, which tells the printer this call IS
  the alias's own definition.
- scaffoldJsonSchema crashed spec new outright for any schema containing a
  bigint or symbol field, since S.toJSONSchema has no JSON representation
  for either — a real "not applicable" case, not a bug, so it now catches
  and scaffolds `_skip: not-applicable` instead of throwing.
- cmdCheck's --write gate required the spec to already be format-valid,
  which made it unable to do the one thing it exists for: a freshly-added
  example (just `input`, per the documented workflow) is format-invalid
  until --write fills in output/error. Removed that precondition; recompute
  failures are now caught locally and fall through to checkSpec's own
  reporting instead of writing anything.

Also extends valueToCode to registry symbols was already done; this pass
additionally exercises it against real schemas with symbol/bigint fields at
every nesting depth (object, tuple, and discriminated-union members).

* Remove @ark/attest type-instantiation benchmark suite

tests/types.bench.ts's 18 hardcoded-expectation scenarios (object/tuple/
union/merge shapes) are superseded by the 13 specs added in the previous
commit — same coverage, but each schema's instantiation count now lives
as ts.instantiations on its own spec instead of a standalone benchmark
file, and drifts get caught the same way any other spec regression does
(pnpm test / spec check) rather than a separate pnpm bench:types run.

Drops the CI "Type performance benchmarks" step, the bench:types/
benchmark:types scripts, and the @ark/attest devDependency. introspect.ts
still explains (in comments) why the spec harness vendors its own
@typescript/vfs-based introspection instead of depending on attest
directly — that rationale doesn't change, only the now-redundant
standalone benchmark it's being compared against goes away.

* Fix CI timeout: raise Vitest testTimeout for golden recomputation

recomputeGoldens (spec_test.ts, spec_errors_test.ts) runs a TS-program
introspection pass plus an esbuild child-process build per spec. The
first spec processed pays the ~1s cold-start cost the spec skill
documents, which pushed CI's slower/more contended runner past Vitest's
5000ms default: 'merge.optional > goldens match live behavior' timed out
at 5093ms, and spec_errors_test.ts's 'stale golden' case already passed
at 4890ms — 110ms of margin. Raises testTimeout to 20s globally rather
than a single per-test override, since both test files exercise the same
slow path.

Also drops vitest.config.mjs's benchmark.exclude for tests/types.bench.ts,
a dangling reference to the file removed in the previous commit.

* Spec: fix review-flagged correctness bugs and a real efficiency gap

Correctness (each verified with a live repro before and after):

- CI never actually ran identityViolations/lintSkips against the real
  specs — pnpm coverage only runs spec_test.ts, which called validate()/
  recomputeGoldens() but never checkSpec(). spec_test.ts now runs the
  same two checks per spec directly, so a drifted identity marker or a
  bad _skip reason fails pnpm test, not just the standalone spec check
  nothing in CI invokes.
- checkSpec (harness.ts) and cmdCheck's --write path (cli.ts) both used
  `if (schema)` to mean "evalSchema didn't throw" — but a ts.schema that
  evaluates without throwing to a falsy, non-schema value ("0", "", NaN)
  is falsy too, so it silently skipped the entire identity/goldens check
  and reported a clean pass. Both now track whether evaluation actually
  threw, separately from the value it produced.
- `spec check` never reported spec.schema.json as missing, only as
  stale, because `existsSync(...) && ...` short-circuits to no-failure
  when the file's gone.
- recomputeGoldens's jsonSchema derivation wasn't guarded like its
  sibling scaffoldJsonSchema — editing an existing schema to add a
  bigint/symbol field without flipping jsonSchema to _skip crashed
  spec_test.ts with a raw internal Sury exception. Extracted the shared,
  try/caught deriveJsonSchema used by both now.
- scaffoldOperations had no such guard either and crashed `spec new`
  with a raw stack trace whenever --ts evaluated successfully to
  something that isn't a usable schema (a plausible typo like "S.strng").
- A mistyped spec id crashed `spec check`/`spec fmt` with a raw ENOENT;
  targets() now fails with a clear "no such spec" message instead.

Efficiency: cmdCheck's --write path recomputed every spec's goldens
twice — once to decide what to write, again inside the checkSpec call
right after — doubling the esbuild+TS-introspection cost per spec.
checkSpec now takes an optional already-computed result so --write's
own recompute is reused instead of redone; verified this still reports
an unrelated problem (e.g. a bad _skip reason) that --write can't fix,
alongside the goldens it did fix.

Also: dedup evalSchema/inlineToValue (byte-identical), make
KEY_ORDER/TS_KEY_ORDER/OP_ORDER exhaustiveness-checked against the
format schema (a missing field is now a compile error, not a silently
misordered key), and fix format.ts's _skip-reasons comment, which had
already drifted out of sync with the description two lines below it.

* Spec: guard introspect.ts against a silently-empty type golden, fix lint

deriveTypeInfo now throws (with the actual compiler diagnostics, when
any) if the __Output/__Input type-alias walk comes up empty, instead of
returning "" — an empty golden would otherwise byte-match itself on every
future recompute and never get flagged as wrong. Also documents why the
shared env/PROBE_FILE state is safe under concurrent calls only because
check() has no internal await.

Adds the missing language identifier on SKILL.md's workflow code fence
(markdown lint).

* Spec harness: parse-don't-validate, clean schema-usability check, canonical example formatting, hyphenated spec ids

- format.ts's validate() now returns the parsed Spec on success instead of
  discarding it; checkSpec and its caller work from that parsed value.
- checkSpec detects a ts.schema that evaluates but isn't a real Sury schema
  via the Standard Schema `~standard.vendor` marker, reporting one clear
  message instead of a cascading second error from identityViolations/
  recomputeGoldens. Removes the now-resolved CONTRIBUTING.md bullet.
- jsonSchema no longer whole-dimension `_skip`s when S.toJSONSchema can't
  represent a direction (e.g. bigint/symbol fields) — it records the thrown
  message per direction instead, since the two directions can differ.
- valueToCode emits canonical formatting (spaces, unquoted object keys where
  valid identifiers); recomputeGoldens always emits it and canonicalize
  reformats recorded example output too, so `not canonical`/`goldens stale`
  cleanly separate formatting-only drift from real value drift.
- Renamed spec ids to a consistent hyphen/no-separator scheme (e.g.
  object.required5 -> object5, union.objects5 -> union5-discriminated) and
  added lintSpecsDir, gating `spec check` on the specs dir containing only
  valid *.yaml ids (letters, digits, -).
- Converted spec_errors_test.ts's snapshots to inline for readability.
- Stripped comments that only restated code across packages/spec/*.ts,
  keeping non-obvious "why" ones; added a repo-wide comment-discipline rule
  to CLAUDE.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec harness: fmt->format rename, diff-enriched staleness messages, input formatting, tighter docs

- Renamed the `fmt` CLI command to `format` throughout (cli.ts, message
  strings, docs) for a clearer, unabbreviated name.
- "not canonical"/"goldens stale" now include a plain unified diff (via
  @vitest/utils/diff, ANSI-free) of what actually differs, instead of just
  asserting that something does.
- canonExample now reformats example `input` the same way it already did
  `output`, so both sides of an example get canonical spacing/unquoted keys.
- Tightened CLAUDE.md's Comments section into a short, checkable rule list,
  and condensed the spec skill's more narrative sections (workflow prose,
  the types/instantiations/bundle-size derivation writeup) into scannable
  bullets — dropping historical benchmark-comparison detail that doesn't
  help someone authoring a spec today.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Spec CLI: route check failures to stderr, test against real CLI output

- red()/green() are now TTY-aware (no raw ANSI when piped/captured), and
  every check-failure path (spec.schema.json, specs dir, per-spec) now
  prints via a shared formatFailure() to stderr instead of stdout; success
  stays on stdout. Verified against the real CLI with stream redirection.
- cli.ts is now safely importable: main() only runs when executed directly
  (entry-point guard via import.meta.url), and a new runCheck(id, raw)
  exercises the exact same read-only check flow, formatting, and stream
  routing a real `spec check <id>` run produces, given spec source text.
- spec_errors_test.ts now asserts on runCheck's actual {stdout, stderr}
  instead of calling checkSpec() directly and snapshotting its bare return
  array — these snapshots are now the literal bytes an author or CI sees.
- harness.ts: extracted parseSpec(raw) from readSpec so both cli.ts and the
  tests can parse spec text without going through a file.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Regenerate specs against merged main (Standard JSON Schema support)

main added Standard JSON Schema support and per-target toJSONSchema output
(#267), which shifted every spec's ts.instantiations/bundleBytes (new
StandardSchema module in the bundle/type surface) and reordered some
jsonSchema key ordering (same content, different property order from
Sury's own serialization) — regenerated via `spec check --write` and
refreshed the one inline snapshot referencing string.yaml's old numbers.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Move spec harness implementation details out of the skill into CONTRIBUTING.md

SKILL.md should read as day-to-day authoring guidance, not an implementation
writeup. Moved the introspect.ts/bundleSize.ts mechanism details and the
ts.instantiations methodology discussion (bare-import baseline vs.
alternatives, and why) into a new "Spec Harness" section in CONTRIBUTING.md,
leaving SKILL.md with just the Dimensions table and a pointer to it.

Also fixed a misplaced comment in harness.ts (reformatIfEvaluable's doc had
drifted onto canonExample's own explanation after an earlier edit) while
re-auditing every packages/spec/*.ts and spec test file comment against
CLAUDE.md's tightened Comments rule.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* Address spec harness review findings (#287)

- Gate spec.schema.json freshness in spec_test.ts (CI previously never checked it)
- Make `spec new` refuse to overwrite an existing spec
- Harden valueToCode: throw on cyclic values, class instances, and symbol keys
  instead of emitting a golden that doesn't evaluate back to the real output
- Single-source SKIP_REASONS in format.ts; derive the schema description from it
- Drop the duplicated spec id prefix from lintSkips error paths
- Fail `spec new` on unknown arguments
- Add ±1% bundleBytes tolerance so toolchain bumps don't go stale across all specs
- Document specs-as-docs rationale for publishing specs/*.yaml


Claude-Session: https://claude.ai/code/session_0184gidPSxhfQftHBhQPganb

Co-authored-by: Claude <noreply@anthropic.com>

* Refactor spec skill: trigger on core changes, drop harness internals (#288)

Claude-Session: https://claude.ai/code/session_01Bdw8tzYEpK3pobVWE69sJG

Co-authored-by: Claude <noreply@anthropic.com>

* Spec harness: bundleBytes measures the schema alone, cli.ts throws if imported

- bundleSize.ts now measures schemaTs itself (not S.parser(schemaTs)) for
  ts.bundleBytes, isolating schema-construction cost from compiled-operation
  cost. Force-recomputed every spec's bundleBytes past the ±1% tolerance
  band, since this is a real methodology change, not toolchain noise.
- Split cli.ts's testable report formatting (red/green, formatFailure,
  runCheck) into a new report.ts. cli.ts is a script again: it throws with a
  clear message if ever imported instead of run directly, rather than
  silently doing nothing; spec_errors_test.ts now imports runCheck from
  report.ts.
- Tightened cli.ts's HELP text and fixed several format.ts field
  descriptions (typeof schema, stale S.parser(schema) wording, a clearer
  ts.schema description).
- Trimmed CONTRIBUTING.md's Spec Harness section to drop PR-changelog-style
  detail (removed-benchmark history, considered-and-rejected alternatives),
  keeping just what's durably useful for future contributors.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

* jsonSchema: one-line string instead of a nested YAML object

jsonSchema.input/.output now hold source text (rendered via the same
valueToCode used for example values — spaces, unquoted keys) instead of a
raw JSON Schema object serialized as nested YAML. Unifies the success case
(the schema itself, as text) and the failure case (S.toJSONSchema's thrown
message, already a string) into one plain string field, and lets a spec's
whole contract fit in far fewer lines for a quicker human overview — a
3-level nested schema went from ~20 lines to 2.

format.ts: jsonSchema.input/.output are now S.string, not S.json. Dropped
the now-unused asJson bridge helper in harness.ts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vtcqfz3GAh6vTpphLiyi1f

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants