Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
73 changes: 41 additions & 32 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -1087,38 +1087,47 @@ way (no live Windows execution available here), that is noted explicitly rather
and `:run_entry_smoke`, none of it needed when there's no interpreter, but not yet verified safe
to skip wholesale).

**Bucket A remaining scope: candidate fix shapes for a general mechanism, not yet chosen
between -- full pros/cons/value-add/risk writeup now in `docs/plan-die-fatal-remediation.md`
(2026-08-18, corrected same day via a direct hand-trace -- see below), grounded against the
current, actual 27-site inventory (not the ~31 estimate this entry used to cite before Items
45/Bucket B/slice 1 landed), with the choice itself registered as an open question for the
maintainer in `docs/open-questions.md`.** Three candidates: (a) targeted `goto`s per site,
continuing the slice-1 pattern; (b) a global `HP_FATAL` flag set by `:die`, checked via `if
defined HP_FATAL goto :fatal_exit` at chosen resumption points; (c) change `:die` itself to halt
the process directly (a real `exit`, not `exit /b`). **Corrected 2026-08-18**: this entry
previously claimed (c)'s main risk was "a bare `exit` closes the console window immediately for
a double-click user with no chance to read the message first" -- traced directly against `:die`'s
actual current body and found this does NOT hold: `:die` already `pause`s BEFORE its existing
`exit /b`, so a real interactive user has already read the message by the time a converted
`exit` would run. The real, verified risk is narrower and different: this repo tracks the OS
process exit code and the self-reported `~bootstrap.status.json` `exitCode` field as two
INDEPENDENT signals (most `:die` sites today fall through to `:success`'s unconditional `exit /b
0`, so the real process exit code is currently always `0` regardless of `state` -- a deliberate
"graceful stop" design), and a hand-sweep of every `tests/*.ps1` file found exactly ONE currently
-gating test (`tests/selfapps_entrysmoke_no_interpreter.ps1:171`) hard-asserts that old behavior
as part of its own pass condition. See `docs/agent-lessons-learned.md`'s `:die` entry and the
plan doc's own corrected candidate-(c) section for the full trace -- (c) is genuinely more
viable than this entry previously suggested, though still the widest-blast-radius option of the
three. The plan doc's own recommendation: continue (a) for the immediate next slice regardless
(the only one of the three that fits this item's own "EXTREME CAUTION, one slice at a time"
constraint without modification); if the durable close-the-whole-class outcome is wanted later,
the choice between (b) and (c) is now a closer call than previously stated and should be scoped
as its own dedicated effort with a small proof-of-concept and CI soak time, not folded into the
ongoing slice work. Given the number of remaining call sites and `:die`'s central, load-bearing
role throughout the file, treat this as EXTREME CAUTION on the same order as the
DLL-bundling/hidden-import repair loops
elsewhere in this backlog -- one incremental slice at a time (slice 1 above is the second proof
this works, after Bucket B), not a single sweeping change across every remaining site.
**Bucket A decision (made 2026-08-21, maintainer call -- see `docs/plan-die-fatal-remediation.md`'s
"Batch Roadmap" section for the full reasoning and `docs/agent-lessons-learned.md`'s `:die` entry
for the choreography this decision preserves): continue (a) as a deliberate STOPGAP, batched by
proven shape rather than one site per PR; (c) (change `:die` itself to halt the process, `exit`
not `exit /b`) is the likely eventual target for a SEPARATE, later, dedicated effort once the
batches below shrink the remaining inventory -- not folded into the ongoing batch work. (b) (a
global `HP_FATAL` flag) remains a candidate for that later effort too but is not preferred over
(c) per the plan doc's own blast-radius-vs-structural-simplicity tradeoff writeup.** All 20
remaining sites (27 total minus the 7 already safe -- 2 with slice 1's own `goto`, 5 inside
`:conda_binary_corrupt`'s self-contained `exit /b` chain) are now individually traced and grouped
into 6 batches by shape and risk (`docs/plan-die-fatal-remediation.md`'s Finding 3), landing in
this order:
- **Batch 1 (next up): the conda-acquisition-probe chain, 5 sites** (`conda.bat not found after
bootstrap`, `'conda'`/`'python'`/`'python -V'` not-found-on-PATH, `Conda not found at:`) --
same proven shape as slice 1, can currently stack 4-5 redundant `[ERROR]`/pause pairs for one
root cause before reaching the real sink; existing test `tests/selfapps_conda_bothfail.ps1`
already reaches this exact chain via its `:try_conda_install`-failure setup, so no new test
hook is needed, only new assertions.
- **Batch 2**: `:conda_create_done`'s "python.exe missing" check (1 site) -- fixes a genuinely
misleading "[BOOT] ... Selected Python provider: Conda (Portable)." success-sounding message
that currently prints right after an `[ERROR]` was already reported.
- **Batch 4/5**: 7 embedded-helper-write-failure sites plus 2 CI-only-exposure sites -- low risk,
low urgency, mechanical once each is individually traced.
- **Batch 3**: the `:determine_entry` double-call (the first `call :die "[ERROR] Could not
determine entry point"` site, inside `:after_env_mode_selection`) -- `:determine_entry` runs
TWICE per normal bootstrap; falling through here wastes the entire
dependency-install/pipreqs/warnfix block (far more intervening work than Batch 1) before
reproducing the same failure at the second call site. MEDIUM risk (a real behavior change, not
just a redundant-pause removal) -- lands on its own, later.
- **Batch 6**: `:tci_both_failed`'s own two failure sites (inside `:try_conda_install`) -- MEDIUM
risk, needs a new caller-side coordination flag (not a drop-in goto, since these sites already
sit inside a `call`ed subroutine with their own `goto :eof`); deliberately deferred until after
Batch 1 lands, since Batch 1's own fix already shrinks this site's fall-through blast radius as
a side effect.
- **The "Active Python interpreter not resolved" sink** (inside `:after_env_mode_selection`,
every other batch's `goto` routes toward it) stays deferred indefinitely -- already
Item-45-backstopped for the one dangerous consequence, a fix here would only trim harmless
wasted work.
Same EXTREME CAUTION discipline as every other high-risk change in this file: one batch lands at
a time, full 8-lane matrix CI proof to completion before the next batch starts, no blanket sweep
across every remaining site in one PR.

## Cold Storage (promising ideas, deliberately shelved -- revisit only if a named trigger fires)

Expand Down
33 changes: 33 additions & 0 deletions docs/agent-interconnect.md
Original file line number Diff line number Diff line change
Expand Up @@ -1063,6 +1063,39 @@ this specific check does not catch. So the fix relocates pause #2 to an already-
cutting pause count. A further Bucket A slice targeting `:after_env_mode_selection`'s own check is
the natural next candidate if reducing to a single pause is ever pursued.

### Batch 1 (CLAUDE.md Item 46 Bucket A): the conda-acquisition-probe chain now also collapses, same limitation as above

**Sibling fix to slice 1 above, applied one step EARLIER in the same overall provider-acquisition
flow -- touch any of these 5 sites, must understand the other 4, since they form one linear
fall-through chain, not 5 independent bugs.** `docs/plan-die-fatal-remediation.md`'s Finding 3
traced this chain in full: `:after_conda_bat_validation`'s own `if not defined CONDA_BAT (...)`
block (`conda.bat not found after bootstrap.`), then (once `CONDA_BAT` update reaches PATH) `where
conda` / `where python` / `python -V`'s own `||` failure blocks, then the REQ-... channel-policy
check (`Conda not found at: %CONDA_BAT%`) -- all 5 share the identical root cause (conda was never
successfully acquired) and, before this fix, could stack up to 4-5 back-to-back `[ERROR]`/pause
pairs before finally reaching `:try_conda_create` with a broken `CONDA_BAT`, dying again there
(already `goto`-fixed by slice 1), and ultimately landing on the same `:after_env_mode_selection`
"Active Python interpreter not resolved" sink slice 1's own fix routes toward. Fixed by adding
`goto :after_env_mode_selection` right after each of the 5 `call :die` lines, same pattern as
slice 1.

**Same "does not reduce to one pause" limitation slice 1 already taught -- do not expect this fix
to eliminate the sink pause.** Whichever of the 5 sites fires first still pauses once (its own
message is the most specific, most actionable one), and execution still reaches the SAME
`:after_env_mode_selection`-guard sink pause described above for slice 1 -- this fix only prevents
the OTHER 4 sites in the chain from ALSO firing for the identical root cause. Worst case goes from
5-6 stacked pauses down to 2, not 1.

**`:try_conda_install`'s own two failure sites (`:tci_both_failed`, CLAUDE.md's Batch 6, not yet
fixed) fall through INTO this exact chain as their caller resumes** -- see that subroutine's own
`call`-frame return via `goto :eof` back to its caller at the Miniconda-install call site, which
then runs `:select_conda_bat` and falls straight into this chain's own top (`conda.bat not found`)
when the install genuinely failed. This means Batch 1's fix already shrinks Batch 6's own
fall-through blast radius as a side effect, even before Batch 6 itself is touched -- confirmed via
`tests/selfapps_conda_bothfail.ps1`, which drives a genuine `:tci_both_failed` failure and (post
this fix) now asserts the resulting cascade collapses to exactly this chain's first site plus the
sink, not the full 5-6-pause worst case.

Regression test: `tests/selfapps_entrysmoke_no_interpreter.ps1` (`self.entrysmoke.no_interpreter_
guard`, `conda-full` lane) -- asserts `[ERROR] python.exe missing from conda environment.` is
ABSENT (proving the second `:handle_conda_failure` call is skipped) and the new
Expand Down
50 changes: 4 additions & 46 deletions docs/open-questions.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,48 +7,7 @@ here and fold the outcome into wherever it actually belongs (CLAUDE.md's Active/
answered questions accumulate here as history; that's what the other docs' own Closed Backlog /
changelog-style sections are for.

## 1. CLAUDE.md Active Backlog Item 46 (Bucket A): which remediation shape for the remaining `:die` non-halting fall-through sites?

Full research, pros/cons/value-add/risk writeup: `docs/plan-die-fatal-remediation.md` (2026-08-18).
Three candidates, none pre-decided:

- **(a) Targeted `goto` per site** -- continue the slice-by-slice pattern already shipped twice
(Bucket B's `:warn_build_incomplete`, Bucket A slice 1's `goto`). Lowest risk, slowest, no
durable protection against a future un-gotoed `call :die` site.
- **(b) A global `HP_FATAL` flag** -- one mechanism, checked at chosen resumption points, protects
current AND future sites. Higher value if done right; real risk the checkpoint placement itself
is incomplete or wrong, invisibly.
- **(c) Make `:die` itself halt the process** -- most complete. **Corrected 2026-08-18** (via a
direct hand-trace, prompted by a maintainer question): the originally-suspected "breaks the
pause-before-exit convention" risk does NOT apply -- `:die` already pauses before its existing
exit line. The real, precisely-bounded risk instead: this repo tracks the OS process exit code
separately from `~bootstrap.status.json`'s own `exitCode` field, and exactly one currently
-gating test (`tests/selfapps_entrysmoke_no_interpreter.ps1`) hard-depends on the OLD
always-exit-0-on-failure behavior. Genuinely more viable than first assessed, though still the
widest-blast-radius option of the three.

The plan doc's own recommendation: continue (a) for the immediate next slice regardless (only
option that fits Item 46's "EXTREME CAUTION, one slice at a time" constraint as-is); if the
durable close-the-whole-class outcome is wanted, the choice between (b) and (c) is now closer than
first assessed (see the plan doc's corrected candidate-(c) section) and should be scoped as its
own dedicated effort with a small proof-of-concept and CI soak time, not folded into the ongoing
slice work.

**Needs the maintainer's call, not something an agent should pick unilaterally** -- this is
exactly the kind of decision Item 46's own process notes flag as requiring EXTREME CAUTION, and
the three options have genuinely different risk/durability tradeoffs rather than one obviously
dominating.

## 2. If (a) is chosen for Item 46: is there a target completion bar for Bucket A?

All ~20 remaining sites (see the plan doc's Finding 1-3 for the current count and severity
breakdown)? Only the ones demonstrated to risk a redundant fallback-chain/consent-prompt replay
(the plan doc's Finding 3 already identifies the `conda.bat not found after bootstrap` site as
structurally identical to the one slice 1 already fixed)? Or opportunistic, one slice per loop
indefinitely, with no fixed end state? Secondary to question 1 -- only needs an answer once (a) is
confirmed as the path forward.

## 3. CLAUDE.md Active Backlog Item 38: which CWD is "correct" for EXE verification -- unify to `dist\`, or leave the two verification points as-is?
## 1. CLAUDE.md Active Backlog Item 38: which CWD is "correct" for EXE verification -- unify to `dist\`, or leave the two verification points as-is?

Two verification points use different working directories today: `:run_exe_smokerun` (run 1,
fresh build) runs from `pushd dist`; `:try_fast_exe`/`:verify_no_exe_interpreter` (every later
Expand Down Expand Up @@ -83,7 +42,7 @@ about the underlying verification-CWD mismatch itself, still open.
CWD-relative-path app between a first and second run, not just wording; an agent choosing wrong
here silently changes what "works" means for a live feature, not just doc content.

## 4. CLAUDE.md Active Backlog Item 35: does the agent have (or can it get) the GitHub repo-admin access needed to actually flip a check to gating?
## 2. CLAUDE.md Active Backlog Item 35: does the agent have (or can it get) the GitHub repo-admin access needed to actually flip a check to gating?

Item 35's real mechanism (an aggregate `selftest-gate` check that could cover every lane with one
required-check entry) is implemented and its precondition has landed, but "gating" in this repo is
Expand All @@ -103,7 +62,7 @@ Without an answer, Item 35's own soak-then-promote slices can get PREPARED (impl
several green runs) but never actually CLOSED -- the last step structurally can't happen without
this.

## 5. CLAUDE.md Active Backlog Item 42, lever 2: is the two-prompt fresh-build flow still worth changing, and if so, reword only or also combine the two prompts?
## 3. CLAUDE.md Active Backlog Item 42, lever 2: is the two-prompt fresh-build flow still worth changing, and if so, reword only or also combine the two prompts?

Lever 1 (console-output tiering) shipped and merged 2026-08-24 (PR #466). Lever 2 -- the two
elective Y/N prompts every successful fresh interactive build shows (`:run_postexec_checkpoint`'s
Expand Down Expand Up @@ -131,7 +90,7 @@ uses an explicit "(optional, safe to skip)"-style cue.
- Should the two prompts stay separate (matching each prompt's own distinct consent-gate mechanism
and existing test coverage -- `self.checkpoint.accept`/`.decline`, `self.optbuild.offer`'s 4
scenarios) with wording-only changes, or should they be combined into a single ask as the item's
own original text floated ("consider combining them into a single, clearly-optional ask")?
own original text floated ("consider combining them into a single, clearly-optional ask")?
Combining is a materially bigger change (restructures which subroutine owns the prompt, how
declining one but not the other is expressed, and every test that currently asserts prompt COUNT
or exact prompt text) than a pure reword, with real risk of a maintainer having a specific
Expand All @@ -141,4 +100,3 @@ uses an explicit "(optional, safe to skip)"-style cue.
mechanism: hide diagnostic-tier lines, keep them in the log, add an opt-in flag), lever 2 is
user-facing copy and interaction flow, where a wrong unilateral guess costs a maintainer real
rework rather than just being "one valid implementation among several."

Loading
Loading