The retrieval benchmark harness now exists (PR #125), but the claim that actually backs the positioning — the Brain improves AI coding — is still only a design in docs/VALIDATION.md §2, not code.
Goal: run the same agent on a held-out set of coding tasks that ship with executable tests, twice (Brain-injected vs not), holding model/prompt/seed constant, and report test pass-rate with n + paired difference + CI. No LLM-as-judge.
Acceptance
See docs/VALIDATION.md "Proposed methodology §2".
The retrieval benchmark harness now exists (PR #125), but the claim that actually backs the positioning — the Brain improves AI coding — is still only a design in
docs/VALIDATION.md §2, not code.Goal: run the same agent on a held-out set of coding tasks that ship with executable tests, twice (Brain-injected vs not), holding model/prompt/seed constant, and report test pass-rate with n + paired difference + CI. No LLM-as-judge.
Acceptance
docs/VALIDATION.md(or explicitly reported as null/negative if that's the truth).See
docs/VALIDATION.md"Proposed methodology §2".