Add weak-coherent-discrimination-design task - #19
Conversation
Companion design task to heterodyne-shot-noise-analysis: plan a shot-noise-limited discrimination experiment for a ~1.4 pW coherent state at 795 nm, graded against a frozen 11-item expert rubric by a quote-verified LLM judge with a fail-closed score validator. Task package authored by Wenjun Ke; submitted from their latest archive with three mechanical CI fixes (taxonomy vocabulary mapping in task.md, ruff findings in score.py and check_grader.py). rubric.yaml and judge.py are untouched; the rubric.sha256 freeze gate still verifies. Registers the task in registry.json and validate_repository.py, following the pattern of #8. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Replace the prose-only custom grader with a multi-module scientific SWE task, human expert plans, native plan-and-trajectory judging, eight deterministic outcome tests, and a mentor skill. Freeze claim-level source provenance and require all planning, research, trajectory, and execution gates to pass. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
Match navigation ledgers to trusted trajectories by chronology and target while allowing ACP to normalize model-visible shell reads, searches, and listings into native event kinds. Failed or invented calls still fail the all-pass trajectory gate. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
Use NAVIGATION.md only as a chronological target and finding index. Score action kinds, batching, counts, and edit timing from the host trajectory so self-authored tool labels cannot either earn credit or create false failures. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
Require the agent to research before producing only PLAN.md, move the complete LLM judge contract to rubric.toml, and compose eight deterministic plan checks with nine all-pass trajectory, research, and planning criteria. Remove the legacy research, navigation, execution, and duplicate verifier-plan artifacts; add the expert oracle plan, future output contract, mentor guidance, contributor documentation, and CI enforcement for the standard planning-task layout. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a3bfde8bb9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Replace the task-local judge and alternate rubric files with the draft native agentic review contract, while keeping the deterministic execution reward separate. Standardize the contributor, maintainer, and validation pipeline around task.md, oracle/plan.md, verifier/rubric.json, and test_outputs.py. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c3bb40a5bc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Remove unnecessary PLAN-specific protocol and duplicated rubric schema machinery. Treat rubric.json as the sole generic task rubric input while retaining only the minimal planning-task layout guidance and task-specific PLAN and trajectory criteria. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
|
@codex review |
Prevent agents from hiding unauthorized edits by committing them. The held-out verifier now diffs the working tree against the content-addressed Docker starter tree with replacement objects disabled, while still allowing only PLAN.md. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2efc2f6f8e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Use the current standalone rubric.json schema, preserve the required pre-plan research gate, and tighten the mentor control and deterministic terminology checks based on native end-to-end evaluation. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
Expose a frozen primary record for author measurements, enumerate the expert plan's public output fields, refresh the frozen starter-tree guard, and clarify the with-skill research and final-audit guidance without adding any rubric mechanism outside rubric.json. Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fb7d969abe
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| - Test the submitted plan contract deterministically. Put process and semantic | ||
| plan criteria only in `rubric.json`, where the native reviewer can use the | ||
| trusted host trajectory. |
There was a problem hiding this comment.
Scope plan-only verifier rules to planning tasks
These rules sit in the general verifier/ section, so contributors creating computational, data-analysis, or artifact tasks are now instructed to test a nonexistent plan artifact and move semantic checks into rubric.json. The same unscoped instruction appears in .agents/skills/task-creator/SKILL.md, contradicting the repository's outcome-based verifier contract; keep the original general verifier guidance here and place these rules exclusively under the planning/research profile.
AGENTS.md reference: AGENTS.md:L43-L44
Useful? React with 👍 / 👎.
| With the frozen efficiency assumptions, balanced heterodyne has 8.824 detected | ||
| photons, a detected-state Helstrom error of 3.68e-5, and an implementable error | ||
| of 1.784e-2 after the image-band penalty. Threshold direct counting has 6.624 | ||
| detected photons, a Helstrom error of 3.32e-4, and a receiver error of 9.14e-4 |
There was a problem hiding this comment.
Specify the efficiencies used in the receiver comparison
At the stated 11.366 incident photons, the two detected-photon values imply total efficiencies of 0.7763 and 0.5828, but neither value nor its optical/detector/mode factorization is provided or sourced anywhere in the plan; the only supplied configuration instead gives 0.90 * 0.80 * 0.95**2 = 0.6498. Consequently, an implementer following the public schema cannot reproduce these Helstrom and receiver-error figures without inventing loss assumptions, so the reference plan should state and classify each architecture's efficiency factors explicitly.
AGENTS.md reference: AGENTS.md:L40-L41
Useful? React with 👍 / 👎.
Outcome
This PR adds the
weak-coherent-discrimination-designplanning/research task, authored by Wenjun Ke (@wenjun-ke), and migrates it to BenchFlow's standalonerubric.jsoncontract.rubric.jsonis generic and independent ofPLAN.md. This task happens to request onePLAN.md, so its task-specific criteria evaluate that plan and the trusted trajectory. No other task file defines rubric configuration, reviewer execution, or scoring logic.Task contract
The agent must:
/app/receiver_lab/PLAN.md;RESEARCH.md/NAVIGATION.md; anddesign.yaml,output/summary.json, andoutput/trials.csvcontract for a future execution.The contributor provides four core human-authored inputs:
task.md— task and final-output contract;oracle/plan.md— reviewed expert plan;oracle/solve.sh— oracle control; andverifier/rubric.json— the sole rubric definition.verifier/test.shandtest_outputs.pyseparately implement 8 deterministic final-artifact checks. The verifier directory contains exactly:The old task-local judge, Reward Kit,
verifier.md, TOML/YAML rubrics, evaluator credential plumbing, navigation ledger, and trajectory truncation logic are removed.Native review architecture
This task's
rubric.jsoncontains 9 yes/no items covering research before the firstPLAN.mdwrite, codebase and QuTiP synthesis, source grounding, receiver comparison, dependency-aware stages, risks, scope, validation, and the future output contract.The JSON contains only the native reviewer configuration, pass threshold, and task criteria. BenchFlow supplies the submitted workspace, held-out oracle reference, and complete trusted trajectory at runtime. The binary artifact reward remains separate from the plan/trajectory review score.
Native validation used the BenchFlow
feat/rubric-reviewimplementation. It auto-discovered this single JSON file, ran all 9 criteria inindividualmode, retained trajectory snapshots, enforced the required research gate, and emitted separatereward,review, andreview_passedmetrics. The native BenchFlow rubric/config test suite passes 88/88 locally.The released BenchFlow 0.6.5 does not contain that review stage; merge/deployment of this task therefore requires the native
rubric.jsonsupport to land. This task does not duplicate the runner while waiting for that release.Provenance and anti-cheat controls
docs/apparatus_record.mdis an agent-visible frozen primary record forAUTHOR-APPARATUS-LOG, including measurement scope and limitations.component_catalog.jsonpoints directly to that record instead of held-out oracle provenance.docs/output_contract.mdand the expert oracle enumerate the concrete nested summary fields and CSV reconciliation contract.b22e09f7dd29b9c7e23c24546cf8f3ea3e42acbf, with Git replacement objects disabled. An agent that edits README, commits the edit, and then writes PLAN.md still fails withonly PLAN.md may change.Complexity
The environment is a partially implemented scientific package with public tests, a frozen component/source catalogue, an author apparatus record, and a pinned QuTiP 5.3.0 snapshot: 491 files and about 79k Python lines. The plan must connect receiver, budget, simulation, reporting, and quantum-reference boundaries to upstream implementations, exports, tests, source provenance, and reproducible outputs.
Validation
Exact head:
fb7d969abeff5c87951eca9bd0077dc49dc63490Task digest:
sha256:e64041d369a31cf3221105d128a0dfa529a93fe3747fa0bb1f74eac6391a7b5bRubric SHA-256:
7af9eba7ebbf764cd16d66fc51817686b77c3d88fa5f7609fda8b5b972e23daereward=1.0, 8/8 deterministic checks.sha256:5992f4d6...):reward=1.0,review=1.0,review_passed=1.0, 9/9 rubric items yes, and 68 useful pre-plan calls after one duplicate was excluded.rubric.jsonis byte-identical.git diff --checkpass locally.Every commit in this PR includes
Co-authored-by: Wenjun Ke <u3597436@connect.hku.hk>.