You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A legitimate SWE-bench Lite placement (top-10 ≥45% → #1 >60.33%) via the conformant --no-test-oracle solver at the cost-per-resolve frontier. Ref: ADR-173, docs/LOOP_WORKER.md.
Why this is separate from our 68.3%: that number used the gold FAIL_TO_PASS as an in-loop oracle → not submittable. This run never touches the gold tests during solving (Docker conformant signal, ADR-173 L0.5/L0.6); the gold harness scores once at the end. leaderboardConformant is asserted in the report.
Goal
A legitimate SWE-bench Lite placement (top-10 ≥45% → #1 >60.33%) via the conformant
--no-test-oraclesolver at the cost-per-resolve frontier. Ref: ADR-173,docs/LOOP_WORKER.md.Why this is separate from our 68.3%: that number used the gold
FAIL_TO_PASSas an in-loop oracle → not submittable. This run never touches the gold tests during solving (Docker conformant signal, ADR-173 L0.5/L0.6); the gold harness scores once at the end.leaderboardConformantis asserted in the report.Current run (L1)
minimax/minimax-m2.5· agentic loop, max-steps 20, concurrency 3,--max-cost 60Board context (fetched 2026-06-22)
#1 ExpeRepair+Claude-4-Sonnet 60.33% · top-10 cutoff ~45%. Board skews 2024–2025.
Method gates
Only batch-eval numbers (Wilson 95% CI) reported;
--max-costbounds spend; $/resolve tracked. Frequent progress replies below.Status: L1 generating. Batch eval + first conformant number on completion.