fix(operator-queue): stale_id marks a genuine re-use, never the entry awaiting its terminal flip (#3024) - #3026
Conversation
… awaiting its terminal flip (#3024) Since #2915, every ask the platform ends (a cancel, a bulk cancel or an expiry) read "Re-used id" on the Resolved card. The per-agent reconcile (step 2) runs before the write-back (step 4). So in the cycle right after the platform ends a row, the agent's file still holds the ORIGINAL pending entry, the one step 4 is about to flip, and the reconcile flagged it stale_id. Nothing cleared the flag: the next cycle the entry reads the ending, and the branch looked only at pending entries. - The terminal sync index carries delivery_state and delivery_detail. - `awaits_terminal_flip(row)` is the write-back's own selection, in Python: a cancelled or expired row whose flip has not landed (delivery_state NULL or undelivered, and not entry_missing). - A pending entry is stale_id only when the write-back will not flip it. A delivered flip, a missing entry and an acknowledged row still flag. - A row flagged before this fix heals to confirmed once its entry reads the row's own ending (audited `reconciled`). A genuine re-use cannot heal: the write-back never flips an entry once its row has left the set. - A real-DB test pins `awaits_terminal_flip` to the SQL of get_terminal_items_for_agent. Fixes #3024 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
CI: 18 checks pass (including journey-smoke, pytest (head) and the regression diff). The two red checks, |
/review ReportBranch: Execution coverage
Fix mutations: 4, all red, listed in the PR body. Critical FindingsNone. Informational FindingsNone. One design refinement compared with the issue's first draft. The heal was originally going to compare Clean categories
Summary
|
vybe
left a comment
There was a problem hiding this comment.
Validated at lane B (no schema), head 2d636a4: READY. The 14 new tests drive the real _sync_agent through reconcile and write-back; with the service change reverted 8 of them fail, and with the db change reverted the parity test fails. All checks green on this head, including pytest (head), regression diff, pg-migrations and gitleaks. Merge-simulated with #3023 on top: clean, fix intact, 85/85 in the file.
Two non-blocking follow-ups: the heal only reaches a mis-flagged row whose entry is still in the agent's file (operator_queue_service.py:1484-1493), and the SQL side of the rule carries LIMIT 200 where the Python side has none (db/operator_queue.py:780-812).
Description
Since #2915 (PR #2989), every operator-queue ask the platform ends — a cancel, a bulk cancel or an expiry — read "Re-used id" (
sync_state = stale_id) on the Operations Resolved card. The Workspace projection showed it asunconfirmed.Cause.
_sync_agentreconciles the agent's file against the rows (step 2) before it writes the platform's decisions back (step 4). In the cycle right after the platform ends a row, the agent's file still holds the original pending entry, which is exactly the entry step 4 is about to update with the ending. The reconcile ran first and misread it as a reuse. Nothing ever cleared the flag: the next cycle the entry reads the ending, and the branch only looked at pending entries.Fix.
delivery_stateanddelivery_detail.awaits_terminal_flip(row)is the write-back's own selection, restated in Python. It is true for a cancelled or expired row whose flip has not landed:delivery_stateNULL orundelivered, and notentry_missing.stale_idonly when the write-back will not flip it. A delivered flip, a missing entry and an acknowledged row still count as a genuine reuse, as bug(operator-queue): a pending approval can sit unseen for days — the human's card goes stale while the container↔platform sync stays silent #2915 intended.confirmedonce its entry reads the row's own ending, audited asreconciled. A genuine reuse cannot heal this way: once its row has left the write-back set, the write-back never touches the entry again, so the reused entry stayspending.awaits_terminal_flipto the SQL ofget_terminal_items_for_agent, so the two cannot drift.Related Issue
Fixes #3024
Journey Impact
Journey Impact: extends: J05
Type of Change
Testing
New tests.
tests/unit/test_2915_operator_queue_sync_honesty.py::TestStaleIdIsOnlyAGenuineReuseadds 14 tests. The rows sit in both the sync index and the write-back set, as they do in production. The earlier harness only ever filled one of them, which is how this shipped.stale_id, across all three awaiting delivery states, and that cycle still delivers the ending into it.stale_id.Wider suites. All 47 unit suites that touch the operator queue: 1,453 passed.
Live, on a local instance running this branch.
diverged → stale_idstamped milliseconds beforewritten_back, in the same cycle) moved toconfirmedwith areconciledaudit row on the first poll cycle after the reload. Nostale_idrow remains on the instance.cancelled/confirmed/delivered. Its only audit row iswritten_back, with nodiverged, and the agent's file entry readscancelled.Mutation (each on the line that applies the fix, restored byte-identically):
awaits_terminal_flip→ the 6test_the_original_entry_awaiting_its_flip_is_not_a_reused_idcases go red.awaits_terminal_flipignoresdelivery_state→test_a_pending_entry_the_write_back_will_not_flip_is_a_reused_id[flip-already-delivered]and the parity test go red.test_a_row_misflagged_since_2915_heals_once_its_entry_reads_the_endingcases go red.test_the_awaiting_flip_rule_is_the_write_back_selectiongoes red./cso --diff: no findings (docs/security-reports/cso-diff-2026-09-25-3024-stale-id-misflag.md).CI note.
schema-parityis currently red on every open PR because the test environment resolves SQLAlchemy 2.1.0, whose defaultpostgresql://driver is psycopg 3 (#3015 caps it). This PR has no migration.Checklist
stale_idrow inoperating-room.md's reconcile table now states the rule, and a new row covers the heal.reconciledaudit row)🤖 Generated with Claude Code