Skip to content

fix(operator-queue): stale_id marks a genuine re-use, never the entry awaiting its terminal flip (#3024) - #3026

Merged
vybe merged 2 commits into
devfrom
feature/3024-stale-id-misflag
Sep 27, 2026
Merged

vybe merged 2 commits into
devfrom
feature/3024-stale-id-misflag

Conversation

@webmixgamer

@webmixgamer webmixgamer commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Description

Since #2915 (PR #2989), every operator-queue ask the platform ends — a cancel, a bulk cancel or an expiry — read "Re-used id" (sync_state = stale_id) on the Operations Resolved card. The Workspace projection showed it as unconfirmed.

Cause. _sync_agent reconciles the agent's file against the rows (step 2) before it writes the platform's decisions back (step 4). In the cycle right after the platform ends a row, the agent's file still holds the original pending entry, which is exactly the entry step 4 is about to update with the ending. The reconcile ran first and misread it as a reuse. Nothing ever cleared the flag: the next cycle the entry reads the ending, and the branch only looked at pending entries.

Fix.

  • The terminal sync index also carries delivery_state and delivery_detail.
  • New awaits_terminal_flip(row) is the write-back's own selection, restated in Python. It is true for a cancelled or expired row whose flip has not landed: delivery_state NULL or undelivered, and not entry_missing.
  • A pending entry is stale_id only when the write-back will not flip it. A delivered flip, a missing entry and an acknowledged row still count as a genuine reuse, as bug(operator-queue): a pending approval can sit unseen for days — the human's card goes stale while the container↔platform sync stays silent #2915 intended.
  • Self-heal. A row flagged before this fix moves to confirmed once its entry reads the row's own ending, audited as reconciled. A genuine reuse cannot heal this way: once its row has left the write-back set, the write-back never touches the entry again, so the reused entry stays pending.
  • One rule, two spellings. A real-DB test pins awaits_terminal_flip to the SQL of get_terminal_items_for_agent, so the two cannot drift.

Related Issue

Fixes #3024

Journey Impact

Journey Impact: extends: J05

Type of Change

  • Bug fix (non-breaking change that fixes an issue)
  • New feature (non-breaking change that adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)
  • Documentation update

Testing

  • I have tested this locally
  • New tests added (if applicable)
  • All existing tests pass
  • Every new test executes the changed path

New tests. tests/unit/test_2915_operator_queue_sync_honesty.py::TestStaleIdIsOnlyAGenuineReuse adds 14 tests. The rows sit in both the sync index and the write-back set, as they do in production. The earlier harness only ever filled one of them, which is how this shipped.

  • The original entry waiting for its flip is not stale_id, across all three awaiting delivery states, and that cycle still delivers the ending into it.
  • A delivered flip, a missing entry and an acknowledged row are still stale_id.
  • A mis-flag heals.
  • A genuine flag does not heal, and neither does an entry the agent closed differently.
  • The parity test runs on a real DB.

Wider suites. All 47 unit suites that touch the operator queue: 1,453 passed.

Live, on a local instance running this branch.

  • The three rows mis-flagged earlier the same day (diverged → stale_id stamped milliseconds before written_back, in the same cycle) moved to confirmed with a reconciled audit row on the first poll cycle after the reload. No stale_id row remains on the instance.
  • A fresh platform cancel reads cancelled / confirmed / delivered. Its only audit row is written_back, with no diverged, and the agent's file entry reads cancelled.

Mutation (each on the line that applies the fix, restored byte-identically):

  • The reconcile flags regardless of awaits_terminal_flip → the 6 test_the_original_entry_awaiting_its_flip_is_not_a_reused_id cases go red.
  • awaits_terminal_flip ignores delivery_state → test_a_pending_entry_the_write_back_will_not_flip_is_a_reused_id[flip-already-delivered] and the parity test go red.
  • The heal branch is removed → both test_a_row_misflagged_since_2915_heals_once_its_entry_reads_the_ending cases go red.
  • The sync index drops the delivery columns → test_the_awaiting_flip_rule_is_the_write_back_selection goes red.

/cso --diff: no findings (docs/security-reports/cso-diff-2026-09-25-3024-stale-id-misflag.md).

CI note. schema-parity is currently red on every open PR because the test environment resolves SQLAlchemy 2.1.0, whose default postgresql:// driver is psycopg 3 (#3015 caps it). This PR has no migration.

Checklist

  • My code follows the project's style guidelines
  • I have updated the documentation (if applicable). The stale_id row in operating-room.md's reconcile table now states the rule, and a new row covers the heal.
  • I have not committed any sensitive data (API keys, credentials, etc.)
  • I have added appropriate logging for new functionality (the heal writes the existing reconciled audit row)

🤖 Generated with Claude Code

… awaiting its terminal flip (#3024)

Since #2915, every ask the platform ends (a cancel, a bulk cancel or an
expiry) read "Re-used id" on the Resolved card. The per-agent reconcile
(step 2) runs before the write-back (step 4). So in the cycle right after
the platform ends a row, the agent's file still holds the ORIGINAL pending
entry, the one step 4 is about to flip, and the reconcile flagged it
stale_id. Nothing cleared the flag: the next cycle the entry reads the
ending, and the branch looked only at pending entries.

- The terminal sync index carries delivery_state and delivery_detail.
- `awaits_terminal_flip(row)` is the write-back's own selection, in Python:
  a cancelled or expired row whose flip has not landed (delivery_state
  NULL or undelivered, and not entry_missing).
- A pending entry is stale_id only when the write-back will not flip it.
  A delivered flip, a missing entry and an acknowledged row still flag.
- A row flagged before this fix heals to confirmed once its entry reads
  the row's own ending (audited `reconciled`). A genuine re-use cannot
  heal: the write-back never flips an entry once its row has left the set.
- A real-DB test pins `awaits_terminal_flip` to the SQL of
  get_terminal_items_for_agent.

Fixes #3024

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@webmixgamer

Copy link
Copy Markdown
Contributor Author

CI: 18 checks pass (including journey-smoke, pytest (head) and the regression diff). The two red checks, pg-migrations and schema-parity, are the SQLAlchemy 2.1.0 test-environment breakage (No module named 'psycopg' before any revision runs). It hits every open PR right now, and #3015 caps the version. I'll re-run both once #3015 lands.

@webmixgamer

Copy link
Copy Markdown
Contributor Author

/review Report

Branch: feature/3024-stale-id-misflag → dev (merge-base 0378f2550, the current dev tip)
Files Changed: 7 (+251 / −11)
Scope: CLEAN. The intent is #3024: every ask the platform ended read "Re-used id". The diff narrows the rule to a genuine reuse, heals the rows flagged before this fix, and updates the reconcile table in operating-room.md.
Plan Completion: all four of the issue's acceptance criteria are done.

Execution coverage

changed symbol executed by live consumer verdict
awaits_terminal_flip TestStaleIdIsOnlyAGenuineReuse: waiting-for-flip ×6, reuse ×3, and the real-DB parity test the terminal branch of _sync_agent ✅ executed
the heal branch …heals_once_its_entry_reads_the_ending ×2, …genuine_reuse_flag_is_not_cleared…, …closed_differently_is_not_healed _sync_agent ✅ executed
the terminal sync index's delivery_state / delivery_detail the real-DB parity test _sync_agent, via db.get_operator_queue_sync_index_for_agent ✅ executed

Fix mutations: 4, all red, listed in the PR body.

Critical Findings

None.

Informational Findings

None.

One design refinement compared with the issue's first draft. The heal was originally going to compare sync_updated_at with delivery_updated_at. Reading get_terminal_items_for_agent showed that a delivered row leaves the write-back set, so the write-back never touches a reused entry and it can never meet the heal condition. The simpler rule is enough, and the issue body was corrected to match.

Clean categories

  • One rule, two spellings: the Python rule restates the write-back's SQL. A parity test on a real DB pins the two together, so they can't drift apart.
  • SQL: two more selected columns, parameterised.
  • Races: the reconcile is leader-locked, and every write is edge-triggered.
  • Incomplete fix (4.14): _sync_agent is the only writer of stale_id, so no other call site needs the change.
  • Existing behaviour kept: a reused id is still never re-admitted, and a reuse of an acknowledged row still flags.
  • Security (/cso --diff): no findings. The report is in docs/security-reports/cso-diff-2026-09-25-3024-stale-id-misflag.md.

Summary

@webmixgamer
webmixgamer requested review from dolho and vybe September 26, 2026 10:22

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validated at lane B (no schema), head 2d636a4: READY. The 14 new tests drive the real _sync_agent through reconcile and write-back; with the service change reverted 8 of them fail, and with the db change reverted the parity test fails. All checks green on this head, including pytest (head), regression diff, pg-migrations and gitleaks. Merge-simulated with #3023 on top: clean, fix intact, 85/85 in the file.

Two non-blocking follow-ups: the heal only reaches a mis-flagged row whose entry is still in the agent's file (operator_queue_service.py:1484-1493), and the SQL side of the rule carries LIMIT 200 where the Python side has none (db/operator_queue.py:780-812).

@vybe
vybe merged commit 84d9399 into dev Sep 27, 2026
26 checks passed
@webmixgamer
webmixgamer deleted the feature/3024-stale-id-misflag branch September 27, 2026 14:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants