Skip to content

fix(#5064): distinct contract for post-quorum local commit failure (committed-but-local-apply-failed) - #5075

Merged
lvca merged 8 commits into
mainfrom
fix/5064-committed-remotely
Jul 7, 2026
Merged

lvca merged 8 commits into
mainfrom
fix/5064-committed-remotely

Conversation

@lvca

@lvca lvca commented Jul 7, 2026

Copy link
Copy Markdown
Member

Implements option A from the #5064 design discussion.

The problem

When the Raft quorum durably commits a transaction but the leader's local phase-2 apply fails, the application received a generic commit failure - told the commit failed while the data was committed cluster-wide. Worse: an application-level retry of the same records inserted duplicates, because the #4940 rollback reset their identities to provisional and the retry re-inserted records the cluster already held.

The fix

  • TransactionCommittedRemotelyException (engine exception package, catchable without an ha-raft dependency): states the transaction IS committed cluster-wide, whether local pages were reconciled, and that it must NOT be retried. reconcileLeaderPagesAfterPhase2Failure now reports its outcome instead of swallowing it.
  • A remotelyCommitted durability regime in TransactionContext, set by the HA layer after quorum commit: a local failure past that point releases resources without rolling back user-held record identities (they are the identities the cluster committed) and without fencing (no orphaned local WAL record exists; the Raft layer reconciles pages from the replicated payload). Slots between the walAppended and fence-refused/rollback regimes - [tx] Phase-2 commit failure uses reset() instead of rollback(), leaving dangling RIDs on user documents #4940/fix(#4936,#4937): the WAL append is the commit's point of no return #5053 semantics unchanged for non-replicated databases. The ALL-quorum recovery path gets the same boundary shift.

Verification

Red-first: Issue5064CommittedRemotelyContractIT (3-node cluster, single-shot post-quorum fault) fails on pre-fix behavior (raw exception escapes, identity reset) and passes with the fix - distinct type, actionable message, identity preserved, all nodes converge on the committed data after step-down. HA phase-2 battery green (Issue4740Phase2ReconcileIT, Issue5018Phase2ConflictMessageTest, DatabaseReconcilerTest); engine commit-path battery green (WalCommitOrdering, Issue4940, Issue4959, ExplicitLocking, IsolationContract).

Conflict risks

TransactionContext (one new field + one new regime branch in the finally) - the follow-up fan-out workers do not touch it.

When the Raft quorum durably committed a transaction but the leader's
LOCAL phase-2 apply failed, the application received a generic commit
failure - told the commit failed while the data was committed
cluster-wide. Worse, an application-level retry of the same records
INSERTED DUPLICATES: the #4940 rollback reset their identities to
provisional, so the retry re-inserted records the cluster already held.

Option A from the design discussion, as approved:

- New TransactionCommittedRemotelyException (engine exception package, so
  applications can catch it without an ha-raft dependency): 'committed
  cluster-wide, do NOT retry, reload the records', with the local
  reconciliation outcome in the message
  (reconcileLeaderPagesAfterPhase2Failure now reports success/failure
  instead of swallowing silently).

- TransactionContext gains a remotelyCommitted durability regime, set by
  the HA layer after quorum commit and before the local phase 2: a local
  failure past that point releases resources WITHOUT rolling back
  user-held record identities (the cluster committed them) and WITHOUT
  fencing (no orphaned local WAL record exists; the Raft layer reconciles
  the pages from the replicated payload). The regime slots between
  walAppended and the fence-refused/rollback regimes in commit2ndPhase's
  finally, composing with #4940/#5053 semantics unchanged for
  non-replicated databases. The ALL-quorum recovery path gets the same
  boundary shift.

Red-first: Issue5064CommittedRemotelyContractIT (3-node cluster,
single-shot post-quorum fault) fails on the pre-fix behavior (raw
ConcurrentModificationException escapes, identity reset); green with the
fix - distinct type, actionable message, identity preserved, and all
three nodes converge on the committed data after the step-down. A
baseline write pre-registers the dictionary name so the internal
dictionary transaction does not consume the single-shot fault.

HA phase-2 battery green (Issue4740Phase2ReconcileIT,
Issue5018Phase2ConflictMessageTest, DatabaseReconcilerTest); engine
commit-path battery green (WalCommitOrdering, Issue4940, Issue4959,
ExplicitLocking, IsolationContract).
@mergify

mergify Bot commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

Tick the box to add this pull request to the merge queue (same as @mergifyio queue).

  • Queue this pull request

@codacy-production

codacy-production Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

🟢 Metrics 0 complexity

Metric Results
Complexity 0

View in Codacy

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

Gemini encountered an error creating the review. You can try again by commenting /gemini review.

@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: PR #5075 - distinct contract for post-quorum local commit failure

Solid, well-scoped fix with excellent inline documentation. The core mechanism is right: making the new exception extend TransactionException (not NeedRetryException) means the db.transaction() retry loop (LocalDatabase:1328, which only catches NeedRetryException | DuplicatedKeyException) will not auto-retry it, which is exactly what stops the duplicate-insert-on-retry problem from #5064. Nice.

A few points worth addressing before merge.

1. remotelyCommitted is never reset - latent correctness hazard (medium)

The new field is only ever set to true (RaftReplicatedDatabase:432 and :472), and is cleared by neither reset() (TransactionContext:1070) nor begin() (:147). Meanwhile the base TransactionContext at stack index 0 is reused across transactions on the same thread - LocalDatabase.begin()/commit() both operate on current.getLastTransaction(), and payload.tx() is that same getLastTransaction() object (RaftReplicatedDatabase:353). So once the flag flips true it stays true for the life of that thread's base context.

Today this is masked because the Raft leader path re-sets it to true before every phase-2, so the value happens to be correct whenever commit2ndPhase's finally runs on the leader. But it's fragile: any future/edge path that reaches the commit2ndPhase finally with committed=false && walAppended=false and the stale flag still true would silently take the reset() branch (:1013) instead of rollback() - skipping the #4940 identity reset and the fence, on a transaction that was never actually remotely committed.

Suggest making the flag single-transaction-scoped defensively: clear it in reset() (the common teardown) and/or set remotelyCommitted = false at the top of begin(). Cheap, and removes the "sticky flag on a pooled object" footgun. Every other mutable regime signal in this method (committed, walAppended) is a method-local; this one being an un-cleared instance field is the odd one out.

2. The new engine-side branch is not actually exercised by the IT (medium)

The IT injects the fault at RaftReplicatedDatabase:436-437, which throws before payload.tx().commit2ndPhase(...) at :439 is ever called. So the new remotelyCommitted branch in TransactionContext.commit2ndPhase's finally (:1013-1020) never executes in the test. The identity-preservation assertion passes for a different reason: commit2ndPhase never runs, so no rollback path runs at all, and the manual begin()/commit() bypasses the retry wrapper. In other words the test validates the HA exception contract (which is valuable) but gives the engine change zero coverage.

The branch is reachable in production (e.g. a genuine ConcurrentModificationException from validateAndBumpVersions at :937, before walAppended). Consider a test that faults from inside commit2ndPhase after entry but before the WAL append, so the new branch is genuinely covered - or a focused TransactionContext-level unit test for the regime.

3. "without fencing" only holds pre-WAL-append (low)

The field comment and PR description say the remotely-committed regime releases "without fencing." But walAppended is checked before remotelyCommitted in the finally, so a phase-2 failure that occurs after writeTransactionToWAL (e.g. publishPages throwing) still takes the walAppended fence branch (:994). Identity is still preserved there (that branch also does reset(), not rollback()), so the user-facing contract is intact - but the "no fence" statement is only true for failures before the WAL append. Worth tightening the comment to say so, or confirming that post-append failures intentionally still fence.

4. Minor

  • applyLocallyAfterMajorityCommit (:467) sets the flag and reconciles/steps-down on failure but does not surface TransactionCommittedRemotelyException. That looks intentional (background ALL-quorum recovery, not user-driven), but it means that path doesn't deliver the same "do not retry" signal - a one-line comment noting why would help the next reader.
  • The message-building ternary at :451-454 reads a little densely; extracting the reconciled/not-reconciled suffix into a local would improve readability, but not blocking.

Verdict

Good fix, correct core design, strong comments. The reset-the-flag point (#1) and the test-coverage gap (#2) are the two I'd want resolved before merge; the rest are polish.

🤖 Generated with Claude Code

@lvca lvca self-assigned this Jul 7, 2026
@lvca lvca added this to the 26.7.2 milestone Jul 7, 2026
…y covered; comments tightened

All four points taken; the two mediums were exact:

1. remotelyCommitted was a sticky flag on the REUSED per-thread base
   context - set by the Raft path, cleared by nothing. Any later path
   reaching the finally with committed=false && walAppended=false and the
   stale flag would silently skip the #4940 rollback on a transaction
   never remotely committed. Cleared at begin() and in reset() (belt and
   braces), with the hazard documented at both.

2. The engine-side regime branch had ZERO test coverage - the IT's fault
   fires before commit2ndPhase runs, so the identity assertion passed
   trivially. New engine test drives the branch for real: phase 1, the
   HA-layer flag set, a conflicting-version injection failing
   validateAndBumpVersions PRE-WAL - asserting the identity survives, no
   fence, and (second half, on the SAME reused context) that a fresh plain
   transaction still rolls back identities per #4940. Red-first: without
   the begin()/reset() clearing the second half fails on the sticky flag.

3. The 'without fencing' comment now states its true scope: only the
   pre-append branch skips the fence; a post-append failure intentionally
   still takes the walAppended fence branch (an orphaned local WAL record
   exists regardless of the remote commit) and preserves identities there
   too via reset().

4. applyLocallyAfterMajorityCommit documents why it does NOT surface the
   new exception (background ALL-quorum recovery, no user caller); the
   message ternary is extracted to a local.

WalCommitOrderingTest 6/6; engine compile + ha-raft compile clean.
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Code Review: PR #5075 - distinct contract for post-quorum local commit failure

Solid fix. The durability-boundary model is clearly reasoned, the failure regimes in commit2ndPhase's finally are well-ordered and documented, and the test coverage (engine unit test for the pre-WAL path + 3-node HA IT for the end-to-end contract) is genuinely good - including the sticky-flag regression guard. A few points below.

🔴 Broken indentation in TransactionContext.java (must fix)

The two inserted assignments lost their indentation and the comments are over-indented, inconsistent with the surrounding 4-space style, and likely to trip the pre-commit hooks.

begin() (lines ~148-152):

        // #5064: the base context is REUSED across transactions on this thread ...
    // silently skip the #4940 rollback. Cleared here AND in reset() (belt and braces).
    remotelyCommitted = false;
this.isolationLevel = isolationLevel;   // <-- no indentation

reset() (lines ~1078-1079):

        remotelyCommitted = false;      // <-- 8-space indent
status = STATUS.INACTIVE;              // <-- no indentation

Both should be a clean 4-space indent (comment block and the two statements). Worth double-checking no similar artifacts slipped in elsewhere.

🟡 Minor: double blank line in WalCommitOrderingTest.java

Two blank lines before the new @Test remotelyCommittedRegimePreservesIdentitiesOnPreWalFailure method - collapse to one to match the rest of the file.

🟢 Design note (not a blocker): "no fence" applies only pre-append

The remotelyCommitted branch sits after the walAppended branch, so it only reaches the no-fence reset() when the local apply fails before the WAL append. In production, a phase-2 failure during page publishing happens after the append and correctly takes the walAppended fence branch (which also preserves identities via reset()). The single-shot IT injects the fault via TEST_PHASE2_COMMIT_FAULT before commit2ndPhase, i.e. the pre-append path - so the post-append + remotely-committed path (the more common real-world case) isn't directly exercised. The comments already acknowledge this is intentional and identity preservation holds either way, so this is fine - just flagging the IT's coverage boundary.

🟢 Verified good

  • remotelyCommitted cannot leak across transactions: cleared in begin() and reset(), and resetAndFireCallbacks() -> reset() clears it on the success path too. Belt-and-braces is justified.
  • TransactionCommittedRemotelyException extends TransactionException keeps it catchable without an ha-raft dependency, is correctly re-thrown as-is by the existing catch (TransactionException) rethrow, then wrapped at the HA layer.
  • reconcileLeaderPagesAfterPhase2Failure now returning its outcome and the message reflecting reconciled vs. not-reconciled is accurate, and degrades gracefully if reconcile fails.
  • Exception message is actionable ("committed cluster-wide", "Do NOT retry: reload the records and continue"), matching the Javadoc contract.

Overall: approve once the indentation is cleaned up.

Automated review by Claude - please verify findings before acting.

The critical was mine: the round-1 automated edit ate the indentation of
the statements following both insertion points (begin() and reset()) and
over-indented the comment block. Both sites restored to clean 4-space
style; a column-0 sweep across both touched files confirms no other
artifact slipped in. The double blank line before the new test method is
collapsed.

The design note (no-fence applies only pre-append; the IT exercises the
pre-append path while post-append cases take the walAppended fence branch
with identities equally preserved) matches the intent documented in round
1 - acknowledged, no change.

WalCommitOrderingTest 6/6, Issue4940Phase2FailureRollbackTest green.
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: PR #5075 - distinct contract for post-quorum local commit failure

Reviewed the full 3-commit diff against the surrounding commit-path code. This is a well-scoped, high-quality fix for a genuine data-integrity bug (application retries inserting duplicates of already-committed records). The TDD/red-first approach, the composition with the existing #4940/#5053 regimes, and the extensive inline rationale are all excellent. A few observations below, mostly minor.

Correctness - looks solid

  • The remotelyCommitted branch is placed after walAppended in commit2ndPhase's finally, so the "only reachable pre-append" claim in the comment is structurally enforced by the if/else ordering. Good.
  • The sticky-flag hazard (base TransactionContext is reused per-thread) is correctly closed by clearing in both begin() and reset(), and resetAndFireCallbacks() -> reset() clears it on the success path too. The second half of the engine test (a fresh tx on the same reused context still rolls back per #4940) is a good regression guard.

1. New IT is missing @Tag("slow") (should fix)

Issue5064CommittedRemotelyContractIT is a 3-node cluster test with leader election and a post-fault step-down - multi-second by nature. Its sibling Issue4740Phase2ReconcileIT (referenced in the PR description) and RaftDivergedFollowerRecoveryIT both carry @Tag("slow"). Per CLAUDE.md, noticeably-slow functional regression tests should be tagged. Please add @Tag("slow") (class level) for CI consistency.

2. Engine test leaks a poisoned page into the shared PageManager (minor)

In remotelyCommittedRegimePreservesIdentitiesOnPreWalFailure, two conflicting pages are pushed into PageManager.INSTANCE.putPageInReadCache(...), but the finally only removes the first (victimPageId[0]); the second (conflicting2 / victim2's pageId) is never removed. Since PageManager.INSTANCE is a process-wide singleton, a stale mismatched-version page can linger and risk cross-test pollution. Suggest tracking and removing victim2's pageId too (mirror the existing victimPageId[0] cleanup).

3. Remote/wire propagation of the new exception (please confirm)

The whole value of the new distinct type is that the application can catch it. The IT exercises the embedded API on the leader node only. When the committing client is remote (HTTP/Postgres/Bolt/redirected follower), does TransactionCommittedRemotelyException survive serialization back to the caller as its distinct type (and importantly, not get mapped to a retryable error)? If there is existing exception-mapping coverage that guarantees this, a one-line pointer would be reassuring; if not, this is the path most likely to silently regress the contract for real deployments.

4. Post-append remotely-committed path is documented but untested (consider)

The comments carefully explain that a remotely-committed tx that fails after the local WAL append takes the walAppended fence branch instead (identities still preserved). Only the pre-append path is exercised by tests. Worth noting: in that post-append case the DB is fenced, so reconcileLeaderPagesAfterPhase2Failure -> applyChanges may itself fail and the thrown message degrades to "reconciliation ALSO failed - steps down". That branch is handled by the ternary, but adding a test for the post-append remotely-committed case would lock in the fence + step-down interaction you describe.

Nits

  • TransactionCommittedRemotelyException only defines the (String, Throwable) constructor - fine for current usage; add a message-only constructor only if other call sites need it later.
  • Message text and release note are clear and actionable ("Do NOT retry: reload the records").

Overall: strong fix, correct approach, good tests. Items 1 and 2 are quick; 3 and 4 are worth a look before merge.

…e contract made explicit

1. The 3-node IT carries @tag(slow) at class level, consistent with its
   siblings (Issue4740Phase2ReconcileIT, RaftDivergedFollowerRecoveryIT).

2. The engine test's finally now evicts BOTH poisoned pages from the
   process-global PageManager (the second conflicting page leaked).

3. Wire propagation - the review's 'please confirm' investigated and
   answered with code, because the answer was 'partially': the new
   exception does NOT extend NeedRetryException, so it was never mapped to
   the retryable 503 - but with no explicit mapping it fell into the
   generic 5xx branch, and HTTP clients and load balancers routinely retry
   5xx: the exact duplicate-insert hazard over the wire. The HTTP handler
   now maps TransactionCommittedRemotelyException to 409 with a
   do-not-retry detail and the exception class in the JSON payload - the
   same non-retryable rationale as the DuplicatedKeyException 409 (#4350).

4. Post-append remotely-committed path: declined as a separate test with
   reasoning - past the local append the walAppended branch runs
   IDENTICALLY regardless of the remote flag (fence + identity-preserving
   reset), and that branch is already covered by the #5053 fence tests;
   the only remotely-committed-specific behavior is pre-append, which both
   the IT and the engine test exercise.

Nit acknowledged: the single (String, Throwable) constructor stays until
another call site needs more.

WalCommitOrderingTest 6/6, the tagged IT green, server module compiles.
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: PR #5075 - distinct contract for post-quorum local commit failure

Thorough, well-scoped fix. The core idea (a new durability regime that sits between walAppended and the #4940 rollback, so a post-quorum local failure releases resources without resetting user-held identities or fencing) is sound, and the layered comments make the intent easy to follow. The sticky-flag hazard on the reused per-thread base context is handled correctly by clearing in both begin() and reset(), and the engine test explicitly guards that regression on the same reused context. Nice.

Correctness - looks right

  • Flag ordering in commit2ndPhase's finally is correct: committed -> walAppended (fence) -> remotelyCommitted (reset, no fence) -> fence-refused -> #4940 rollback. A post-append remotely-committed failure intentionally falls into the walAppended fence branch, which also preserves identities via reset(). Consistent.
  • setRemotelyCommitted(true) is set only after replicateTransaction returns without throwing (quorum reached at RaftReplicatedDatabase.java:385), so the flag can never be set on a transaction that did not reach quorum. Good.
  • The catch (TransactionCommittedRemotelyException) in AbstractServerHttpHandler is correctly ordered before the generic catch (TransactionException) (line 291 vs 354), and since the new type does not extend NeedRetryException it is not mapped to the retryable 503. The 409 mapping (matching the DuplicatedKeyException/HTTP API returns 503 Service Unavailable for DuplicatedKeyException; should be 409 Conflict #4350 precedent) is the right call to keep clients and load balancers from re-driving the write.
  • RID consistency holds: the leader assigns the optimistic RIDs and ships them in the replication payload, so the locally-preserved identities are exactly the cluster-committed ones.

Test coverage gaps (main feedback)

  1. The HTTP 409 mapping is untested. The AbstractServerHttpHandler change is the wire-facing half of the fix - and the commit message calls out that a wrong status code is "the exact duplicate-insert hazard over the wire" - yet no server-module test asserts a 409 (rather than 5xx) is returned when the exception surfaces. A focused handler test (or asserting the status via the HTTP layer in the IT) would lock in the behavior against a future refactor that reorders the catch blocks.
  2. applyLocallyAfterMajorityCommit (the ALL-quorum recovery path) is unexercised for the new flag. It now sets remotelyCommitted too, but deliberately does NOT surface the exception. Nothing asserts that a failure there steers the finally away from the identity rollback while staying silent to callers. Given the divergence stakes on that path, a targeted test would be worthwhile even if the main path is well covered.

Minor notes

  • TransactionCommittedRemotelyException embeds the full payload.tx().toString() in the message. Fine for diagnostics; just be aware it can be verbose in logs/JSON payloads.
  • The remotelyCommitted branch skips the modified-record reload that the #4940 rollback performs. That is correct here (the in-memory content already matches what the cluster committed, so there is nothing to restore), but it is an implicit asymmetry - a one-line note in the branch comment ("modified records are intentionally NOT reloaded; their in-memory content is what the cluster committed") would match the care taken elsewhere in this method.

Style / conventions

  • Matches CLAUDE.md: final on new params/locals, imports (no FQNs), no em dashes, @Tag("slow") on the 3-node IT consistent with siblings, release note added. Both poisoned pages are evicted from the process-global PageManager in the engine test's finally - good hygiene.

Overall this is a high-quality change with strong engine-level test coverage and clear reasoning. The primary ask is closing the two coverage gaps above (especially the HTTP 409 mapping, since that is the client-facing contract the whole PR is about).

Note: the Maven suite could not be run in this sandboxed environment, so the correctness assessment is from static analysis; the author reports the engine and HA batteries pass.

…recovery path

Closes both coverage gaps from the round-4 review:

- HTTP 409 mapping: new server-module unit test drives the real handleRequest
  catch chain (mocked exchange, per the reviewer's sanctioned fallback) asserting
  409 + do-not-retry detail, plus catch-order guards (plain TransactionException
  stays 500, NeedRetryException stays 503). A new wire-level IT method in
  Issue5064CommittedRemotelyContractIT proves the exception reaches that catch
  RAW through the real Raft commit path over HTTP.

- applyLocallyAfterMajorityCommit: package-private (no new production hook; the
  failure is injected via the payload's transaction) with a unit test asserting
  the flag is set BEFORE the apply, the failure stays silent to callers, no
  rollback is added, and the reconcile + step-down remedy fires. Composes with
  the engine WalCommitOrderingTest that pins the flag's finally semantics; a
  real ALL-quorum IT would hinge on Ratis watch timeouts (nondeterministic).

Also found while closing the gaps: a follower forwarding a write lost the
do-not-retry contract (reconstructLeaderException collapsed the 409 body to a
generic TransactionException). Added the (String) constructor + factories entry
+ reconstruction unit test.

Minor notes: took the suggested branch comment (modified records intentionally
not reloaded); acknowledged the verbose tx-toString note (kept for diagnostics).
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Code Review: PR #5075 - distinct contract for post-quorum local commit failure

Reviewed against CLAUDE.md. Overall this is a well-crafted, high-quality fix. The problem (a durably-committed transaction reported as a generic failure, with identity reset turning app-level retries into duplicate inserts) is real and the fix is surgical: a single new durability regime plus a distinct, catchable exception type. Layered test coverage (engine unit, HA unit, HA IT, server unit, follower-forwarding reconstruct) is exemplary.

Correctness - verified points

  • Catch ordering is correct. TransactionCommittedRemotelyException precedes the generic TransactionException arm in AbstractServerHttpHandler, so the specific 409 mapping is reached before the generic 500. The plainTransactionExceptionStillMapsTo500 test pins this order against future refactors - nice defensive test.
  • The exception is deliberately NOT a NeedRetryException (it extends TransactionException directly), which is exactly right: retryability is the whole hazard. needRetryExceptionStillMapsTo503 guards the split.
  • Flag lifecycle is clean. remotelyCommitted is cleared in begin(), reset(), and transitively via reset() in resetAndFireCallbacks() on the success path. The reused-base-context sticky-flag hazard is explicitly regression-tested in WalCommitOrderingTest.
  • Regime precedence is sound. walAppended is checked before remotelyCommitted, so a post-append failure still fences (intentional: an orphaned local WAL record exists regardless of remote commit), and that branch also uses reset() rather than rollback(), so identities are preserved there too.
  • Follower-forwarding path preserved. The new LEADER_EXCEPTION_FACTORIES entry keeps the do-not-retry contract across a 409 wire response instead of collapsing to a generic TransactionException.

Minor points / questions

  1. reconcile after a successful commit2ndPhase. In both the main phase-2 path and the ALL-quorum path, saveConfiguration() runs after commit2ndPhase() succeeds. If only saveConfiguration() throws, the catch calls reconcileLeaderPagesAfterPhase2Failure, which replays the WAL against pages that were already published by the successful commit2ndPhase. This edge is pre-existing (not introduced here) and WAL replay from the same payload should be idempotent, but since this PR hardens exactly this failure surface, a one-line note confirming reconcile is safe when phase-2 already succeeded would help a future reader who might assume the catch only fires on a pre-publish failure.

  2. setRemotelyCommitted is public. Necessary for the cross-module (ha-raft) call, and the Javadoc clearly scopes it HA-layer only, after quorum commit, before phase 2. No action needed - just noting it is an intentional widening of the API surface that relies on convention rather than enforcement.

  3. Comment density. The remotelyCommitted branch carries a ~12-line paragraph, matching the surrounding walAppended/fence branches - consistent with the file style, just confirming it fits.

Style / conventions

  • Uses final on new params/fields, imports the class rather than FQN, existing wildcard exception import covers the new type - matches CLAUDE.md.
  • Slow tagging respected: Issue5064CommittedRemotelyContractIT carries @tag("slow").
  • Test cleanup in WalCommitOrderingTest correctly evicts both poisoned pages from the process-global PageManager.INSTANCE in finally - good attention to shared static state.
  • No System.out, no new dependencies, no author attribution added.

Performance and security

  • Cost is a single boolean field plus one branch in the commit finally - no measurable overhead, no added GC pressure.
  • Exception messages contain only the transaction toString() and a reconcile-outcome phrase; no sensitive data leaked. 409-instead-of-5xx actively reduces a data-integrity hazard (duplicate inserts from client/load-balancer retries).

Verdict

Solid and effectively ready to merge; nothing here is a blocker. The fix is tightly scoped, well-reasoned, and thoroughly tested. Point 1 is the only thing worth a glance before merge.

Automated review - please treat as advisory.

…dy succeeded

Round 5 was 'effectively ready to merge; nothing is a blocker' with one
point worth a glance, taken as the requested one-line note: the phase-2
catch also fires when commit2ndPhase succeeded and only the subsequent
saveConfiguration() threw - reconciling then replays the payload WAL
against already-published pages, which is safe by the #4926 replay
semantics (equal version re-applies the same absolute bytes,
idempotently; lower versions are skipped). Pre-existing edge, now
documented at the site.

Comment-only change; the full suite is running on main in parallel, so no
local Maven run was performed for this push (CI compiles the PR).
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: PR #5075 - distinct contract for post-quorum local commit failure

Reviewed the full diff, the commit2ndPhase finally-block regime ordering, the HTTP catch chain, the Raft commit path, and the exception hierarchy. This is a carefully-scoped, well-reasoned change. Below is constructive feedback, ordered by priority.

Strengths

  • Correct finally-block placement. The remotelyCommitted branch is inserted after walAppended and before the fence-refused and #4940 rollback branches. That ordering is the crux: a post-append failure still fences (an orphaned local WAL record exists regardless of remote commit), while a pre-append failure skips the identity rollback that would otherwise turn an app retry into duplicate inserts. The inline comment explaining why is excellent.
  • Sticky-flag hazard handled and tested. The reused base TransactionContext clears the flag in both begin() and reset() (belt-and-braces), and WalCommitOrderingTest deliberately re-drives a fresh transaction on the same context object to prove a plain local failure still rolls back (#4940). resetAndFireCallbacks() -> reset() also clears it on the success path, so no leak.
  • Exception hierarchy is right for the catch chain. TransactionCommittedRemotelyException extends TransactionException (not NeedRetryException), so it deliberately misses the 503 NeedRetryException arm and lands on the new 409 arm, which correctly precedes the generic TransactionException 500 arm.
  • Test coverage is thorough: engine unit (WalCommitOrderingTest), HA unit (Issue5064ApplyLocallyAfterMajorityCommitTest with InOrder verification that the flag is set before the apply), HA cluster IT (embedded + over-HTTP, verifying convergence on 3 nodes after step-down), server HTTP mapping unit test, and the follower-forwarding reconstruction test. The red-first IT approach is the right call.

Suggestions

1. (low/medium - defense in depth) The two wrapper catch arms lack the committed-remotely unwrap branch. AbstractServerHttpHandler has established precedent that DuplicatedKeyException can arrive wrapped - so it grows an unwrap branch in both the CommandExecutionException arm (line ~338) and the generic TransactionException arm (line ~377), each citing "some code paths wrap it." TransactionCommittedRemotelyException is also a TransactionException, so if any path ever wraps it in a plain TransactionException (the auto-commit wrapper in DatabaseAbstractHandler wraps "any Exception thrown by execute()") or a CommandExecutionException, it would fall through to the else -> 500, i.e. a retryable status - the exact hazard this PR exists to prevent. The IT confirms the current Raft path throws it raw (commit happens outside the wrapped lambda), so this is not a live bug - but for symmetry with DuplicatedKeyException and to make the do-not-retry contract robust against future call paths (scripts, batch, nested commands), consider adding an instanceof TransactionCommittedRemotelyException branch to both wrapper arms mapping to 409.

2. (minor - message accuracy) "the local apply failed" can be slightly inaccurate. As the new NOTE (#5075 review) comment acknowledges, the catch in RaftReplicatedDatabase.commit() also fires when commit2ndPhase succeeded and only the subsequent saveConfiguration() threw. In that case the local record apply actually succeeded; only the schema-config save failed. The outcome contract (committed cluster-wide, don't retry) is still correct, so this is purely cosmetic, but the phrase "the local apply failed" in the user-facing message overstates it for that sub-case. Not worth reworking the message, just noting it.

3. (nit) Log level for an expected, handled condition. The new 409 arm logs at getUserSevereErrorLogLevel() (SEVERE). This matches the DuplicatedKeyException arm, so it is consistent, but a post-quorum local failure is a genuinely rare and recoverable event; SEVERE is fine given it is not on any hot path.

Verdict

Solid, well-tested change that closes a real correctness gap (duplicate inserts on retry after a post-quorum local failure). Only suggestion #1 is worth acting on before merge, and even that is defense-in-depth rather than a live defect given the IT proves raw propagation on the current path.

Automated review; please validate against your own judgment.

@codecov

codecov Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 70.27027% with 11 lines in your changes missing coverage. Please review.
✅ Project coverage is 65.24%. Comparing base (fba3902) to head (155358e).
⚠️ Report is 17 commits behind head on main.

Files with missing lines Patch % Lines
...rcadedb/server/ha/raft/RaftReplicatedDatabase.java 36.36% 7 Missing ⚠️
...ception/TransactionCommittedRemotelyException.java 50.00% 2 Missing ⚠️
...java/com/arcadedb/database/TransactionContext.java 85.71% 0 Missing and 1 partial ⚠️
...server/http/handler/AbstractServerHttpHandler.java 93.33% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##               main    #5075      +/-   ##
============================================
- Coverage     65.62%   65.24%   -0.38%     
- Complexity      842      846       +4     
============================================
  Files          1687     1691       +4     
  Lines        135322   135516     +194     
  Branches      28959    28989      +30     
============================================
- Hits          88799    88418     -381     
- Misses        34413    34974     +561     
- Partials      12110    12124      +14     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

…pping

Round 6's defense-in-depth point taken: TransactionCommittedRemotelyException
is a TransactionException, so when a future call path wraps it (script
execution and command planners wrap in CommandExecutionException; the
auto-commit wrapper in DatabaseAbstractHandler wraps any Exception in a
plain TransactionException) it would have fallen through to the generic
500 - a status HTTP clients and load balancers treat as retryable, which
is the exact duplicate-insert hazard the type exists to prevent. Both
wrapper catch arms in AbstractServerHttpHandler now unwrap it to the same
non-retryable 409, mirroring the DuplicatedKeyException precedent (#4350).

Not a live bug: the IT proves the current Raft path throws it raw (the
commit happens outside the wrapped lambda) - this pins the contract
against future wrapped paths, exactly as the review framed it.

Red-first: two new wrapped-variant cases in
Issue5064CommittedRemotelyHttpStatusTest fail without the unwrap branches
(both degrade to 500) and pass with them; the message-accuracy and
log-level notes are acknowledged as-is (outcome contract correct in the
saveConfiguration sub-case; SEVERE consistent with the sibling arm).

Issue5064CommittedRemotelyHttpStatusTest 5/5, ErrorResponseNestedExceptionTest green.
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: fix(#5064) - distinct contract for post-quorum local commit failure

Reviewed the full diff and traced the interacting code paths (TransactionContext.commit2ndPhase finally-block regimes, the Raft phase-2 path, the ALL-quorum recovery path, the HTTP catch chain, and reconstructLeaderException). This is a carefully-built change with genuinely strong test coverage - the multi-round iteration history shows the important hazards were caught and closed. Overall: looks correct and ready to merge; comments below are mostly confirmations plus a few minor points.

What is done well

  • The core semantic is right. The remotelyCommitted flag is only set after quorum is confirmed (line 437 is reachable only past the replication try/catch at 408-414 that rolls back on replication failure, and past the if (!leader) guard). So the branch never skips the [tx] Phase-2 commit failure uses reset() instead of rollback(), leaving dangling RIDs on user documents #4940 rollback for a transaction that was not actually committed cluster-wide. The RIDs preserved are exactly the ones the replicated WAL carried to the followers, so identities genuinely match across the cluster.
  • Sticky-flag hazard handled. Clearing the field in both begin() and reset() on the reused per-thread context is the correct fix, and the second half of remotelyCommittedRegimePreservesIdentitiesOnPreWalFailure is a real red-first guard for it (verifies a fresh tx on the same context still rolls back per [tx] Phase-2 commit failure uses reset() instead of rollback(), leaving dangling RIDs on user documents #4940).
  • Defense-in-depth on the wire mapping is the standout. Mapping to 409 (not a retryable 5xx), the wrapped-exception unwrap arms in both CommandExecutionException and TransactionException catch blocks, and the reconstructLeaderException factory entry together close the duplicate-insert-over-the-wire hazard end to end. Catch ordering is correct: the specific TransactionCommittedRemotelyException arm precedes the generic TransactionException arm (line 291 before 364), and since it deliberately does not extend NeedRetryException, the 503/retry arm never captures it.
  • Test coverage spans every layer (engine finally-branch, HA embedded IT, HA wire IT, ALL-quorum unit test with InOrder verification, HTTP status unit test, follower-reconstruction unit test) and correctly tags the 3-node IT @Tag("slow").

Minor points / questions (non-blocking)

  1. Post-append + remotelyCommitted fences, then reconcile runs against a fenced DB. When the local failure happens after the WAL append, the walAppended branch fences the database (intended, per the comment). The Raft catch then calls reconcileLeaderPagesAfterPhase2Failure -> applyChanges on that now-fenced database. If the fence rejects applyChanges, reconciled comes back false and the message says "reconciliation ALSO failed" even though the pages are handled by the step-down/rejoin path anyway. Both outcomes step down, so it is not a correctness bug, but the message could be slightly misleading in that window. Worth a one-line confirmation that applyChanges behaves sanely on a fenced DB (or a note that this path intentionally relies on rejoin).

  2. User message says "reload the records" while the in-code comment (correctly) notes modified records are intentionally not reloaded because their in-memory content already equals what the cluster committed. That is fine as advice, but the exception Javadoc and the message are the user-facing contract - it may be worth a word that reloading is optional/for-freshness rather than required for correctness.

  3. Message builds payload.tx().toString() twice (line 462 plus inside recoverLeadershipAfterPhase2Failure(payload.tx().toString()) at 455). Already acknowledged as kept-for-diagnostics; only flagging as a micro-nit given the performance mantra - the failure path is exceptional, so it is genuinely negligible.

  4. saveConfiguration()-only failure surfaces as TransactionCommittedRemotelyException. Confirmed correct: when commit2ndPhase succeeded and only the config save threw, the tx data is committed, so the do-not-retry contract is accurate; the added [storage] Torn 64KB page write + version-based replay skip loses committed deltas even with full WAL fsync #4926-idempotency note on the reconcile replay is a good catch.

Nice work - the red-first discipline and the wire-level defense-in-depth are exactly right for a data-duplication hazard.

Automated review; verified statically against the surrounding code paths rather than by running the suite.

…ing precise

All four round-7 points were non-blocking; the two wording items taken:

1. Verified and documented at the reconcile site: applyChanges operates at
   the FileManager/PageManager level and never passes through
   checkDatabaseIsOpen (the fence's only choke point besides the
   pre-append guard), so a post-append fence does NOT reject
   reconciliation - it is the same page-level machinery recovery replay
   uses on reopen. The 'reconciliation ALSO failed' message therefore
   only appears on genuine replay failures, not fence rejections.

2. The exception javadoc now states reloading held records is optional
   (for freshness): their in-memory content already matches what the
   cluster committed, aligning the user-facing contract with the in-code
   comment.

Points 3 (toString twice on the exceptional path) and 4 (the
saveConfiguration sub-case) were the reviewer's own confirmations.

Comment/javadoc-only change.
@claude

claude Bot commented Jul 7, 2026

Copy link
Copy Markdown

Review: PR #5075 - distinct contract for post-quorum local commit failure

Reviewed the full diff. This is a careful, well-scoped fix and the seven review rounds show. The core idea - a remotelyCommitted durability regime that slots between the walAppended fence branch and the #4940 rollback branch - is the right shape for the problem, and the failure taxonomy in the finally block reads cleanly.

Strengths

  • Correct branch ordering. In TransactionContext.commit2ndPhase finally, remotelyCommitted sits after walAppended (so a post-append failure still fences, intentionally) and before the fence-refused/#4940 branches. Both the new branch and the fence-refused branch converge on reset(), which is safe on a fenced DB, so the interaction with a pre-existing fence is benign.
  • Flag lifecycle is airtight. Cleared in begin() and reset() (belt-and-braces), and the success path clears it via resetAndFireCallbacks() -> reset(). The sticky-flag hazard on the reused base context is explicitly regression-tested (remotelyCommittedRegimePreservesIdentitiesOnPreWalFailure runs a second plain-local failure on the same context and asserts the #4940 rollback still fires).
  • HTTP catch ordering is correct. TransactionCommittedRemotelyException extends TransactionException, and the specific unwrapped arm precedes the generic TransactionException arm; the wrapped arms in both CommandExecutionException and TransactionException are symmetric. plainTransactionExceptionStillMapsTo500 pins the ordering against a future refactor.
  • Wire round-trip preserved. The LEADER_EXCEPTION_FACTORIES entry means a follower forwarding a write reconstructs the exact type, keeping the do-not-retry signal; reconstructLeaderExceptionCommittedRemotelyKeepsDoNotRetryContract asserts it is not a NeedRetryException.
  • 409 vs 5xx choice is well-justified and consistent with the #4350 DuplicatedKeyException precedent - a retryable status here is exactly what would re-drive the duplicate insert.
  • Test coverage is genuinely layered: engine-level branch (WalCommitOrderingTest), ALL-quorum recovery unit test (Mockito InOrder proving the flag is set before commit2ndPhase), HTTP mapping unit test (unwrapped + two wrapping levels + negative cases), and a 3-node @Tag("slow") IT proving the raw exception reaches the catch through the real Raft path and that all nodes converge on the committed data.

Minor observations (non-blocking)

  1. Two-level exception nesting gap (HTTP handler). The wrapped arms unwrap exactly one level (e.getCause()). A CommandExecutionException -> TransactionException -> TransactionCommittedRemotelyException chain would fall through to the generic 500 and lose the 409. This is the same limitation the adjacent DuplicatedKeyException handling has, so it is consistent rather than a regression - but if a future path ever double-wraps, the contract silently degrades. A small unwrap-to-root-cause helper shared by the retryable/non-retryable types would close it for both. Not required here.
  2. setRemotelyCommitted is public engine API. Documented as HA-layer-only and guarded by tests, but nothing structurally prevents non-HA callers from flipping a transaction into the no-rollback regime. Acceptable given the module boundary (ha-raft is the only caller); worth keeping in mind if the engine ever exposes this more broadly.
  3. WalCommitOrderingTest poisons the process-global PageManager.INSTANCE read cache. The finally evicts both victim pages, the right mitigation and called out in the comments. Flagging only that the safety depends on the victim page ids being assigned before the throwing commit2ndPhase call - which the current ordering does, so it is safe as written.
  4. Idempotent reconcile on the saveConfiguration()-only failure. Good catch in the comment that the catch also fires when commit2ndPhase succeeded and only saveConfiguration() threw, and that replaying the payload WAL is safe by #4926 semantics. Consider a one-line note in the IT that this sub-case exists, since the current IT injects the fault before commit2ndPhase and so does not exercise the succeeded-then-saveConfiguration-failed path.

Style/conventions
Adheres to repo conventions: final on locals/params, no fully-qualified names, AssertJ assertThat(...), @Tag("slow") on the IT, Apache headers on new files, no System.out, no Claude authorship, release note added. Compliant with CLAUDE.md.

Overall: LGTM. The observations above are refinements, not blockers.

@lvca

lvca commented Jul 7, 2026

Copy link
Copy Markdown
Member Author

Round 8 dispositions (no push - all four observations are non-blocking refinements and the branch is converged):

  1. Two-level nesting gap: acknowledged as-is, per the review's own framing - the one-level unwrap is deliberately consistent with the adjacent DuplicatedKeyException handling. A shared unwrap-to-root-cause helper for retryable/non-retryable types is a reasonable future refactor but belongs to its own change touching both types, not this PR.
  2. setRemotelyCommitted is public engine API: accepted trade-off, documented at the field and setter. The engine has no visibility mechanism narrower than public that reaches ha-raft (different module); the sharp javadoc plus the regression tests are the guard.
  3. Poisoned-page eviction ordering: confirmed - the ids are captured before the throwing call, as the review verified.
  4. saveConfiguration sub-case in the IT: the sub-case is documented at the production catch site (the [storage] Torn 64KB page write + version-based replay skip loses committed deltas even with full WAL fsync #4926-idempotency note), which is where a maintainer touching that code looks first; the IT's fault-injection point is pre-commit2ndPhase by design.

Eight rounds, the last four all endorsements with polish. Ready to merge.

@codacy-production

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

🟢 Metrics 0 complexity

Metric Results
Complexity 0

View in Codacy

🟢 Coverage 75.68% diff coverage · -7.53% coverage variation

Metric Results
Coverage variation ✅ -7.53% coverage variation
Diff coverage ✅ 75.68% diff coverage

View coverage diff in Codacy

Coverage variation details
Coverable lines Covered lines Coverage
Common ancestor commit (fba3902) 135322 100946 74.60%
Head commit (155358e) 167319 (+31997) 112222 (+11276) 67.07% (-7.53%)

Coverage variation is the difference between the coverage for the head and common ancestor commits of the pull request branch: <coverage of head commit> - <coverage of common ancestor commit>

Diff coverage details
Coverable lines Covered lines Diff coverage
Pull request (#5075) 37 28 75.68%

Diff coverage is the percentage of lines that are covered by tests out of the coverable lines that the pull request added or modified: <covered lines added or modified>/<coverable lines added or modified> * 100%

NEW Get contextual insights on your PRs based on Codacy's metrics, along with PR and Jira context, without leaving GitHub. Enable AI reviewer
TIP This summary will be updated as you push new changes.

@lvca
lvca merged commit b6f09c2 into main Jul 7, 2026
22 of 31 checks passed
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…y covered; comments tightened

All four points taken; the two mediums were exact:

1. remotelyCommitted was a sticky flag on the REUSED per-thread base
   context - set by the Raft path, cleared by nothing. Any later path
   reaching the finally with committed=false && walAppended=false and the
   stale flag would silently skip the #4940 rollback on a transaction
   never remotely committed. Cleared at begin() and in reset() (belt and
   braces), with the hazard documented at both.

2. The engine-side regime branch had ZERO test coverage - the IT's fault
   fires before commit2ndPhase runs, so the identity assertion passed
   trivially. New engine test drives the branch for real: phase 1, the
   HA-layer flag set, a conflicting-version injection failing
   validateAndBumpVersions PRE-WAL - asserting the identity survives, no
   fence, and (second half, on the SAME reused context) that a fresh plain
   transaction still rolls back identities per #4940. Red-first: without
   the begin()/reset() clearing the second half fails on the sticky flag.

3. The 'without fencing' comment now states its true scope: only the
   pre-append branch skips the fence; a post-append failure intentionally
   still takes the walAppended fence branch (an orphaned local WAL record
   exists regardless of the remote commit) and preserves identities there
   too via reset().

4. applyLocallyAfterMajorityCommit documents why it does NOT surface the
   new exception (background ALL-quorum recovery, no user caller); the
   message ternary is extracted to a local.

WalCommitOrderingTest 6/6; engine compile + ha-raft compile clean.

(cherry picked from commit 0770941)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
The critical was mine: the round-1 automated edit ate the indentation of
the statements following both insertion points (begin() and reset()) and
over-indented the comment block. Both sites restored to clean 4-space
style; a column-0 sweep across both touched files confirms no other
artifact slipped in. The double blank line before the new test method is
collapsed.

The design note (no-fence applies only pre-append; the IT exercises the
pre-append path while post-append cases take the walAppended fence branch
with identities equally preserved) matches the intent documented in round
1 - acknowledged, no change.

WalCommitOrderingTest 6/6, Issue4940Phase2FailureRollbackTest green.

(cherry picked from commit b24d4ae)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…e contract made explicit

1. The 3-node IT carries @tag(slow) at class level, consistent with its
   siblings (Issue4740Phase2ReconcileIT, RaftDivergedFollowerRecoveryIT).

2. The engine test's finally now evicts BOTH poisoned pages from the
   process-global PageManager (the second conflicting page leaked).

3. Wire propagation - the review's 'please confirm' investigated and
   answered with code, because the answer was 'partially': the new
   exception does NOT extend NeedRetryException, so it was never mapped to
   the retryable 503 - but with no explicit mapping it fell into the
   generic 5xx branch, and HTTP clients and load balancers routinely retry
   5xx: the exact duplicate-insert hazard over the wire. The HTTP handler
   now maps TransactionCommittedRemotelyException to 409 with a
   do-not-retry detail and the exception class in the JSON payload - the
   same non-retryable rationale as the DuplicatedKeyException 409 (#4350).

4. Post-append remotely-committed path: declined as a separate test with
   reasoning - past the local append the walAppended branch runs
   IDENTICALLY regardless of the remote flag (fence + identity-preserving
   reset), and that branch is already covered by the #5053 fence tests;
   the only remotely-committed-specific behavior is pre-append, which both
   the IT and the engine test exercise.

Nit acknowledged: the single (String, Throwable) constructor stays until
another call site needs more.

WalCommitOrderingTest 6/6, the tagged IT green, server module compiles.

(cherry picked from commit 9478122)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…recovery path

Closes both coverage gaps from the round-4 review:

- HTTP 409 mapping: new server-module unit test drives the real handleRequest
  catch chain (mocked exchange, per the reviewer's sanctioned fallback) asserting
  409 + do-not-retry detail, plus catch-order guards (plain TransactionException
  stays 500, NeedRetryException stays 503). A new wire-level IT method in
  Issue5064CommittedRemotelyContractIT proves the exception reaches that catch
  RAW through the real Raft commit path over HTTP.

- applyLocallyAfterMajorityCommit: package-private (no new production hook; the
  failure is injected via the payload's transaction) with a unit test asserting
  the flag is set BEFORE the apply, the failure stays silent to callers, no
  rollback is added, and the reconcile + step-down remedy fires. Composes with
  the engine WalCommitOrderingTest that pins the flag's finally semantics; a
  real ALL-quorum IT would hinge on Ratis watch timeouts (nondeterministic).

Also found while closing the gaps: a follower forwarding a write lost the
do-not-retry contract (reconstructLeaderException collapsed the 409 body to a
generic TransactionException). Added the (String) constructor + factories entry
+ reconstruction unit test.

Minor notes: took the suggested branch comment (modified records intentionally
not reloaded); acknowledged the verbose tx-toString note (kept for diagnostics).

(cherry picked from commit e3725ea)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…dy succeeded

Round 5 was 'effectively ready to merge; nothing is a blocker' with one
point worth a glance, taken as the requested one-line note: the phase-2
catch also fires when commit2ndPhase succeeded and only the subsequent
saveConfiguration() threw - reconciling then replays the payload WAL
against already-published pages, which is safe by the #4926 replay
semantics (equal version re-applies the same absolute bytes,
idempotently; lower versions are skipped). Pre-existing edge, now
documented at the site.

Comment-only change; the full suite is running on main in parallel, so no
local Maven run was performed for this push (CI compiles the PR).

(cherry picked from commit b98daea)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…pping

Round 6's defense-in-depth point taken: TransactionCommittedRemotelyException
is a TransactionException, so when a future call path wraps it (script
execution and command planners wrap in CommandExecutionException; the
auto-commit wrapper in DatabaseAbstractHandler wraps any Exception in a
plain TransactionException) it would have fallen through to the generic
500 - a status HTTP clients and load balancers treat as retryable, which
is the exact duplicate-insert hazard the type exists to prevent. Both
wrapper catch arms in AbstractServerHttpHandler now unwrap it to the same
non-retryable 409, mirroring the DuplicatedKeyException precedent (#4350).

Not a live bug: the IT proves the current Raft path throws it raw (the
commit happens outside the wrapped lambda) - this pins the contract
against future wrapped paths, exactly as the review framed it.

Red-first: two new wrapped-variant cases in
Issue5064CommittedRemotelyHttpStatusTest fail without the unwrap branches
(both degrade to 500) and pass with them; the message-accuracy and
log-level notes are acknowledged as-is (outcome contract correct in the
saveConfiguration sub-case; SEVERE consistent with the sibling arm).

Issue5064CommittedRemotelyHttpStatusTest 5/5, ErrorResponseNestedExceptionTest green.

(cherry picked from commit 9b568a8)
robfrank pushed a commit that referenced this pull request Aug 14, 2026
…ing precise

All four round-7 points were non-blocking; the two wording items taken:

1. Verified and documented at the reconcile site: applyChanges operates at
   the FileManager/PageManager level and never passes through
   checkDatabaseIsOpen (the fence's only choke point besides the
   pre-append guard), so a post-append fence does NOT reject
   reconciliation - it is the same page-level machinery recovery replay
   uses on reopen. The 'reconciliation ALSO failed' message therefore
   only appears on genuine replay failures, not fence rejections.

2. The exception javadoc now states reloading held records is optional
   (for freshness): their in-memory content already matches what the
   cluster committed, aligning the user-facing contract with the in-code
   comment.

Points 3 (toString twice on the exceptional path) and 4 (the
saveConfiguration sub-case) were the reviewer's own confirmations.

Comment/javadoc-only change.

(cherry picked from commit 155358e)
@robfrank
robfrank deleted the fix/5064-committed-remotely branch August 26, 2026 09:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant