Conversation
Synchronous SQL statements could be submitted twice when a worker crashed after submission because retries had no persisted external ID. Generated-by: GitHub Copilot CLI (GPT-5.6 Sol) Signed-off-by: 1fanwang <1fannnw@gmail.com>
Persist the Databricks statement ID before publishing its XCom so a worker failure can reconnect on retry. Generated-by: GitHub Copilot CLI (GPT-5.6 Sol) Signed-off-by: 1fanwang <1fannnw@gmail.com>
|
Hello @1fanwang - thank you for your contributions to Apache Airflow! The Airflow community has introduced a limit of 5 open pull requests at a time for contributors without write access to the repository. You currently have 34 open pull requests, so - as a one-time step of introducing the limit - we closed the ones where maintainers have not engaged yet:
These pull requests stay open because maintainers are already engaged in them - they count towards your limit:
This is not a judgement of you or of your changes. We never told contributors before that opening many pull requests at once was a problem, so there is nothing to feel bad about - and nothing is lost: your branches, commits and the review history stay where they are. What we ask you to do is to make your first prioritization decision: choose which of the pull requests above matter most to you, and reopen them (up to 5 open at a time, including the ones still open) with the "Reopen pull request" button or While your pull requests are waiting for review, the most valuable thing you can do is help in other ways - reviewing other contributors' pull requests, helping with issues, and taking part in the discussions on the devlist and Slack. Why we introduced the limit, what it means for you and how to reopen or restore a pull request is explained in https://github.com/apache/airflow/blob/main/contributing-docs/32_open_pull_request_limit.rst. Drafted-by: Claude Code (Opus 5); reviewed by @potiuk before posting |
When a worker dies after submitting a Databricks SQL statement, a synchronous task retry submits the statement again. The original statement can still be running, so users can get duplicate work and cost.
Before this change, the retry submits
statement-2whilestatement-1remains active. After this change, Airflow 3.3+ storesstatement-1in task state and reconnects to it. A retry also returns an already successful statement without resubmitting, and starts fresh only after failure, cancellation, closure, or a missing statement.This uses the AIP-103
ResumableJobMixincontract only whenwait_for_termination=Trueanddeferrable=False. The statement ID reaches task state before its XCom is published, so an XCom error cannot reopen the duplicate-submission window. Fire-and-forget and deferrable execution keep their existing paths.Testing
The driver below calls
execute()twice with the real task-state accessor and supervisor messages. Its local Databricks hook forces statement-ID XCom publication to fail after the first submission.Live task-state and XCom ordering proof
Save the driver as
dev/databricks_xcom_ordering_e2e.py, then run it against the previous PR head:Run the same driver against this branch:
Please check the type of change your PR introduces:
Relevant Issue(s)
None.
Description
See above.
How did you test it?
See the live proof above.
Did you add documentation?
Yes. The SQL statements guide covers the recovery behavior and version floor.
Does this introduce a breaking change?
No.
Checklist
Do you use Generative AI to contribute to this PR?
AI-assisted review checklist
PR Checklist