Skip to content

Introduced retries to the TableauSensor - #52770

Merged
potiuk merged 2 commits into
apache:mainfrom
dominikhei:tableau-sensor-retries
Apr 6, 2026
Merged

potiuk merged 2 commits into
apache:mainfrom
dominikhei:tableau-sensor-retries

Conversation

@dominikhei

Copy link
Copy Markdown
Contributor

closes: #32799
I opened this as a draft since I’m unsure about the pattern used and made some compromises —> would appreciate feedback.

@hussein-awala , I like your idea of retrying on certain error codes (e.g., 408, 5xx). However, requests to Tableau’s get_by_id endpoint don’t expose status codes, and many exceptions don’t either. Error messages have them in plain text like:

409093: Resource Conflict
Job for 'Main Transactions Dashboard' is already queued. Not queuing a duplicate.

Parsing strings for status codes feels fragile. We could catch exceptions that have response.status_code and retry on likely transient codes, but many don’t.
My suggestion: add an optional retries_on_failure parameter to the sensor. It retries a set number of times on any failure, then raises an AirflowException. This keeps logic simple. The only other sensor that allows for configuring retries by itself is the BashSensor.

Handling job cancellations on failure, as described in the issue, seems better suited for on_failure_callback.
Additionally the TableauOperator comes with a param blocking_refresh which could lead to the same problem, I wanted to address the sensor first and get some feedback.

Comment thread providers/tableau/src/airflow/providers/tableau/sensors/tableau.py Outdated
@dominikhei
dominikhei marked this pull request as ready for review July 8, 2025 21:30
@dominikhei

Copy link
Copy Markdown
Contributor Author

@eladkal what is your thought on the adjusted name of the retries parameter? (max_status_retries instead of retries_on_failure )

@eladkal

eladkal commented Jul 27, 2025

Copy link
Copy Markdown
Contributor

The sensor currently doesn't have deferrable capability. Do we expect this functionality to work well with defer mode when implemented?

@dominikhei

Copy link
Copy Markdown
Contributor Author

The sensor currently doesn't have deferrable capability. Do we expect this functionality to work well with defer mode when implemented?

In that case, if deferrable == True the retries could be passed to the trigger, or? E.g looking at the BatchSensor example:

            self.defer(
                timeout=timeout,
                trigger=BatchJobTrigger(
                    job_id=self.job_id,
                    aws_conn_id=self.aws_conn_id,
                    region_name=self.region_name,
                    waiter_delay=int(self.poke_interval),
                    waiter_max_attempts=self.max_retries,
                ),

In this case the trigger probably should retry on any error though, to not make it overly complex, as other Tableau Operators / Sensors will also pass in retries.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 5 days if no further activity occurs. Thank you for your contributions.

@github-actions github-actions Bot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Sep 12, 2025
@github-actions github-actions Bot closed this Sep 18, 2025
@dominikhei

Copy link
Copy Markdown
Contributor Author

The sensor currently doesn't have deferrable capability. Do we expect this functionality to work well with defer mode when implemented?

In that case, if deferrable == True the retries could be passed to the trigger, or? E.g looking at the BatchSensor example:

            self.defer(
                timeout=timeout,
                trigger=BatchJobTrigger(
                    job_id=self.job_id,
                    aws_conn_id=self.aws_conn_id,
                    region_name=self.region_name,
                    waiter_delay=int(self.poke_interval),
                    waiter_max_attempts=self.max_retries,
                ),

In this case the trigger probably should retry on any error though, to not make it overly complex, as other Tableau Operators / Sensors will also pass in retries.

@eladkal what is your take? I would like to finish this PR.

@potiuk potiuk reopened this Feb 15, 2026
@potiuk

potiuk commented Feb 15, 2026

Copy link
Copy Markdown
Member

I will reopen it, as it seems close to be complete.

@potiuk
potiuk force-pushed the tableau-sensor-retries branch from b0f814a to 525f7ac Compare February 15, 2026 20:38
@potiuk

potiuk commented Feb 15, 2026

Copy link
Copy Markdown
Member

@eladkal - any comments?

@eladkal

eladkal commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

@eladkal - any comments?

My comments were more of questions. I don't have time to look into it but regardless this is something we can always change in the future so feel free to proceed

@github-actions github-actions Bot removed the stale Stale PRs per the .github/workflows/stale.yml policy file label Mar 13, 2026
@dominikhei

Copy link
Copy Markdown
Contributor Author

@eladkal - any comments?

My comments were more of questions. I don't have time to look into it but regardless this is something we can always change int he future so feel free to proceed

Thanks a lot for the reply, will test my code again and proceed

@potiuk

potiuk commented Apr 6, 2026 •

Copy link
Copy Markdown
Member

@dominikhei This PR has a few issues that need to be addressed before it can be reviewed — please see our Pull Request quality criteria.

Issues found:

  • ⚠️ Unresolved review comments: This PR has 1 unresolved review thread from maintainers: @eladkal (MEMBER): 1 unresolved thread. Please review and resolve all inline review comments before requesting another review. You can resolve a conversation by clicking 'Resolve conversation' on each thread after addressing the feedback. See pull request guidelines.

Note: Your branch is 1381 commits behind main. Some check failures may be caused by changes in the base branch rather than by your PR. Please rebase your branch and push again to get up-to-date CI results.

What to do next:

  • The comment informs you what you need to do.
  • Fix each issue, then mark the PR as "Ready for review" in the GitHub UI - but only after making sure that all the issues are fixed.
  • There is no rush — take your time and work at your own pace. We appreciate your contribution and are happy to wait for updates.
  • Maintainers will then proceed with a normal review.

There is no rush — take your time and work at your own pace. We appreciate your contribution and are happy to wait for updates. If you have questions, feel free to ask on the Airflow Slack.


Note: This comment was drafted by an AI-assisted triage tool and may contain mistakes. Once you have addressed the points above, an Apache Airflow maintainer — a real person — will take the next look at your PR. We use this two-stage triage process so that our maintainers' limited time is spent where it matters most: the conversation with you.

@potiuk
potiuk merged commit cfe840f into apache:main Apr 6, 2026
86 checks passed
shivaam pushed a commit to shivaam/airflow that referenced this pull request Apr 8, 2026
* Introduced retries to the TableauSensor to account for transient errors

* Renamed param retries_on_failure to max_status_retries to ensure distinction to airflow retries
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Tableau - problem when fetching job status

3 participants