Skip to content

Fail deferrable KPO instead of silent SUCCESS when a GC'd pod takes its XCom with it - #73749

Open
developer-rpai wants to merge 1 commit into
apache:mainfrom
developer-rpai:fix-73117-kpo-gcd-pod-xcom
Open

developer-rpai wants to merge 1 commit into
apache:mainfrom
developer-rpai:fix-73117-kpo-gcd-pod-xcom

Conversation

@developer-rpai

Copy link
Copy Markdown
Contributor

Problem

A deferrable KubernetesPodOperator with do_xcom_push=True can be marked SUCCESS without pushing any return_value XCom when the pod is garbage-collected between the trigger firing and the worker resuming the task. Downstream tasks then fail unrecoverably at templating time (TypeError: the JSON object must be str, bytes or bytearray, not NoneType); retrying the downstream task can never recover because the upstream XCom does not exist.

Root cause

In trigger_reentry() (providers/cncf/kubernetes/operators/pod.py), the 404 handler added by #66716 returns silently whenever the trigger event status is "success" — regardless of do_xcom_push. The trigger observed the main container finish, but the XCom sidecar died with the GC'd pod, so the returned None becomes a successful task with no XCom.

Fix

Keep the silent-success path only when do_xcom_push is off (there, only pod logs are lost). When XCom was expected, raise PodNotFoundException instead, so the task fails (and retries, when configured) and recreates the pod along with its XCom sidecar — mirroring the existing behavior for non-success events.

Note: a prior fix attempt (#73131) was closed unmerged during today's enforcement of the new 5-open-PR limit for non-committers, before any maintainer review. This PR re-implements the same approach (fail when do_xcom_push is set) with its own regression test.

Tests

  • Added test_async_trigger_reentry_raises_pod_not_found_on_success_when_xcom_push: GC'd pod + success event + do_xcom_push=True -> PodNotFoundException. Fails on main before this change.
  • Updated test_async_trigger_reentry_returns_when_pod_gcd_on_success to pin do_xcom_push=False, making the preserved silent-success contract explicit.
  • A source-level reproduction script confirmed the unguarded if event["status"] == "success": return pattern before the fix and the do_xcom_push guard after.
  • Sandbox limits: the full provider unit-test suite could not be run in this environment (Airflow + provider test dependencies are not installable here); relying on CI for the suite. py_compile passes on both touched files.

Impact

Correctness: prevents silent data loss (missing XCom) and a wrong task status (SUCCESS for a task whose result was destroyed). No behavior change when do_xcom_push is False.

Fixes: #73117


Was generative AI tooling used to co-author this PR?
  • Yes — Muse (Meta)

Generated-by: Muse (Meta) following the guidelines

A deferrable KubernetesPodOperator with do_xcom_push=True was silently
marked SUCCESS when its pod was garbage-collected between the trigger
firing and task re-entry: the 404 handler in trigger_reentry() returned
None for any "success" event, so no return_value XCom was ever pushed and
downstream tasks failed unrecoverably at templating time.

Keep the existing silent-success behavior when do_xcom_push is off (only
pod logs are lost there). When XCom was expected, raise PodNotFoundException
so the task fails/retries and recreates the pod along with its XCom sidecar.

Fixes: apache#73117

Signed-off-by: Rakesh Ramakrishna Pai <developer-rpai@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:providers provider:cncf-kubernetes Kubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Deferrable KubernetesPodOperator marks task SUCCESS without pushing XCom when the pod is gone at re-entry (regression from #66716, provider 10.17.1+)

1 participant