Clean up stale KubernetesPodOperator pods#70140
Open
fat-catTW wants to merge 2 commits into
Open
Conversation
fat-catTW
requested review from
hussein-awala,
jedcunningham and
jscheffl
as code owners
July 20, 2026 16:53
fat-catTW
force-pushed
the
fix-kpo-zombie-pod-cleanup
branch
from
July 21, 2026 03:59
27b62f0 to
b31e182
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why:
KubernetesPodOperator pods can remain in Kubernetes after their TaskInstance is no longer active when a sidecar container keeps the pod running.
This can leave stale pods consuming cluster resources even though Airflow has already moved on from the task attempt.
Related: #58968
Solution:
Add an optional KubernetesPodOperator zombie pod cleanup path for KubernetesExecutor.
When enabled, the executor periodically lists pods with the
kubernetes_pod_operator=Truelabel and compares them with unfinished TaskInstances in the Airflow DB.A pod is deleted only when Airflow can identify the TaskInstance it belongs to from the pod labels, and either:
A pod is not deleted when:
kpo_zombie_pod_cleanup_max_deletes_per_loopfor the current loopPods selected for deletion are processed oldest first by
metadata.creation_timestamp.The cleanup is disabled by default. Users can enable and tune it with the following
[kubernetes_executor]config options:kpo_zombie_pod_cleanup_enabledkpo_zombie_pod_cleanup_intervalkpo_zombie_pod_cleanup_max_deletes_per_loopkpo_zombie_pod_deletion_grace_period_secondsThis is scoped to KubernetesExecutor for now, with the cleanup logic kept in a reusable helper so it can be wired into other scheduler-side cleanup paths later.
Was generative AI tooling used to co-author this PR?
[X]Yes (please specify the tool below)Generated-by: [Codex] following the guidelines
{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.