Skip to content

[Bug] Quota exceeded error during merge flush causes connector to silently stop delivering data #481

Description

@priyanshu-ctds

Issue:
When a BigQueryException with a quota exceeded error occurs during a merge flush, recordSuccessfulFlush is never called, leaving the failed batch permanently in the internal batches map. All subsequent flush attempts for the next batch then block indefinitely inside prepareToFlush waiting for the stuck batch to be removed, with no timeout, no log and no recovery. This causes the connector to appear running but stop delivering data silently, and persists even after the quota resets.

Fix:
Quota exceeded errors should be added as a retryable error in the merge flush retry logic alongside the existing serialize access and job internal errors. Since QueryUsagePerDay quota replenishes after 24 hours, configuring bigQueryRetry and bigQueryRetryWait such that the total retry window exceeds 24 hours will allow the flush to succeed naturally once the quota resets, without any task failure or data loss. This avoids the batch state getting broken and subsequent batches getting blocked indefinitely in prepareToFlush.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions