Issue:
When a BigQueryException with a quota exceeded error occurs during a merge flush, recordSuccessfulFlush is never called, leaving the failed batch permanently in the internal batches map. All subsequent flush attempts for the next batch then block indefinitely inside prepareToFlush waiting for the stuck batch to be removed, with no timeout, no log and no recovery. This causes the connector to appear running but stop delivering data silently, and persists even after the quota resets.
Fix:
Quota exceeded errors should be added as a retryable error in the merge flush retry logic alongside the existing serialize access and job internal errors. Since QueryUsagePerDay quota replenishes after 24 hours, configuring bigQueryRetry and bigQueryRetryWait such that the total retry window exceeds 24 hours will allow the flush to succeed naturally once the quota resets, without any task failure or data loss. This avoids the batch state getting broken and subsequent batches getting blocked indefinitely in prepareToFlush.
Issue:
When a
BigQueryExceptionwith a quota exceeded error occurs during a merge flush,recordSuccessfulFlushis never called, leaving the failed batch permanently in the internal batches map. All subsequent flush attempts for the next batch then block indefinitely insideprepareToFlushwaiting for the stuck batch to be removed, with no timeout, no log and no recovery. This causes the connector to appear running but stop delivering data silently, and persists even after the quota resets.Fix:
Quota exceeded errors should be added as a retryable error in the merge flush retry logic alongside the existing serialize access and job internal errors. Since
QueryUsagePerDayquota replenishes after 24 hours, configuringbigQueryRetryandbigQueryRetryWaitsuch that the total retry window exceeds 24 hours will allow the flush to succeed naturally once the quota resets, without any task failure or data loss. This avoids the batch state getting broken and subsequent batches getting blocked indefinitely inprepareToFlush.