Systems integration

API integrations: share the quota budget across every worker

October 6, 2026 · Systems integration

Coordinate API usage across background jobs and interactive workflows with scoped budgets, request-cost awareness and provider-specific backoff.

Interactive requests, scheduled synchronization and historical backfills enter one shared quota controller. It accounts for provider scope and request cost, then admits work to the API or delays it. Response quota signals feed back into the controller.
Original explanatory diagram created for Quarro, October 6, 2026 · Original vector artwork composed from geometric shapes and text; no stock images, logos, screenshots, or third-party visual assets.

Why this workflow needs an operating contract

When several workers share an API quota, each worker needs a view of the shared budget. Coordinate admission by the provider’s actual quota scope, account for request cost, and use response signals to adjust pacing. Preserve capacity for important work and make deferred jobs visible instead of hiding them behind repeated retries. Picture a reporting integration with a morning dashboard refresh, an interactive repair tool and a historical backfill. Each process works in isolation during testing. At launch, they run under the same account or project and compete for the same allowance. More workers can then produce more throttling without producing more completed work.

Identify what the provider is counting

A quota can apply per project, user, app installation, store, endpoint or combination of those scopes. Read the current documentation for the exact API and authentication method. Treat a request-rate limit, a concurrency limit and a paid usage allowance as separate constraints. The <a href="https://developers.google.com/workspace/sheets/api/limits">Google Sheets API documentation</a> describes separate per-minute project and per-user-per-project quotas for reads and writes. <a href="https://docs.github.com/en/rest/using-the-rest-api/rate-limits-for-the-rest-api">GitHub’s REST API documentation</a> describes shared user budgets and additional secondary limits. These examples show why a single “requests per second” setting may be incomplete. Request count is not always the correct unit. <a href="https://shopify.dev/docs/api/usage/limits">Shopify documents different limit mechanics by API</a>, including calculated query cost for its GraphQL Admin API. A small lookup and a larger query need not consume the same allowance. Use the documentation and returned usage metadata for the API actually in use.

Put a coordinator in front of shared work

Define a budget key that matches the quota boundary. Jobs sharing that boundary should acquire capacity through one coordinated mechanism, even if they run on different machines. Jobs with genuinely independent quotas should not block each other unnecessarily. For a simple low-volume service, a single worker may be enough. A distributed system may need an atomic shared counter, token bucket or other admission-control mechanism. Specify how the controller behaves after restarts and how it prevents several workers from spending the same remaining capacity at once. - Keep reads, writes and specialized endpoint limits separate when the provider does - Estimate request cost conservatively, then reconcile available provider feedback - Track in-flight concurrency independently from requests completed in a time window - Account for other applications using the same quota, not just your own process - Preserve a recoverable job record when work must wait Do not describe an internal counter as the provider’s exact remaining balance. External activity and distributed responses can make it approximate. Design for a limit error even when the estimate says capacity remains.

Make priority a business decision

Agree on which work should finish first. A customer-facing correction may deserve capacity ahead of a historical backfill, while a time-sensitive operational sync may outrank an optional dashboard refresh. Keep the categories small enough that operators can understand them. Set a maximum wait age or escalation rule for low-priority work so it cannot starve indefinitely. Make backfills resumable and allow them to pause at a checkpoint. Show operators what is waiting, why it is waiting, and whether the delay affects a business deadline. The priority system should not override the provider’s limits. Reserved capacity is an internal planning choice within the allowed budget. When sustained demand is too high, reduce unnecessary calls, change freshness expectations or reassess the integration design.

Handle throttling according to the API

A 429 response is a common throttling signal, but status codes and retry instructions differ by provider. GitHub can return 403 or 429 for rate-limit conditions and documents how to use Retry-After or reset information. Its response headers are the authoritative current signal; a separate status request can disagree. Preserve the response context rather than treating every 403 as throttling. Google Sheets recommends exponential backoff for quota errors. Use bounded retries and randomized delay where the provider recommends it, and ensure the delay is coordinated across workers. Otherwise every worker may wake at the same instant and trigger another burst. For a write whose outcome is uncertain, resolve whether it succeeded before replaying it when duplication would matter. A quota controller manages capacity; it does not establish that every operation is safe to repeat.

Measure completed work and waiting time

- Record requests, estimated or returned cost, throttles and concurrency by quota scope - Track queue age and oldest waiting job for each priority - Distinguish provider-required waiting from authentication errors and invalid requests - Alert on sustained freshness lag or a deadline risk, not every routine backoff - Test simultaneous workers, a large backfill, a restart and unexpected external usage The useful result is predictable progress under a constrained allowance. Start by mapping the busiest integration’s quota scopes and the workflows competing for them. That often exposes a coordination problem before an expensive capacity change is considered. For related reliability controls, see <a href="https://quarro.org/blog/queue-backed-workflows-need-retry-ownership/">retry ownership for queue-backed workflows</a> and <a href="https://quarro.org/blog/change-sync-needs-checkpoints-not-polling/">checkpointed change synchronization</a>. These decisions are part of maintainable <a href="https://quarro.org/">systems and API integration work</a>.

Sources