Dmitriy Kononov.
Let’s talkContact

News and analysis

API rate limits: sharing a request budget across workflows

Share account-level API limits across workflows, handle HTTP 429, prioritise important tasks and monitor synchronisation delays.

DevelopmentPublished:

An overnight export finishes without visible errors, but staff cannot update customer records the next morning. A shared external API limit may be the cause: batch processing, the user interface and administrative checks all consume the same resource. Each workflow looks restrained in isolation, yet their combined traffic exceeds the allowed request rate.

Treat API limits as a shared team budget. Know both the request allowance and its scope: account, endpoint, key, plan or another boundary. That determines whether multiple workers need coordination and how available throughput should be allocated.

What the Twilio Content API update illustrates

Twilio's September 30, 2026 publication describes account-level Content API limits and HTTP 429 responses when they are exceeded. Values differ by operation: Fetch Content is listed at 100 requests per second, while List Content is listed at 20. Update Content is also listed at 20 requests per second.

These figures apply to the specified Content API endpoints. They should not be extended to Voice, Verify, other Twilio products or other providers. Maintain a registry for your integration: operation, current limit, applicable scope and documentation link.

Different access keys do not necessarily provide independent budgets. If the provider counts at account level, several keys still participate in one allowance. This is particularly relevant when different teams maintain integrations and background scripts run independently of the customer application.

Request rate and concurrency solve different problems

Limiting concurrent workers protects execution capacity, but does not necessarily enforce request rate. One worker can send many short requests in succession; several slow workers may generate little traffic over the same interval. Measure the actual outbound flow at the provider boundary.

The n8n rate-limiting guide explicitly explains that queue mode and concurrency controls do not remove third-party API limits. It also separates outbound request pacing from inbound traffic protection: a public webhook needs a separate limiting layer when that control is required.

For workflow automation, first identify the resource that is scarce. The execution queue, a provider's API limit and a customer task's deadline may require different mechanisms. Adding workers can simply consume the available allowance faster.

Allocate capacity around operational needs

Imagine a shop updating message templates, displaying their status to an operator and running a full catalogue reconciliation at the same time. This is a design example, not a measurement from a particular implementation. The operations have different urgency: an interactive request should not wait indefinitely behind a large administrative export.

Separate traffic into classes: interactive actions, required synchronisation and batch processing. For each, define an acceptable delay, whether it can be deferred and the consequences of missing its deadline. Reserve capacity for important operations, then specify how batch work can use spare capacity.

With multiple processes, budget accounting must be coordinated. A local delay in every script may work with one instance but no longer reflect aggregate load after scaling. A shared dispatcher or coordinated counter store are possible implementation choices. The choice depends on architecture and failure-handling requirements.

Do not aim to run continuously at the published ceiling. Short bursts, administrative calls and changing response times introduce variation. Choose a reserve based on observations and the consequences of delay, rather than a universal percentage.

A person beside a server and lines marked with check marks.
K. Limpitsouni / unDraw · License

Handle HTTP 429 as a scheduled task

Retain the task and determine when it can be attempted again. If the response contains Retry-After, interpret it according to the provider's contract. n8n describes waiting, batching and increasing delays between attempts. These are tools for configuring behaviour, not a promise that every workflow automatically respects a provider's limits.

Without an explicit retry time, consider a bounded backoff policy with a small random spread so workers do not all restart together. Define an endpoint for the policy: a task deadline, an attempt limit or a handover to an operator. Repeating an outdated action indefinitely has little value, even if another attempt is technically possible.

Read and write requests have different consequences. Before repeating a write, determine whether the action may already have happened. Shared-budget management does not replace duplicate-result protection; that belongs to a separate part of the API integration contract.

Reduce traffic while preserving useful freshness

Before adding workers, check whether the API supports paginated retrieval, batch operations or local storage of slowly changing reference data. Caching needs an update rule: yesterday's staff list may be acceptable for a report but insufficient for an access decision.

The custom API description treats caching and request limits as architectural elements. It is useful context, not evidence of measured throughput. A new workflow still needs its own freshness requirements and source-change handling.

Monitor queue age, 429 counts by operation, the share of urgent tasks and recovery time after a batch run. A low error count alongside a synchronisation backlog lasting days does not establish acceptable integration performance. Monitoring should reflect the time needed to deliver a useful result.

Test combined load before scaling

Exercise the live interface, planned export and administrative checks together in an appropriate environment. Confirm that a 429 leaves a visible task, priority operations retain capacity and the accumulated queue eventually shrinks. Record who can change priorities during an incident.

Before increasing volume, also check the API contract and version: staying within the allowance does not make an invalid response schema correct. If requests go to AI models, complement request allocation with a model registry and cost controls. Rate, financial spend and result quality are connected, but need separate measures.

Sources