Dmitriy Kononov.
Let’s talkContact

News and analysis

Background jobs: retry processing without duplicating the outcome

Designing queues, retries and idempotency around business outcomes, visible delays and a clear process for jobs that cannot complete.

InfrastructurePublished:

A retry can recover a workflow after a brief outage. It can also send a document again or create a duplicate record. A background queue therefore needs to deliver the work and keep repeated processing within agreed boundaries. That design decision should precede the choice of queue library.

Start with the result a person expects

A background operation needs a clear contract. What does accepting a job mean? When is the result available? How does the user learn about a delay or a final failure? If the interface announces that a document has been sent immediately after placing a job in the queue, it may be promising an action that has not happened yet.

Distinguish states such as accepted, running, completed and requiring intervention. The appropriate number and names depend on the process. What matters is that the team can identify the current outcome and the action needed next. A critical process also needs a connection to the original request or document.

Understand why processing can happen again

Amazon SQS standard queues provide at-least-once delivery, so their consumers need to allow for repetition. This is one specific delivery contract; check the properties of the queue you select. Different guarantees do not remove every question about external effects or uncertain responses.

Imagine that an external service accepts an operation but its response is lost. The worker sees a timeout and may choose to retry. A missing response does not prove that the action failed to happen. Retrying without checking can duplicate the result. Establish whether the provider accepts an operation key or offers a way to retrieve a previously created outcome.

Idempotency means that repeating an action does not create an additional intended effect. Bind the key to the business operation rather than generating a fresh key on every worker attempt. A key alone does not prevent every duplicate: the implementation must record state, check that the input matches and account for concurrent processing.

Separate retryable failures from work needing a correction

Temporary unavailability and invalid input need different treatment. The former may justify another attempt within a defined retry budget. The latter usually requires a person or a separate correction workflow. Repeatedly executing an invalid job consumes resources and obscures the underlying problem.

Use a controlled delay between attempts. During a widespread failure, simultaneous retries can increase pressure on an already unavailable component. Choose the strategy around provider limits, completion deadlines and queue volume. A payment, a report and a reference-data update do not necessarily deserve the same number of attempts.

If a worker stops after an external action but before recording its local status, reconciliation may be necessary. Some processes can support a safe automatic retry. Others should pause with the uncertainty clearly visible. A blanket promise of exactly one execution conceals this boundary between local state and external effects.

Failed jobs need an operational owner

A dead-letter queue in SQS isolates messages that have not been processed successfully. Moving a message there does not resolve its cause. Define notification, diagnosis, ownership and a controlled route back into processing.

Before replaying a job, check the cause, earlier side effects and whether the work is still valid. A document may have changed, an order may have closed, or the external result may already exist. An operator view should provide context and previous attempt evidence while respecting access restrictions on customer information.

Measure more than the message count. The age of the oldest waiting operation, processing duration and the proportion of jobs ending in failure can reveal different problems. An overdue result in a particular business workflow matters more than a queue size judged against the same threshold under every load. Agree intervention criteria with the process owner.

In workflow automation, this design keeps exceptions visible between software and people. A useful first stage can cover one operation, its retry path and human review of an uncertain outcome. Expand automation to neighbouring actions after those paths have been exercised and their limits understood.

I connect this question with my Dent-Picks work, where I developed backend logic and synchronization between systems. An incoming IVR call automatically creates a GoHighLevel lead. That transition provides a clear acceptance outcome: a record the team can use for further work. The public case does not establish a queue implementation or retry guarantees; those need verification in the particular system.

My guide to connecting documents, CRM and tasks develops these requirements through request identifiers, retries after timeouts and visible failed transitions. It offers a basis for discussing background processing before choosing a queue library or its infrastructure.

Sources

Sources checked on 7 October 2026.