Pattern: Long-Running Tasks
Some work takes too long to finish inside a web request: generating a report, transcoding a video, importing a million rows, sending a campaign to a hundred thousand people, running a model. If the request waits for the work, the connection times out, the server thread is held, the user stares at a spinner, and a retry starts the whole thing again. The standard answer is to accept the request quickly, do the work elsewhere, and let the client find out the result later. This chapter turns that sentence into a design.
1. Recognise the signal
Use this pattern when any of these hold:
- The work can take longer than a request timeout, which is typically tens of seconds.
- The duration is unpredictable or depends on input size.
- The work is CPU-heavy, memory-heavy or calls slow external services.
- The user does not need the result in the same response.
- The work must survive a server restart or deploy.
- You need to retry failures without involving the user.
If the work finishes in a few hundred milliseconds, just do it synchronously. The pattern adds moving parts, and you should not pay for them without a reason.
2. The basic shape
- The client calls the API to start the task.
- The API validates the request, creates a job record with a unique identifier and state "queued", enqueues the work, and returns immediately with the job identifier.
- A pool of workers takes jobs from the queue and executes them, updating the job record as they go.
- The client learns the outcome by polling the job's status, by a callback (webhook), or by a push message.
- The result is stored where the client can fetch it, usually in object storage for large outputs, with a link in the job record.
Return HTTP 202 Accepted with the job identifier and where to check status, not a 200 with a result that does not exist yet:
POST /v1/reports -> 202 { "job_id": "j-81", "status_url": "/v1/jobs/j-81" }
GET /v1/jobs/j-81 -> 200 { "state": "running", "progress": 0.42 }
GET /v1/jobs/j-81 -> 200 { "state": "succeeded", "result_url": "https://..." }
3. The job as a state machine
Give each job explicit states, and allow only valid transitions:
queued to running to succeeded, or to failed, with optional cancelled, and retrying between attempts.
Store with the job: who requested it, the input (or a reference to it), timestamps, the attempt count, the worker that holds it, progress, an error message and the result location. The job store is the source of truth about the task, which is why it belongs in a durable database and not in the queue.
A terminal state never changes. That makes status queries cheap to cache, and makes the rules for retry and cleanup simple.
4. Making it reliable
Queue semantics and acknowledgement
A worker takes a message and, if it does not finish and acknowledge it in time, the queue makes the message visible again for another worker. This is at-least-once delivery: a job can run more than once, for example if a worker crashes after finishing but before acknowledging.
So the work must be idempotent, or protected by an idempotency key:
- Produce results to a location derived from the job identifier, so a second run overwrites the first with the same content.
- Record completed steps in the job store, so a retry skips what is done.
- Use unique constraints or identifiers when writing side effects, such as an email send keyed by job and recipient.
Timeouts and heartbeats
Set a visibility timeout longer than the work should take, or have the worker heartbeat to extend its lease while it makes progress. If a worker dies, the lease expires and another worker picks the job up. Without a lease, a crashed worker leaves the job "running" forever.
Retries with backoff
Distinguish transient failures (a timeout, a dependency briefly down) from permanent ones (invalid input). Retry the first with exponential backoff and jitter, up to a limit. Fail the second immediately. After the limit, move the job to a dead-letter state or queue for inspection, with the error recorded, and alert on its depth.
Poison jobs and isolation
A job that crashes the worker every time can take down the pool. Limit attempts, run risky work in isolated processes with resource limits, and separate queues for different job types, so a flood of one kind of work does not starve the others.
Resumability
For long jobs, checkpoint progress: process a large import in chunks and record the last completed chunk. After a crash the job resumes from the checkpoint instead of starting over. Splitting a large job into many smaller tasks (a fan-out), each one idempotent, also lets many workers share the load, and a final step combines the results (a fan-in).
5. Telling the client the result
| Method | How it works | Good for | Cost |
|---|---|---|---|
| Polling | Client asks for status periodically | Simple clients, browsers | Wasted requests, delay up to the interval |
| Webhook (callback) | Server calls a URL the client registered | Server-to-server integrations | Receiver must be reachable and verify and dedupe calls |
| Push (WebSocket, server-sent events, mobile push) | Server pushes an update to the user | Interactive apps | Connection infrastructure |
| Email or message | Notify when done | Very long jobs | Not programmatic |
For polling, return a hint of when to check next, and use backoff in the client so that thousands of waiting clients do not hammer the status endpoint. For webhooks, sign the request so the receiver can verify it, retry with backoff on failure, and include the job identifier, so receivers can deduplicate.
Offer progress where possible. A percentage or a stage name ("encoding", "uploading") makes waiting tolerable and helps debugging.
6. Scheduling, priority and fairness
- Priority queues. Interactive or paid-tier jobs go ahead of bulk work. Use separate queues with dedicated workers, so low-priority volume cannot block urgent work, rather than a single queue with priorities that starve the low end.
- Per-tenant fairness. One customer submitting a million jobs must not delay everyone else. Limit concurrent jobs per tenant, and take work from tenants in rotation.
- Rate limiting. Respect the limits of downstream services. Cap worker concurrency for calls to a rate-limited API.
- Delayed and scheduled jobs. Use a queue that supports a delay, or a scheduler that enqueues at the due time. For recurring jobs, make sure only one instance of the schedule runs, because two schedulers produce duplicates. Elect a leader or use a coordination service.
- Autoscaling. Scale workers by queue depth or age of the oldest job. Account for slow starts and for work that must finish before a worker is removed.
7. Cancellation and cleanup
Let users cancel. The API marks the job cancelling, and workers check the flag between steps and stop cleanly. Know that a job that is already finishing may complete anyway, so treat cancellation as best-effort and say so.
Apply retention to the job table and to stored results: delete or archive old jobs and outputs after a set period. Use time-limited signed links for results.
8. When a simple queue is not enough: workflows
Single jobs are simple. Real processes have steps, branches, waits and compensation: "charge, then reserve stock, then ship, and refund if shipping fails, wait up to a week for a signature". Hand-rolling this on queues leads to scattered state and hard-to-trace failures. A workflow engine (durable execution) records each step's progress in durable storage, so that a crashed workflow resumes exactly where it stopped, with built-in retries, timers and visibility. The next chapter covers when to move to one.
9. A worked example
Problem. Users export their data as a downloadable archive. Exports can take from seconds to half an hour, and 5,000 may be requested at the start of the month.
Reasoning.
- The export is long, variable and heavy, so it is asynchronous.
POST /exportscreates an export job (state queued) and returns 202 with an identifier. A unique key per user and period prevents duplicate exports from double-clicks.- The job goes on a dedicated export queue, so exports do not compete with interactive work. Worker count scales with queue depth, capped to protect the database.
- The worker reads the data in chunks, writes each chunk to temporary storage and records a checkpoint. A crash resumes from the last chunk.
- At the end, it assembles the archive in object storage and marks the job succeeded with a time-limited signed link.
- The client polls with backoff, and the user also receives an email and a push notification with the link when the job finishes.
- A failed job retries a few times with backoff, then goes to the dead-letter state with the error, and the user sees "failed, try again", with support alerted on the dead-letter depth.
- Fairness: a limit of one active export per user, and a limit on concurrent exports per tenant, so the month-start spike drains in order without starving others.
- Old exports are deleted after seven days.
What I would say about the trade-off. "The user waits for an email instead of a page, which is fine for exports. In return the API stays fast, the work survives restarts, and the month-start spike drains at a rate the database can handle."
10. Interview questions and model answers
Q: A request takes two minutes. How do you handle it? I return 202 with a job identifier straight away, store the job in a durable store, put it on a queue, and let workers do the work. The client polls the status, or gets a webhook or a push when done.
Q: A worker crashes halfway through. What happens? The queue's lease expires, and another worker picks the job up. Because delivery is at-least-once, the work is idempotent, and long jobs checkpoint progress so they resume rather than restart.
Q: How do you stop one customer from monopolising the workers? Per-tenant concurrency limits and fair scheduling, separate queues by job class, and rate limits on submission.
Q: How should the client learn the job finished? Polling with backoff for simple clients, a signed webhook for server integrations, and push for interactive apps. The job store is the source of truth in each case.
Q: What is the difference between a queue and the job store? The queue carries work to workers and may lose track of history. The job store records the state, progress and result of each job durably, and is what the API reads.
Q: When do you move beyond a simple queue? When a process has several steps, waits, branches or compensation, I use a workflow engine that persists progress and handles retries and timers.
11. Common mistakes
- Doing the long work inside the request.
- Keeping job state only in the queue or in a worker's memory.
- Non-idempotent jobs on an at-least-once queue.
- No lease or heartbeat, so a crashed worker leaves jobs stuck forever.
- Retrying permanent failures forever, and no dead-letter handling.
- One shared queue where bulk work starves interactive work.
- Clients polling in a tight loop with no backoff.