Draft — not published. This page is noindex and not in the sitemap until it's approved.
Guide · Python backend
Celery Task Recovery: What Happens When Things Break
Key takeaways
- Default early acks mean a dead worker takes its task with it. Late acks fix that — only with idempotent tasks.
- Retries need exponential backoff and a cap; then a failed state someone actually monitors.
- Queues can't help with events that were never enqueued — downtime needs a backfill job.
- The database row is the source of truth for job state; Redis is just the broker.
By Prasanna Patil · updated
The failure modes, in order of how often they bite
- A worker dies mid-task. The default early acknowledgement already told the broker the task succeeded.
- A task fails repeatedly. Naive retries hammer a struggling third-party API and bury the real signal.
- Your service was down. Nothing was enqueued during the outage, so the queue is empty and the data is missing.
Each has a specific fix. They compose into the settings block below — the one I ship on every Celery service, including the WhatsApp document ingestion at Zuneko Labs.
The configuration that survives a dead worker
app.conf.task_acks_late = True # ack AFTER the task finishes
app.conf.worker_prefetch_multiplier = 1 # one task in flight per worker process
app.conf.task_reject_on_worker_lost = True # re-deliver tasks interrupted by a dead worker
@app.task(bind=True, autoretry_for=(TransientError,),
retry_backoff=True, retry_backoff_max=600, max_retries=5)
def handle_message(self, message_id: str):
row = db.get(message_id)
if row.status == "done": # idempotency guard: a re-delivery is a no-op
return
...
db.save_reply(message_id, reply, status="done")Late acks without idempotency is how you double-send messages. The status check on the row is the idempotency guard — with it, re-delivery is safe; without it, late acks are a liability.
The gap no queue setting fixes
If your service (or the network in front of it) is down, the provider's webhook never reaches you. The queue holds nothing because nothing was ever enqueued. For the WhatsApp document ingestion at Zuneko Labs I added a recovery job: on a schedule, it asks the source of truth what's missing — messages without a processed row in the window — and backfills them. The database is what makes this possible; this is a big part of why job state lives in PostgreSQL and not in Redis.
Monitoring: what to alert on
- Rows stuck in
queuedpast a threshold — a dead worker pool or a lost broker. - Tasks exhausted retries and entered
failed— needs a human, not another retry. - Backfill jobs finding work — a signal your ingestion path has a gap worth investigating.
Limitations of this guide
It covers Redis-broker Celery with a database-backed state model. Kafka-style event sourcing, exactly-once pipelines, and very long-running jobs (minutes to hours) need different machinery.
Frequently asked questions
Do late acks hurt throughput?
Slightly — with prefetch_multiplier=1 a worker holds one unacked task instead of hoarding a batch. For slow tasks (inference, LLM calls) that's what you want anyway; the hoarding was starving the other workers.
How long should the backfill window look?
Overlap it generously — look back past the longest outage you've had, not just since the last successful run. The idempotency guard makes reprocessing cheap; missing a window is what's expensive.
Redis or RabbitMQ as the broker?
Redis is fine for most workloads and is what I default to, with the caveat that it's not a durable log. If your requirements include broker-level durability guarantees you can't reconstruct from the database, that's the case for RabbitMQ.
Losing tasks when workers die?
Describe the queue setup and what disappears.