Guide · Python backend
FastAPI, Celery and Redis: A Production Architecture Guide
Key takeaways
- FastAPI validates and enqueues; Celery workers do the slow work; Redis is only the broker.
- The database row — not Redis — is the source of truth for every job.
- Late acks plus idempotent tasks mean a dead worker loses nothing.
- Downtime loses events before they're enqueued; you need a backfill job, not just a queue.
By Prasanna Patil · updated
The problem this solves
A webhook or API call arrives, and handling it means calling an LLM, running a model, or writing to several tables. Done inline, the request takes seconds; the caller times out and retries; you process the same event twice; and under load, the web workers are all stuck waiting.
I hit this building WhatsApp AI agents for a client at Shivohini TechAI: message processing, AI inference and database writes were all on the request thread. Moving them to Redis + Celery fixed the timeouts. The CCTV analytics backend I built for Smart India Hackathon used the same pattern — a Celery worker pool processed video frames so the Django API never blocked.
The architecture
client / webhook
│ POST
▼
┌─────────────┐ 1. validate (Pydantic)
│ FastAPI │ 2. insert row: status = "queued"
│ │ 3. task.delay(id) ──────────► Redis (broker)
└─────────────┘ 4. return 202 + id │
▼
┌──────────────────┐
│ Celery workers │ inference, LLM calls,
│ (N processes) │ 3rd-party APIs, DB writes
└──────────────────┘
│
update row: "done" / "failed"The database row is the source of truth for the job, not Redis. If Redis restarts or a message is lost, you can still find every job that never finished and re-enqueue it.
A minimal version
# tasks.py
from celery import Celery
app = Celery("worker", broker="redis://redis:6379/0")
app.conf.task_acks_late = True # ack after the task finishes, not when it starts
app.conf.worker_prefetch_multiplier = 1 # don't let one worker hoard slow tasks
@app.task(bind=True, autoretry_for=(TimeoutError,), retry_backoff=True, max_retries=5)
def handle_message(self, message_id: str):
msg = db.get(message_id)
if msg.status == "done": # idempotent: a re-delivered task is a no-op
return
reply = llm.generate(msg.text, context=history_for(msg.user_id))
db.save_reply(message_id, reply, status="done")
# api.py
@router.post("/webhook", status_code=202)
async def webhook(event: IncomingEvent):
message_id = db.insert_if_new(event.id, event.payload) # unique on provider's event id
if message_id:
handle_message.delay(message_id)
return {"id": event.id}Failure modes that bite in production
- The provider retries the webhook. Store the provider's event ID with a unique constraint and ignore duplicates, or the same message gets two replies.
- A worker dies mid-task. With the default early acknowledgement the task is gone. Late acks plus idempotent tasks mean it is re-delivered and finishes once.
- Downtime means missed events. Queues don't help if the event never arrived. For the WhatsApp document ingestion at Zuneko Labs I added a recovery job that backfills files missed during downtime or network failures.
- Models loaded per task. Load a model once per worker process, not inside the task, or each task pays the load time and memory spikes.
- Retries that hammer a failing API. Use exponential backoff and a retry cap, then mark the job failed so someone can look at it.
Trade-offs
- You now operate a broker and a worker pool. For a few jobs a minute, FastAPI's
BackgroundTasksor a scheduled job may be enough. - Results arrive asynchronously, so clients need polling, a callback, or a push channel to find out a job is done.
- Redis as a broker is simple and fast, but it isn't a durable log — keep the job state in PostgreSQL.
Limitations of this guide
It covers one service with one queue. Multiple queues with priorities, very long-running jobs, and exactly-once delivery across systems each need more than this.
Frequently asked questions
Should job state live in Redis or PostgreSQL?
PostgreSQL. Redis is the broker, not the ledger — it doesn't persist a durable history. Keep a row per job in PostgreSQL and re-enqueue from it if Redis loses messages.
What does task_acks_late actually change?
With the default early ack, a worker that dies mid-task takes the task with it — the broker already considers it done. Late acks tell the broker to re-deliver the task, which is only safe when tasks are idempotent.
When is Celery overkill?
For a few jobs a minute, FastAPI's BackgroundTasks or a scheduled job is enough. The queue earns its keep when work is slow, spiky, or must survive worker restarts.
How do clients know when an async job is done?
Polling a status endpoint keyed on the job ID, a webhook callback, or a push channel. The 202 response should always include an ID the client can track.
Stuck on a slow or flaky backend?
Describe the setup and where it fails.