Skip to content

Guide · Python backend

FastAPI, Celery and Redis: A Production Architecture Guide

Key takeaways

  • FastAPI validates and enqueues; Celery workers do the slow work; Redis is only the broker.
  • The database row — not Redis — is the source of truth for every job.
  • Late acks plus idempotent tasks mean a dead worker loses nothing.
  • Downtime loses events before they're enqueued; you need a backfill job, not just a queue.

By Prasanna Patil · updated

01

The problem this solves

A webhook or API call arrives, and handling it means calling an LLM, running a model, or writing to several tables. Done inline, the request takes seconds; the caller times out and retries; you process the same event twice; and under load, the web workers are all stuck waiting.

I hit this building WhatsApp AI agents for a client at Shivohini TechAI: message processing, AI inference and database writes were all on the request thread. Moving them to Redis + Celery fixed the timeouts. The CCTV analytics backend I built for Smart India Hackathon used the same pattern — a Celery worker pool processed video frames so the Django API never blocked.

02

The architecture

client / webhook
      │  POST
      ▼
┌─────────────┐   1. validate (Pydantic)
│  FastAPI    │   2. insert row: status = "queued"
│             │   3. task.delay(id)  ──────────►  Redis (broker)
└─────────────┘   4. return 202 + id                    │
                                                         ▼
                                               ┌──────────────────┐
                                               │  Celery workers   │  inference, LLM calls,
                                               │  (N processes)    │  3rd-party APIs, DB writes
                                               └──────────────────┘
                                                         │
                                               update row: "done" / "failed"
Request path on the left returns in milliseconds; everything slow happens on the right.

The database row is the source of truth for the job, not Redis. If Redis restarts or a message is lost, you can still find every job that never finished and re-enqueue it.

03

A minimal version

# tasks.py
from celery import Celery

app = Celery("worker", broker="redis://redis:6379/0")
app.conf.task_acks_late = True            # ack after the task finishes, not when it starts
app.conf.worker_prefetch_multiplier = 1   # don't let one worker hoard slow tasks

@app.task(bind=True, autoretry_for=(TimeoutError,), retry_backoff=True, max_retries=5)
def handle_message(self, message_id: str):
    msg = db.get(message_id)
    if msg.status == "done":              # idempotent: a re-delivered task is a no-op
        return
    reply = llm.generate(msg.text, context=history_for(msg.user_id))
    db.save_reply(message_id, reply, status="done")

# api.py
@router.post("/webhook", status_code=202)
async def webhook(event: IncomingEvent):
    message_id = db.insert_if_new(event.id, event.payload)   # unique on provider's event id
    if message_id:
        handle_message.delay(message_id)
    return {"id": event.id}
Illustrative sketch — names and settings are simplified.
04

Failure modes that bite in production

  • The provider retries the webhook. Store the provider's event ID with a unique constraint and ignore duplicates, or the same message gets two replies.
  • A worker dies mid-task. With the default early acknowledgement the task is gone. Late acks plus idempotent tasks mean it is re-delivered and finishes once.
  • Downtime means missed events. Queues don't help if the event never arrived. For the WhatsApp document ingestion at Zuneko Labs I added a recovery job that backfills files missed during downtime or network failures.
  • Models loaded per task. Load a model once per worker process, not inside the task, or each task pays the load time and memory spikes.
  • Retries that hammer a failing API. Use exponential backoff and a retry cap, then mark the job failed so someone can look at it.
05

Trade-offs

  • You now operate a broker and a worker pool. For a few jobs a minute, FastAPI's BackgroundTasks or a scheduled job may be enough.
  • Results arrive asynchronously, so clients need polling, a callback, or a push channel to find out a job is done.
  • Redis as a broker is simple and fast, but it isn't a durable log — keep the job state in PostgreSQL.
06

Limitations of this guide

It covers one service with one queue. Multiple queues with priorities, very long-running jobs, and exactly-once delivery across systems each need more than this.

07

Frequently asked questions

Should job state live in Redis or PostgreSQL?

PostgreSQL. Redis is the broker, not the ledger — it doesn't persist a durable history. Keep a row per job in PostgreSQL and re-enqueue from it if Redis loses messages.

What does task_acks_late actually change?

With the default early ack, a worker that dies mid-task takes the task with it — the broker already considers it done. Late acks tell the broker to re-deliver the task, which is only safe when tasks are idempotent.

When is Celery overkill?

For a few jobs a minute, FastAPI's BackgroundTasks or a scheduled job is enough. The queue earns its keep when work is slow, spiky, or must survive worker restarts.

How do clients know when an async job is done?

Polling a status endpoint keyed on the job ID, a webhook callback, or a push channel. The 202 response should always include an ID the client can track.

Stuck on a slow or flaky backend?

Describe the setup and where it fails.