Skip to content

Guide · Computer vision

Building Real-Time CCTV Analytics with YOLO

Key takeaways

  • Sample frames down before the model — the detector is rarely the bottleneck.
  • Batch frames from several cameras into one model call for real-time throughput.
  • Track identities (DeepSORT/ByteTrack) so deduplication can work at all.
  • Deduplicate alerts in a Redis window per camera + track — noise kills operator trust.

By Prasanna Patil · updated

01

What we're building

A service that watches several CCTV feeds, detects people and crowd-density spikes, and alerts an operator — once per incident, not once per camera per frame. The numbers below come from the system I built for Smart India Hackathon 2023 (a finalist entry) and from the person detection and tracking pipeline I built on CCTV feeds at Zuneko Labs.

02

The pipeline

RTSP feeds ─► OpenCV reader ─► sample to N fps ─► batch ─► YOLO ─► tracker ─► events (MongoDB)
 (per camera)   (reconnects)                       (GPU/CPU)  (DeepSORT/    │
                                                              ByteTrack)    ▼
                                                                    dedup (Redis, 30 s) ─► alert (WhatsApp / SMS)
  1. Read each stream in its own process and reconnect on failure — cameras drop more often than you expect.
  2. Sample. Don't process every frame. Decide the frame rate from what you need to catch.
  3. Batch frames from several cameras into one model call.
  4. Detect with YOLO; track so a person keeps one ID across frames.
  5. Store each event: timestamp, camera ID, track ID, bounding box, confidence.
  6. Deduplicate, then alert.
03

Throughput: what actually happened

Running four simultaneous streams in real time on a single machine, the frame queue backed up after about 90 seconds. Two changes fixed it: dropping to 5 fps per stream, and batching frames through the model instead of sending them one at a time.

The cost was precision on fast motion. For crowd monitoring that was acceptable — steady throughput mattered more than per-frame accuracy. For something like detecting a fall, it wouldn't be.

04

Alert deduplication

Overlapping cameras see the same person, and a detector fires on every frame. Without deduplication operators get flooded and start ignoring alerts. I used a Redis key per camera ID and track ID with a 30-second expiry: long enough to suppress duplicates, short enough not to hide a genuine repeat.

def maybe_alert(event):
    key = f"alert:{event.camera_id}:{event.track_id}:{event.kind}"
    # SET NX EX: only the first event in the window creates the key
    if redis.set(key, 1, nx=True, ex=30):
        send_alert(event)          # Twilio SMS / WhatsApp Business API
Illustrative sketch.

This deduplication logic and the alert pipeline were specifically called out in the hackathon judges' feedback.

05

Failure modes

  • Queue backlog. If processing is slower than arrival, latency grows until alerts are useless. Watch queue depth, and drop frames rather than fall behind.
  • Stream drops. RTSP connections die silently; the reader needs timeouts and reconnects.
  • Identity switches. Without tracking, one person becomes many IDs and dedup by track fails.
  • False positives from lighting changes, reflections and posters of people.
06

Limitations

  • Accuracy depends heavily on camera placement, lighting and resolution; test on your own footage.
  • Facial recognition carries legal and privacy obligations — in India, the Digital Personal Data Protection Act, 2023. Decide what you actually need to store.
  • The numbers here come from one machine and four streams; more cameras need more hardware or edge devices.
07

Frequently asked questions

What frame rate do I need for real-time detection?

Far less than the camera produces. On the four-stream build, 5 fps per stream was enough for crowd monitoring once frames were batched — decide from what you need to catch, not from the camera's maximum.

How do I stop duplicate alerts from overlapping cameras?

Deduplicate at the alert layer with a sliding window per camera ID and track ID — 30 seconds in Redis was enough on this build. Tracking must come first, or there's no stable ID to deduplicate on.

YOLOv8 or a custom model?

Pre-trained YOLO handles people and common objects well. Anything specific to your site — uniforms, restricted zones, particular vehicles — needs labelled footage and fine-tuning.

Can this run on the edge instead of a server?

Yes — export the detector through ONNX to TensorRT and run on an NVIDIA Jetson. That trades model size and update convenience for on-site processing and lower bandwidth.

Planning a video-analytics project?

Tell me about the cameras, the hardware and what you need to detect.