Guide · Computer vision
Building Real-Time CCTV Analytics with YOLO
Key takeaways
- Sample frames down before the model — the detector is rarely the bottleneck.
- Batch frames from several cameras into one model call for real-time throughput.
- Track identities (DeepSORT/ByteTrack) so deduplication can work at all.
- Deduplicate alerts in a Redis window per camera + track — noise kills operator trust.
By Prasanna Patil · updated
What we're building
A service that watches several CCTV feeds, detects people and crowd-density spikes, and alerts an operator — once per incident, not once per camera per frame. The numbers below come from the system I built for Smart India Hackathon 2023 (a finalist entry) and from the person detection and tracking pipeline I built on CCTV feeds at Zuneko Labs.
The pipeline
RTSP feeds ─► OpenCV reader ─► sample to N fps ─► batch ─► YOLO ─► tracker ─► events (MongoDB)
(per camera) (reconnects) (GPU/CPU) (DeepSORT/ │
ByteTrack) ▼
dedup (Redis, 30 s) ─► alert (WhatsApp / SMS)- Read each stream in its own process and reconnect on failure — cameras drop more often than you expect.
- Sample. Don't process every frame. Decide the frame rate from what you need to catch.
- Batch frames from several cameras into one model call.
- Detect with YOLO; track so a person keeps one ID across frames.
- Store each event: timestamp, camera ID, track ID, bounding box, confidence.
- Deduplicate, then alert.
Throughput: what actually happened
Running four simultaneous streams in real time on a single machine, the frame queue backed up after about 90 seconds. Two changes fixed it: dropping to 5 fps per stream, and batching frames through the model instead of sending them one at a time.
The cost was precision on fast motion. For crowd monitoring that was acceptable — steady throughput mattered more than per-frame accuracy. For something like detecting a fall, it wouldn't be.
Alert deduplication
Overlapping cameras see the same person, and a detector fires on every frame. Without deduplication operators get flooded and start ignoring alerts. I used a Redis key per camera ID and track ID with a 30-second expiry: long enough to suppress duplicates, short enough not to hide a genuine repeat.
def maybe_alert(event):
key = f"alert:{event.camera_id}:{event.track_id}:{event.kind}"
# SET NX EX: only the first event in the window creates the key
if redis.set(key, 1, nx=True, ex=30):
send_alert(event) # Twilio SMS / WhatsApp Business APIThis deduplication logic and the alert pipeline were specifically called out in the hackathon judges' feedback.
Failure modes
- Queue backlog. If processing is slower than arrival, latency grows until alerts are useless. Watch queue depth, and drop frames rather than fall behind.
- Stream drops. RTSP connections die silently; the reader needs timeouts and reconnects.
- Identity switches. Without tracking, one person becomes many IDs and dedup by track fails.
- False positives from lighting changes, reflections and posters of people.
Limitations
- Accuracy depends heavily on camera placement, lighting and resolution; test on your own footage.
- Facial recognition carries legal and privacy obligations — in India, the Digital Personal Data Protection Act, 2023. Decide what you actually need to store.
- The numbers here come from one machine and four streams; more cameras need more hardware or edge devices.
Frequently asked questions
What frame rate do I need for real-time detection?
Far less than the camera produces. On the four-stream build, 5 fps per stream was enough for crowd monitoring once frames were batched — decide from what you need to catch, not from the camera's maximum.
How do I stop duplicate alerts from overlapping cameras?
Deduplicate at the alert layer with a sliding window per camera ID and track ID — 30 seconds in Redis was enough on this build. Tracking must come first, or there's no stable ID to deduplicate on.
YOLOv8 or a custom model?
Pre-trained YOLO handles people and common objects well. Anything specific to your site — uniforms, restricted zones, particular vehicles — needs labelled footage and fine-tuning.
Can this run on the edge instead of a server?
Yes — export the detector through ONNX to TensorRT and run on an NVIDIA Jetson. That trades model size and update convenience for on-site processing and lower bandwidth.
Planning a video-analytics project?
Tell me about the cameras, the hardware and what you need to detect.