CCTV Analytics for Crime & Crowd Management
Built for Smart India Hackathon 2023 — a government problem statement on public safety monitoring. We reached the finals.
By Prasanna Patil, AI Engineer
result
Finalist
Smart India Hackathon 2023
streams
4×
simultaneous CCTV feeds
alert dedup
30 s
Redis window per camera + track
Summary
A system that watches several CCTV feeds at once, detects people and crowd-density spikes, and alerts operators over SMS and WhatsApp — without sending five copies of the same alert when overlapping cameras see the same person. It covered crowd counting, crime detection and facial recognition.
The problem statement
The Ministry of Home Affairs wanted a system that could analyse CCTV footage in real time to detect crowd density spikes, suspicious behaviour, and persons of interest — and alert law enforcement without a human watching every feed.
My role
I worked on the backend and the computer-vision pipeline: the APIs and background workers for real-time video analysis, the OpenCV pipeline, and the alert dispatch with Twilio and WhatsApp.
Architecture
Django backend with a Celery worker pool for processing video frames off the main thread. MongoDB for event storage (timestamps, bounding boxes, confidence scores, camera IDs). OpenCV for frame extraction, YOLO for person detection, a separate model for anomaly classification.
The hard part wasn't detection — it was the alert pipeline. Twilio for SMS, WhatsApp Business API for operator notifications. Had to make sure duplicate events (same person, multiple cameras) didn't fire duplicate alerts. Solved it with a Redis-based deduplication window.
Technology choices
- Celery workers so frame processing never blocks the API.
- MongoDB because each detection event is a document — timestamp, camera, boxes, scores.
- Redis for the short-lived deduplication keys; they expire on their own after the window.
- Twilio + WhatsApp Business API so alerts reach operators on their phones.
What broke under pressure
Processing four simultaneous streams in real time on a single machine was too slow — frame queue backed up after about 90 seconds.
Fixed it by dropping to 5fps per stream and batching frames through the model instead of one at a time. Lost some precision on fast motion, which was acceptable for crowd monitoring. The tradeoff was worth it — consistent throughput mattered more than per-frame accuracy at low motion speeds.
Result
Finalist. Presented to a panel that included government officials. The deduplication logic and alert pipeline were specifically called out in feedback. The final build handled four simultaneous feeds on one machine at 5 fps per stream.
Limitations
- At 5 fps, fast movement loses precision. Fine for crowds; not for, say, tracking a running person.
- Everything ran on a single machine, so the stream count was capped by that hardware.
- It was built and evaluated as a hackathon prototype, not a deployed city system.
What I'd improve next
Export the detector to ONNX/TensorRT to get back to a higher frame rate on the same hardware, and replace the fixed 30-second dedup window with logic based on track lifecycles instead of a flat timer.
Have cameras and a problem like this?
Tell me about the feeds, the hardware and what you need to catch.