Service
Computer vision for CCTV and video: detection, tracking and alerts
I build video pipelines that turn camera feeds into events someone can act on — who or what was detected, where, and when — and that keep up in real time on the hardware you actually have.
Who this is for
- Organisations with existing CCTV who want automated monitoring instead of people watching every feed.
- Teams that need people detection, tracking or ANPR as part of a larger product.
- Teams with a working vision model that has to run faster, or on edge hardware.
Problems I solve
Nobody can watch every camera, so incidents are found after the fact.
Run detection on each feed and raise an alert only for the events that matter — crowd density spikes, restricted-area entries, persons of interest.
The same incident fires a flood of alerts from overlapping cameras.
Deduplicate at the alert layer. On the Smart India Hackathon build I used a Redis sliding window of 30 seconds per camera ID and track ID.
The pipeline falls behind real time.
Measure, then trade frame rate for throughput. Four simultaneous streams on one machine backed up after about 90 seconds; dropping to 5 fps and batching frames through the model fixed it.
The same person gets a new ID every few frames.
Add multi-frame tracking (DeepSORT or ByteTrack) on top of detection, as in the real-time person detection and tracking pipeline I built on CCTV feeds at Zuneko Labs.
What you get
- A detection and tracking pipeline (OpenCV + YOLO + DeepSORT/ByteTrack) for your feeds.
- An event store and API: timestamps, camera IDs, bounding boxes, confidence scores.
- Alerting over WhatsApp, SMS (Twilio) or webhooks, with deduplication.
- Model export and optimisation (ONNX, TensorRT) for edge inference on NVIDIA Jetson.
- A throughput report on your hardware: streams, frame rate and latency you can expect.
Technology, and where it fits
- OpenCV
- Frame extraction and pre-processing from RTSP/CCTV streams.
- YOLO
- Person and object detection.
- DeepSORT / ByteTrack
- Keeping identities consistent across frames.
- ONNX, TensorRT, NVIDIA Jetson
- Faster inference and edge deployment.
- Celery + Redis
- Frame processing off the main thread; alert deduplication windows.
- MongoDB, Twilio, WhatsApp Business API
- Event persistence and incident notifications.
Evidence
- CCTV analytics for crime and crowd management
Smart India Hackathon 2023 finalist: four feeds, YOLO detection, Redis-deduplicated alerts.
- Current work at Zuneko Labs
Real-time person detection and tracking on CCTV feeds; ANPR/CCTV analytics.
Trade-offs worth knowing
Frame rate vs. accuracy
Processing every frame is rarely needed. Lower frame rates and batching buy throughput, at the cost of precision on fast motion — acceptable for crowd monitoring, not for every use case.
Edge vs. server
Edge devices (Jetson) cut bandwidth and keep video on-site, but limit model size. A central GPU server is simpler to update but needs the streams sent to it.
Pre-trained vs. custom models
Pre-trained YOLO handles people and common objects well. Anything specific to your site usually needs labelled footage and fine-tuning, which takes time to collect.
Limitations
- Accuracy depends on camera angle, lighting and resolution. I won't promise a number before testing on your footage.
- Facial recognition and surveillance come with legal and privacy obligations (in India, the Digital Personal Data Protection Act, 2023). You own that decision; I'll build in data minimisation where you want it.
- I don't supply or install camera hardware.
Frequently asked questions
Do I need to replace my existing cameras?
No. The pipeline connects to standard RTSP streams from your existing CCTV. If your cameras can stream, they can be analysed.
How many cameras can one server handle?
It depends on the hardware and frame rate. On the Smart India Hackathon build, one machine handled four simultaneous streams at 5 fps per stream. You get a measured throughput report before deployment.
How accurate is the detection?
Accuracy depends on camera angle, lighting and resolution, so I test on your footage before promising numbers. Pre-trained YOLO handles people and common objects well; site-specific detection needs labelled footage and fine-tuning.
What about privacy and surveillance law?
Surveillance in India is subject to the Digital Personal Data Protection Act, 2023. You own the compliance decision; I build in data minimisation — retention limits, region masking — where you want it.
How we'd work together
- A short call to understand the problem, your current system and your constraints.
- A written scope: what will be built, what won't, and how we'll know it works.
- Build in small milestones you can review, with the code in your repository.
- Handover: documentation, deployment notes and a walkthrough.
Open to full-time roles as well as freelance and consulting work.
Discuss a project
Tell me what you're building and where it's getting stuck. Freelance, consulting and full-time conversations are all welcome.