1. The ATMS Scale Bottleneck: 1 Camera Every Kilometre
On October 17, 2023, the National Highways Authority of India (NHAI) issued circular 11.53/2023 revising the Advanced Traffic Management System (ATMS) specifications. The mandate required intelligent Pan-Tilt-Zoom (PTZ) traffic monitoring cameras spaced roughly every one kilometre along access-controlled expressways and national corridors.
For a 300 km highway corridor, this introduces over 300 continuous RTSP video streams arriving at the Traffic Management Centre (TMC). Human operators suffer cognitive vigilance decrements within 20 minutes of watching multiple surveillance monitors:
- The False Sense of Coverage: A stationary vehicle in the fast lane, a pedestrian walking in the median, or an animal crossing at night is missed if an operator is observing a different quadrant of video tiles.
- Raw Video vs. Structured Audit Trails: Handing a control room an unindexed video feed forces retroactive investigation after a fatal collision has already occurred.
TMCS (Traffic Monitoring Camera System) was engineered at Arya Omnitalk (Arvind Group) to serve as the autonomous detection foundation beneath the ATMS: watching every video stream continuously, executing 7 deterministic incident verification rules across 11 object classes, and emitting cryptographically signed, version-traceable incident events to operators within seconds.
2. High-Level 8-Stage Processing Pipeline
Nothing becomes an alert in TMCS until every stage in the processing graph agrees. The pipeline isolates hardware communication, computer vision inference, geometric tracking, temporal validation, and cryptographic delivery into strict decoupled layers:
Dynamic capability negotiation across Hikvision, Dahua, Axis, and CP PLUS PTZ units with credential vaulting and non-blocking asynchronous state machines.
25–30 FPS hardware-accelerated GStreamer/MediaMTX decode with frame timestamp monotonicity enforcement and 0.1 ms circuit-breaker dead-camera ejection.
Checks PTZ positioning. When an operator moves the camera or recalls a preset, analytics pause until mechanical motion ceases, loading the preset's dedicated lane polygons.
YOLOv8 detects rigid vehicles (cars, trucks, buses, bikes); RT-DETR detects deformable objects (pedestrians, cows, fallen tire debris). Class-agnostic NMS merges candidates.
Tracks objects across variable-interval video frames using state vector $[x_c, y_c, a, h, \dot{x}_c, \dot{y}_c, \dot{a}, \dot{h}]$. Retains occluded objects for up to 60 seconds.
7 spatial-temporal incident rules evaluate candidate states: stopped vehicles, wrong-way driving, pedestrians on highway, animals, debris, queues, and illegal parking.
State machine manages candidate → confirmed → active → cleared phases. Prevents alert fatigue by collapsing redundant bounding triggers into a single persistent incident entity.
Generates immutable Write-Once-Read-Many (WORM) packages containing the raw frame, SVG bounding overlay, SHA-256 digest, and SSE JSON streams dispatching directly to highway patrol.
3. Dual-Model Detection: YOLOv8 + RT-DETR
Highway computer vision presents an intrinsic architectural trade-off: CNN detectors excel at rigid rectangular vehicles at high speeds, while Deformable Transformer detectors excel at capturing non-rigid, arbitrary-aspect-ratio objects like cattle, fallen cardboard boxes, or crouching pedestrians.
Running two separate forward passes on edge GPU servers would cut framerates below the mandatory 25 FPS NHAI threshold. We resolved this through a Class-Specialized Dual-Engine Scheduler executed via TensorRT:
4. Mathematical Formulation: Δt-Aware Kalman Tracking & Occlusion
Standard Multi-Object Tracking algorithms (such as SORT or standard ByteTrack) assume an idealized constant sampling time ($\Delta t = \text{constant}$). In real-world RTSP streaming over cellular 4G or fiber switches, packet jitter and network lag cause inter-frame intervals to oscillate between 28 ms and 75 ms.
A constant $\Delta t$ assumption causes the Kalman covariance matrix $\mathbf{P}$ to desynchronize, resulting in lost tracks during high-speed vehicle maneuvers. TMCS calculates dynamic delta-time timestamps $\Delta t_k = t_k - t_{k-1}$ per frame:
$$\mathbf{x} = \begin{bmatrix} x_c & y_c & a & h & \dot{x}_c & \dot{y}_c & \dot{a} & \dot{h} \end{bmatrix}^T$$
$$\mathbf{F}(\Delta t_k) = \begin{bmatrix}
\mathbf{I}_{4 \times 4} & \Delta t_k \cdot \mathbf{I}_{4 \times 4} \\
\mathbf{0}_{4 \times 4} & \mathbf{I}_{4 \times 4}
\end{bmatrix}, \quad
\mathbf{Q}(\Delta t_k) = \mathbf{G} \mathbf{Q}_c \mathbf{G}^T$$
$$\mathbf{x}_{k|k-1} = \mathbf{F}(\Delta t_k) \mathbf{x}_{k-1|k-1}$$
$$\mathbf{P}_{k|k-1} = \mathbf{F}(\Delta t_k) \mathbf{P}_{k-1|k-1} \mathbf{F}^T(\Delta t_k) + \mathbf{Q}(\Delta t_k)$$
Where:
- $(x_c, y_c)$ represents bounding box centroid coordinates, $a = w/h$ is the aspect ratio, and $h$ is bounding box height.
- $\mathbf{Q}(\Delta t_k)$ is the piecewise continuous white noise process covariance scaling with $\Delta t_k^3$ and $\Delta t_k^2$.
- Category-Aware Retention: Fast vehicles are dropped if unobserved for 30 frames (1.2 s). However, Pedestrians and Animals (which frequently disappear behind 16-wheel semi-trucks) have their state propagation extended to 60 seconds ($1,500\text{ frames}$) before track eviction, preventing track ID fragmentation when the person emerges on the other side.
5. Rider-Aware Occupant Tagging (-95%+ False Alarms)
One of the largest operational headaches in Indian highway surveillance is the "Pedestrian Alarm Storm." Because object detectors output independent classes for `person` and `motorcycle`, a motorcycle carrying two passengers regularly fires two simultaneous "Pedestrian on Highway" alerts to the control room.
TMCS eliminates this by computing geometric spatial containment before passing detections to the incident rule engine:
$$\text{Containment}(P_i, V_j) = \frac{\text{Area}(B_{P_i} \cap B_{V_j})}{\text{Area}(B_{P_i})}$$
$$\text{IsOccupant}(P_i) = \begin{cases}
\text{True}, & \text{if } \max_{j} \text{Containment}(P_i, V_j) \ge 0.55 \text{ and } y_{\text{bottom}, P_i} \le y_{\text{bottom}, V_j} + 0.15 h_{V_j} \\
\text{False}, & \text{otherwise}
\end{cases}$$
If a detected person's bounding box is more than 55% contained inside a two-wheeler, tractor, or auto-rickshaw bounding box, and the person's feet do not extend below the vehicle's road contact patch, the person is tagged as an `occupant`. Only unparented persons can trigger the "Pedestrian on Highway" incident rule, cutting false control room alarms by over 95%.
6. PTZ Safety & ONVIF Circuit Breaker
Unlike static gantry cameras (like VIDES), PTZ cameras move physically. If an operator uses a joystick to pan the camera 90 degrees to inspect a toll plaza lane, a naive analytics engine would evaluate moving road pixels against old static ROI polygons, triggering dozens of false "Wrong-Way Driving" alarms.
TMCS continuously monitors ONVIF PTZ status vectors. Whenever Pan, Tilt, or Zoom velocity is non-zero, the tracking matrix is instantly flushed, and all incident rules enter an atomic PAUSE state.
Once mechanical motion ceases and the camera locks onto a target preset (e.g., "Preset 4: Northbound Curve"), TMCS fetches that specific preset's calibrated lane geometry via a single batched SQL query in <10 ms.
If an operator manually moves a PTZ camera for inspection and forgets to return it, an automated watchdog timer expires after 120 seconds, issuing an ONVIF AbsoluteMove command to return the camera to its primary surveillance preset.
7. Production Implementation: Tracker & Wrong-Way Rule
import math
import numpy as np
from typing import List, Dict, Any, Optional
class TMCSIncidentEngine:
def __init__(self, lane_heading_deg: float, speed_threshold_kmh: float = 15.0):
self.lane_heading_deg = lane_heading_deg
self.lane_vector = np.array([
math.cos(math.radians(lane_heading_deg)),
math.sin(math.radians(lane_heading_deg))
])
self.speed_threshold_kmh = speed_threshold_kmh
self.wrong_way_candidates: Dict[int, int] = {} # track_id -> consecutive_frames
def filter_rider_occupants(self, detections: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
"""
Suppresses false pedestrian alerts by tagging riders inside two-wheelers.
"""
persons = [d for d in detections if d['class_name'] == 'pedestrian']
vehicles = [d for d in detections if d['class_name'] in ['motorcycle', 'bicycle', 'auto_rickshaw']]
filtered = []
for p in detections:
if p['class_name'] != 'pedestrian':
filtered.append(p)
continue
# Check geometric containment against all active vehicles
p_box = p['bbox'] # [x1, y1, x2, y2]
p_area = (p_box[2] - p_box[0]) * (p_box[3] - p_box[1])
is_occupant = False
for v in vehicles:
v_box = v['bbox']
# Compute intersection
ix1 = max(p_box[0], v_box[0])
iy1 = max(p_box[1], v_box[1])
ix2 = min(p_box[2], v_box[2])
iy2 = min(p_box[3], v_box[3])
if ix2 > ix1 and iy2 > iy1:
intersection_area = (ix2 - ix1) * (iy2 - iy1)
if (intersection_area / p_area) >= 0.55:
is_occupant = True
break
if not is_occupant:
filtered.append(p) # Retain genuine pedestrian walking on road
return filtered
def evaluate_wrong_way_rule(self, track_id: int, trajectory_pts: List[np.ndarray], fps: float) -> bool:
"""
Vector angle dot-product against calibrated lane direction vector.
Requires 15 consecutive frames of opposite velocity to trigger incident.
"""
if len(trajectory_pts) < 10:
return False
# Calculate smoothed velocity vector across last 10 points
start_pt = trajectory_pts[-10]
end_pt = trajectory_pts[-1]
displacement = end_pt - start_pt
distance_px = np.linalg.norm(displacement)
if distance_px < 15.0: # Below motion noise threshold
return False
vel_vector = displacement / distance_px
# Dot product: cos(theta) = u . v
cos_similarity = np.dot(vel_vector, self.lane_vector)
# Opposite direction condition: angle > 135 degrees (cos < -0.707)
if cos_similarity <= -0.707:
self.wrong_way_candidates[track_id] = self.wrong_way_candidates.get(track_id, 0) + 1
else:
self.wrong_way_candidates[track_id] = max(0, self.wrong_way_candidates.get(track_id, 0) - 1)
# Confirm incident if sustained for 15 consecutive evaluated frames (~0.5s)
return self.wrong_way_candidates[track_id] >= 15
8. Internal Engineering Benchmarks
We benchmarked TMCS across thousands of hours of live highway footage from operational national corridors:
| Evaluation Parameter | Legacy NVR Video Analytics | TMCS Production Incident Layer | System Impact |
|---|---|---|---|
| PTZ Motion Behavior | Fires hundreds of false alarms during panning | Atomic track pause & preset ROI reload | Zero false alerts during camera adjustment |
| Rider vs Pedestrian Disambiguation | None (2 alerts per passing motorcycle) | Geometric spatial containment filter | -95%+ reduction in false pedestrian alerts |
| Occlusion Handling | Drops track after 15–30 frames | Δt-aware Kalman retention up to 60s | Tracks cows/people emerging from behind trucks |
| Evidence Authentication | Mutable JPEG files on filesystem | WORM storage with SHA-256 digest | Court-admissible tamper-evident proof |
9. SHA-256 Court-Grade Evidence Architecture
Under Indian evidence law and police dispute standards, electronic evidence must satisfy strict tamper-evident chain-of-custody protocols. If an operator can open an incident snapshot in GIMP or Photoshop and alter a license plate or vehicle position, the evidence is discarded in court.
{
"event_id": "tmcs_inc_2026_0312_98412",
"incident_type": "STOPPED_VEHICLE",
"camera_id": "PTZ_NH48_KM_142_NORTH",
"preset_id": 2,
"timestamp_utc": "2026-03-12T14:22:18.421Z",
"model_checkpoint_sha256": "8f3b2a19c...d4e8",
"rule_config_version": "v2.4.1",
"raw_frame_sha256": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
"svg_overlay_sha256": "4b227777d4dd1fc61c6f884f48641d02b4d121d3fd328cb08b5531fcacdabf8a"
}
TMCS writes evidence to Write-Once-Read-Many (WORM) storage. The underlying raw video frame is never compressed or edited. Bounding boxes, track vectors, and timestamp metadata are stored as a separate cryptographically linked SVG overlay layer, ensuring 100% legal admissibility.
10. Critical Edge Failure Modes & Production Hardening
Deploying autonomous surveillance across live national highways revealed severe physical and software failure modes: