RoadTrace: Video-Based Driving Behavior Assessment¶

Rajesh Nandipaty · August 3, 2025 · A local pipeline for analyzing dashcam footage and generating scored driving-safety reports.

Vehicles are detected and tracked across frames using a Kalman filter, ranged using a ground-plane homography, and evaluated against physically defined event rules for tailgating, hard braking, cut-ins, and rolling stops. The pipeline produces an annotated video and a quantitative driver-safety score, providing an insurance-telematics-style analysis without relying on an insurer or external service. This notebook demonstrates the complete pipeline using a synthetic drive with known ground truth. Real dashcam footage is unsuitable for distribution in a portfolio, while a manually constructed synthetic sequence would not provide a meaningful test of the detection and downstream analysis stages. Instead, the scene is rendered for visual inspection, while object detections are supplied through a scripted detection stream derived from the same ground truth.

All stages following detection operate normally, including tracking, geometric ranging, ego-vehicle dynamics, event detection, and safety scoring. Because the underlying ground truth is known, the Evaluation section can quantitatively measure ranging error and event-detection accuracy rather than relying solely on qualitative claims. The production front end uses a YOLOv8 detector exported to ONNX. Its performance on real-world dashcam footage remains the primary component requiring validation with real driving data and is identified as a next step.

1. Configuration¶

The fixed camera configuration and event thresholds used throughout the pipeline. Thresholds are treated as configuration rather than implementation logic: each event rule evaluates a physical quantity against a corresponding value defined here.

In [22]:
import os, sys, json, subprocess
from pathlib import Path
sys.path.insert(0, os.path.abspath("."))
# Make ffmpeg (bundled with imageio) available for any re-encode step.
import imageio_ffmpeg
os.environ["PATH"] = os.path.dirname(imageio_ffmpeg.get_ffmpeg_exe()) + os.pathsep + os.environ["PATH"]

import numpy as np, pandas as pd
from driveaudit.geometry import Calibration
from driveaudit.events import Rules

CACHE, OUT = Path("cache"), Path("output")
OUT.mkdir(exist_ok=True)
rules = Rules()
print("Event Thresholds:")
print(f"  Tailgating     Headway < {rules.headway_s} s for >= {rules.headway_persist_s} s while moving > {rules.min_speed_mps} m/s.")
print(f"  Hard braking   Deceleration beyond {rules.hard_brake_mps2} m/s^2.")
print(f"  Cut in         Vehicle enters ego lane within {rules.cut_in_range_m} m.")
print(f"  Rolling stop   Stop sign within {rules.stop_sign_range_m} m and speed never below {rules.stop_speed_mps} m/s.")
print(f"  Score          Each {rules.score_half_life:.0f} penalty points per 10 min halves the score.")
Event Thresholds:
  Tailgating     Headway < 2.0 s for >= 1.5 s while moving > 5.0 m/s.
  Hard braking   Deceleration beyond -3.0 m/s^2.
  Cut in         Vehicle enters ego lane within 25.0 m.
  Rolling stop   Stop sign within 15.0 m and speed never below 1.0 m/s.
  Score          Each 150 penalty points per 10 min halves the score.

2. Footage and Calibration¶

The synthetic drive is generated once and cached, following the same approach used by the reference project for its retrieval stage. Camera calibration uses a flat-road pinhole model defined by the focal length in pixels, camera height, and image row corresponding to the horizon. These parameters are sufficient to convert the image-row position of a vehicle’s tyre-contact point into an estimated road distance.

In [23]:
from scenarios.synthetic_drive import render, CALIB
import cv2

if not (CACHE / "synthetic-drive.mp4").exists():
    render(CACHE)                      # Regenerate footage, detections, speed log, ground truth.
calib = Calibration.load(CACHE / "calibration.json")

cap = cv2.VideoCapture(str(CACHE / "synthetic-drive.mp4"))
info = dict(fps=cap.get(cv2.CAP_PROP_FPS), frames=int(cap.get(cv2.CAP_PROP_FRAME_COUNT)),
            width=int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)), height=int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)))
cap.set(cv2.CAP_PROP_POS_FRAMES, 150); ok, raw = cap.read(); cap.release()
cv2.imwrite(str(OUT / "raw_frame.png"), raw)
print("Clip:", info, f"= {info['frames']/info['fps']:.0f} s")

# Calibration round-trip: place a point at 15 m, project to image, recover it.
u, v = calib.image_point(0.0, 15.0)
X, Z = calib.ground_point((u-30, v-40, u+30, v))
print(f"Calibration round-trip: 15.0 m -> row {v:.1f} -> {Z:.2f} m recovered.")
Clip: {'fps': 15.0, 'frames': 1800, 'width': 1280, 'height': 720} = 120 s
Calibration round-trip: 15.0 m -> row 416.7 -> 15.00 m recovered.
In [24]:
from IPython.display import Image, display
display(Image(str(OUT / "raw_frame.png"), width=680))
No description has been provided for this image

3. Detection (Perception)¶

Detection establishes what objects are present in each frame. In the production pipeline, a YOLOv8 model is exported to ONNX and executed on the CPU without a PyTorch runtime. The detector decodes bounding boxes and class scores, then filters the results to relevant vehicle classes and stop signs. For this demonstration, the detector is replaced by a scripted detection stream derived from the known ground truth. Each frame contains one detection record per visible object, with controlled positional jitter and occasional dropouts introduced to approximate the noise and intermittency of a real object detector. This preserves the downstream pipeline while keeping the synthetic evaluation reproducible and its ground truth explicit.

In [25]:
det = pd.read_parquet(CACHE / "scripted_detections.parquet")
names = {2: "car", 3: "motorcycle", 5: "bus", 7: "truck", 11: "stop sign"}
print(f"{len(det):,} detections over {det['frame'].nunique():,} frames.")
print("By class:", {names.get(k, k): int(v) for k, v in det["cls"].value_counts().items()})
print(f"Mean confidence: {det['score'].mean():.2f}")
2,201 detections over 1,756 frames.
By class: {'car': 2054, 'stop sign': 147}
Mean confidence: 0.86

4. Tracking (State Estimation)¶

Tracking establishes object identity across consecutive frames. Each track maintains a linear Kalman filter based on a constant-velocity motion model. The model produces a prior estimate of the object state, the current detection provides the measurement, and the two are combined into a posterior estimate according to their respective covariances. This implements the prior-measurement-posterior estimation cycle described in the estimation framework.

Detections are associated with predicted tracks using Intersection over Union (IoU) and the Hungarian algorithm. A track is confirmed only after a minimum number of consistent detections, reducing the influence of transient or spurious observations. Executing the audit runs the tracker, geometric ranging, ego-vehicle dynamics, and event-detection stages across all 1,800 frames.

In [26]:
from driveaudit.pipeline import run_audit
res = run_audit(CACHE / "synthetic-drive.mp4", det, calib, rules, OUT,
                stride=1, speed_log=str(CACHE / "speed_log.csv"),
                write_video=False, log=lambda *a: None)   # Video pre-rendered separately.
m, events, tracks = res["metrics"], res["events"], res["tracks"]
print(f"Processed {len(m):,} frames in {res['wall_s']:.0f} s -> {len(events)} candidate events.")
print("\nTrack lifespans:")
for tid, g in tracks.groupby("track_id"):
    print(f"  Track #{tid}: Frames {g.frame.min():>4}-{g.frame.max():<4}  Range {g.range_m.min():.0f}-{g.range_m.max():.0f} m")
Processed 1,800 frames in 4 s -> 4 candidate events.

Track lifespans:
  Track #1: Frames    0-127   Range 7-58 m
  Track #2: Frames    0-1799  Range 18-66 m
  Track #3: Frames   92-277   Range 20-60 m

5. Geometry and Ground Plane¶

A monocular camera cannot recover metric depth without an additional geometric constraint. Here, the road provides that constraint by being modeled as a planar surface at a known height. The pinhole-camera relations below use the vehicle’s tyre-contact point to estimate its metric distance and lateral offset relative to the ego vehicle. The figure compares the recovered lead-vehicle distance with the known ground-truth distance over the full drive. Because the synthetic scene is generated from known geometry, this comparison provides a direct quantitative evaluation of the ranging model.

In [27]:
Z = calib.focal_px * calib.camera_height_m / (500 - calib.horizon_row)
print(f"  A box bottom at row 500 -> {Z:.1f} m ahead.")
print(f"  Z = focal ({calib.focal_px:.0f}) * height ({calib.camera_height_m}) / (row - horizon ({calib.horizon_row:.0f}))")

from scenarios.evaluate import fig_distance
gt = pd.read_csv(CACHE / "ground_truth.csv")
geo = fig_distance(m, gt, calib, OUT / "distance_accuracy.png")
print("\n", geo)
display(Image(str(OUT / "distance_accuracy.png"), width=760))
  A box bottom at row 500 -> 7.6 m ahead.
  Z = focal (1000) * height (1.3) / (row - horizon (330))

 {'distance_mae_m': 0.945, 'distance_rmse_m': 1.7, 'n_compared': 1750, 'max_true_range_m': 60.0}
No description has been provided for this image

6. Ego Dynamics and Time Headway¶

Distance alone is not a sufficient safety measure. Time headway, defined as the distance to the lead vehicle divided by ego-vehicle speed, provides a more meaningful measure because the same physical gap represents different levels of risk at different speeds. For example, a two-metre gap has very different implications at 30 km/h and 120 km/h.

For this demonstration, ego-vehicle speed is obtained from a supplied GPS-style trajectory log. In a deployment without an external speed signal, the pipeline can instead estimate ego speed from motion observed in the bird’s-eye-view road representation. The figures show the ego-speed profile with the hard-braking event marked, followed by the time-headway trace with detected event windows highlighted.

In [28]:
from scenarios.evaluate import fig_speed, fig_headway
fig_speed(m, events, OUT / "ego_speed.png")
fig_headway(m, events, rules, OUT / "headway_timeline.png")
hw = m["headway_s"].dropna()
print(f"Median headway while following: {hw.median():.1f} s")
print(f"Share of following time below {rules.headway_s:.0f} s: {(hw < rules.headway_s).mean():.0%}")
display(Image(str(OUT / "ego_speed.png"), width=760))
display(Image(str(OUT / "headway_timeline.png"), width=760))
Median headway while following: 4.0 s
Share of following time below 2 s: 12%
No description has been provided for this image
No description has been provided for this image

7. Event Detection¶

Each event rule is defined by a threshold on a physical quantity together with a persistence requirement. This prevents a single noisy frame from triggering an event and merges same-type events separated by only a brief interruption. Each detected event records the peak value that triggered the rule, providing a quantitative basis for subsequent annotation and validation.

In [29]:
show = events[["event_id", "type", "start_s", "end_s", "peak_value", "unit", "track_id"]].copy()
show["start_s"] = show["start_s"].round(1); show["end_s"] = show["end_s"].round(1)
show
Out[29]:
event_id type start_s end_s peak_value unit track_id
0 E001 tailgating 5.9 17.4 1.164654 s headway 2.0
1 E002 cut_in 10.0 10.0 21.435535 m range 3.0
2 E003 hard_braking 18.1 19.6 -9.166667 m/s2 NaN
3 E004 rolling_stop 41.3 44.0 14.000000 m/s minimum speed NaN

8. Safety Scoring¶

This is the final aggregation step: the preceding stages provide the measurements and detected events on which the score is based. Each event type is assigned a penalty weight. The weighted penalties are aggregated, normalized to a rate per ten minutes, and mapped to a 0–100 score using an exponential decay function. This produces a gradual reduction in score rather than forcing a disproportionately low score for a short clip containing several events.

The resulting score is a heuristic safety index, not a calibrated measure of accident risk. The underlying event table remains the primary evidence, while the score provides a compact summary of those observations. For this demonstration, the reported score is calculated from the unconfirmed event candidates.

In [30]:
from driveaudit.events import safety_score
score = safety_score(events, res["info"]["duration_s"], rules)
print(json.dumps(score, indent=2))
{
  "score": 45.6,
  "penalty_points": 34,
  "counts": {
    "tailgating": 1,
    "cut_in": 1,
    "hard_braking": 1,
    "rolling_stop": 1
  },
  "penalty_per_10min": 170.0,
  "exposure_minutes": 2.0
}

9. Annotated Rendering¶

The pipeline generates an annotated video containing per-vehicle bounding boxes, track IDs, estimated ranges, a heads-up display of ego speed and time headway, and event banners. Calibration-derived ego-lane guides are also projected into the image to provide spatial context. The figure shows a representative frame from the detected tailgating interval. The complete annotated video is available at output/driveaudit-annotated.mp4.

In [31]:
display(Image(str(OUT / "annotated_frame.png"), width=820))
No description has been provided for this image

10. Evaluation Against Ground Truth¶

Because the scene is generated from known trajectories, three aspects of the pipeline can be evaluated quantitatively rather than inferred from qualitative inspection.

  • Distance error: estimated lead-vehicle range is compared directly with the known ground-truth distance.
  • Event accuracy: precision and recall are measured against the same event rules applied to the ground-truth signals. This isolates errors introduced by tracking and geometric estimation from errors attributable to the event definitions themselves.
  • Track stability: identity switches are counted using global frame-by-frame matching between the estimated tracks and ground-truth identities. These measurements provide a direct assessment of the pipeline components that can be evaluated under the synthetic ground-truth setup.
In [32]:
from scenarios.evaluate import build_gt_metrics, match_events, id_switches
from driveaudit.events import detect_events
speed_log = pd.read_csv(CACHE / "speed_log.csv")
gt_metrics, gt_lc = build_gt_metrics(gt, speed_log, calib, int(info["fps"]), rules)
gt_events = detect_events(gt_metrics, gt_lc, int(info["fps"]), rules)
by_type, overall = match_events(events, gt_events)

print(f"Distance:  MAE {geo['distance_mae_m']} m   RMSE {geo['distance_rmse_m']} m  (out to {geo['max_true_range_m']} m)")
print(f"Tracking:  {id_switches(tracks, gt)} id switches over {gt[gt['cls'].isin([2,3,5,7])]['actor_id'].nunique()} vehicles")
print(f"Events:    Precision {overall['precision']}   Recall {overall['recall']}")
df = pd.DataFrame(by_type).T
df[["detected", "reference", "true_positive"]] = df[["detected", "reference", "true_positive"]].astype(int)
df
Distance:  MAE 0.945 m   RMSE 1.7 m  (out to 60.0 m)
Tracking:  0 id switches over 3 vehicles
Events:    Precision 1.0   Recall 1.0
Out[32]:
detected reference true_positive precision recall
cut_in 1 1 1 1.0 1.0
hard_braking 1 1 1 1.0 1.0
rolling_stop 1 1 1 1.0 1.0
tailgating 1 1 1 1.0 1.0

11. Summary¶

A machine-readable summary of the completed run, written to output/summary.json.

In [33]:
summary = {
    "clip": {"name": "synthetic-drive.mp4", "duration_s": round(res["info"]["duration_s"], 1),
             "fps": int(info["fps"]), "frames": int(len(m)), "resolution": [calib.frame_width, calib.frame_height]},
    "detection": {"source": "scripted ground-truth stream (detector bypassed for the demo)",
                  "production_detector": "YOLOv8n exported to ONNX, CPU inference"},
    "geometry": geo,
    "tracking": {"id_switches": id_switches(tracks, gt),
                 "distinct_tracks": int(tracks["track_id"].nunique())},
    "events": {"detected": {k: int(v) for k, v in events["type"].value_counts().items()},
               "precision": overall["precision"], "recall": overall["recall"], "by_type": by_type},
    "score": score,
    "speed_source": res["speed_source"],
}
(OUT / "summary.json").write_text(json.dumps(summary, indent=2))
print(json.dumps(summary, indent=2))
{
  "clip": {
    "name": "synthetic-drive.mp4",
    "duration_s": 120.0,
    "fps": 15,
    "frames": 1800,
    "resolution": [
      1280,
      720
    ]
  },
  "detection": {
    "source": "scripted ground-truth stream (detector bypassed for the demo)",
    "production_detector": "YOLOv8n exported to ONNX, CPU inference"
  },
  "geometry": {
    "distance_mae_m": 0.945,
    "distance_rmse_m": 1.7,
    "n_compared": 1750,
    "max_true_range_m": 60.0
  },
  "tracking": {
    "id_switches": 0,
    "distinct_tracks": 3
  },
  "events": {
    "detected": {
      "tailgating": 1,
      "cut_in": 1,
      "hard_braking": 1,
      "rolling_stop": 1
    },
    "precision": 1.0,
    "recall": 1.0,
    "by_type": {
      "cut_in": {
        "detected": 1,
        "reference": 1,
        "true_positive": 1,
        "precision": 1.0,
        "recall": 1.0
      },
      "hard_braking": {
        "detected": 1,
        "reference": 1,
        "true_positive": 1,
        "precision": 1.0,
        "recall": 1.0
      },
      "rolling_stop": {
        "detected": 1,
        "reference": 1,
        "true_positive": 1,
        "precision": 1.0,
        "recall": 1.0
      },
      "tailgating": {
        "detected": 1,
        "reference": 1,
        "true_positive": 1,
        "precision": 1.0,
        "recall": 1.0
      }
    }
  },
  "score": {
    "score": 45.6,
    "penalty_points": 34,
    "counts": {
      "tailgating": 1,
      "cut_in": 1,
      "hard_braking": 1,
      "rolling_stop": 1
    },
    "penalty_per_10min": 170.0,
    "exposure_minutes": 2.0
  },
  "speed_source": "gps"
}

12. Limitations¶

  • The detector is not evaluated in this demonstration. The synthetic run bypasses YOLO using a scripted detection stream. Consequently, these results validate the tracking, geometry, event-detection, and scoring stages, but not the detector itself. On real footage, detector precision and recall will constrain the performance of the complete pipeline. Evaluation against a labelled dataset such as BDD100K is therefore the first additional validation step.
  • Range estimation assumes a flat road and fixed camera mount. Hills, road crown, speed bumps, and camera pitch caused by braking violate the flat-road model. Range error also increases with distance because small image-row errors near the horizon can correspond to large changes in estimated distance. The sub-metre MAE observed here represents a best-case result on level synthetic terrain.
  • Ego-speed estimates depend on the input source. When a GPS-derived speed signal is available, the pipeline uses that signal directly. The bird's-eye-view motion estimator provides a smoothed fallback, but its accuracy can degrade on roads with limited visual texture and during stop-and-go traffic.
  • The safety score is a heuristic, not a calibrated risk model. The event weights and exponential decay function are engineering choices rather than parameters calibrated against insurance claims, collisions, or other outcome data. The score is intended to rank driving sessions consistently; it does not represent a probability of collision or other quantified risk. The event table remains the auditable artifact.
  • Detected events are candidates until confirmed. The event rules are designed to balance false positives and recall and therefore expose an annotator-confirmation step. A production system should report the score derived from confirmed events rather than treating every detected candidate as a validated event.