Road Scene Perception
Road scene perception converts traffic-environment sensor data into spatial understanding: drivable surface, lanes, vehicles, riders, pedestrians, signs, traffic lights, human keypoints, orientation, and tracks. Inputs are camera frames, calibration, timestamps, ego-motion, optional lidar or radar, and scenario metadata. Targets are pixel classes, instance masks, boxes, keypoints, and tracks used by planning, mapping, safety review, or driver assistance.
Framing
The segmentation side is semantic segmentation; the actor and vulnerable-road-user side combines object detection, pose estimation, and person tracking and track aggregation. Evaluation should include mIoU, boundary quality, small-object recall, keypoint error, latency, tracking continuity, and scenario slices such as rain, night, glare, occlusion, and new cities. Cityscapes is the canonical street-scene segmentation artifact: the paper reports stereo video from 50 cities, 5,000 finely annotated images, and 20,000 additional coarse annotations.
Worked Segmentation Check
This 4x4 mask example computes IoU for road, car, and person classes:
| class | intersection pixels | union pixels | IoU |
|---|---|---|---|
| road | 6 | 6 | 1.000 |
| car | 6 | 7 | 0.857 |
| person | 3 | 4 | 0.750 |
The mean IoU is . The mean looks high, but person IoU is the weakest class. For safety-sensitive perception, detection and segmentation metrics should be inspected per class and per scenario, not averaged away.
Failure Modes
Road-scene models fail under domain shift: new city geometry, rare signs, weather, camera exposure, construction layouts, and unusual pedestrian poses. Pose detection also fails when bodies are truncated or occluded by vehicles. Use targeted test sets for vulnerable road users and keep uncertain outputs from silently feeding a planner.
References
Nav