Comparing Monocular and Multi-View 3D Pedestrian Reconstruction Strategies for Urban Intersection Sensing (In Print)
Overview of the two compared strategies for deriving pedestrian motion descriptors: Pipeline~A (monocular learned, left) and Pipeline~B (multi-view geometric, right).
Publication Details
- Venue
- Graphics, Patterns and Images (SIBGRAPI)
- Year
- 2026
- Publication Date
- October 2, 2026
Materials
Abstract
Understanding how pedestrians move and interact at urban intersections is central to traffic-safety assessment, mobility planning, and smart-city sensing. Such analyses rely on physically grounded descriptions of motion—3D body pose, metric trajectories, and pose-derived descriptors such as stature and walking speed. Recovering these from fixed roadside cameras is difficult: viewpoints are oblique and long-range, occlusions are frequent, and neither ground-truth 3D pose nor cross-camera identity labels are typically available. We compare two strategies for deriving pedestrian motion descriptors from a multi-camera intersection dataset: a monocular learned pipeline that combines text-prompted segmentation, tracking, and learned 3D body estimation, and a calibrated multi-view geometric pipeline that combines 2D pose estimation, tracking, cross-camera association, and triangulation. Because 3D ground truth is unavailable, we evaluate both with reference-free criteria relevant to intersection analysis: pedestrian coverage, anthropometric plausibility, walking-speed distributions, computational cost, and failure modes. The two strategies exhibit complementary trade-offs. The learned pipeline is simpler to deploy and covers more pedestrians, but suffers from scale ambiguity and systematic height compression; the geometric pipeline yields metric reconstructions more efficiently, but requires calibration, discards pedestrians seen by only one camera, and can produce severe outliers in dense crowds. We distill these findings into practical guidance for selecting a reconstruction strategy and configuring camera infrastructure according to the target analysis task.
Cite this publication (BIBTEX)
@article{2026-MonoVsMulti3D,
title={Comparing Monocular and Multi-View 3D Pedestrian Reconstruction Strategies for Urban Intersection Sensing (In Print)},
author={Kauan Kauan Mariani Ferreira and Matheus Fillype Ferreira de Carvalho and Joel Perca and João Rulff and Jorge Poco},
journal={Graphics, Patterns and Images (SIBGRAPI) },
year={2026},
url={null},
date={2026-10-02}
}