From 2D Image Planes to 3D Occupancy Grids
Traditional computer vision pipelines relied heavily on 2D bounding boxes to identify cars, pedestrians, and cyclists. However, real-world autonomous driving requires understanding general obstacles regardless of category, such as fallen tree branches, construction debris, or spilled cargo.
Vector Space Transformation
Modern visual neural networks stream multi-camera video into a unified spatial-temporal Transformer. By projecting image features into a 3D voxel space, the onboard AI generates an Occupancy Grid—a volumetric map denoting whether any coordinate in 3D space is free or blocked.