Dense Prediction

Dense Prediction

Dense Prediction at CVAIL combines learning-based perception with physically informed reasoning so models can operate under real-world constraints instead of idealized settings. Our work on dense prediction covers segmentation, depth, correspondence, and structured outputs that help machines reason about entire images rather than isolated objects.

Publications

Beyond Appearances: Material Segmentation with Embedded Spectral Information from RGB-D imagery

Abstract

In the realm of computer vision, material segmentation of natural scenes represents a challenge, driven by the complex and diverse appearances of materials. Traditional approaches often rely on RGB images, which can be deceptive given the variability in appearances due to different lighting conditions. Other methods that employ polarization or spectral imagery offer more reliable material differentiation, but their cost and accessibility restrict everyday usage. This work proposes a deep learning framework that uses paired RGB-D and spectral data during training to embed spectral information through a Spectral Feature Mapper (SFM) layer, enabling material segmentation from standard RGB-D images after training. The method also generates a 3D point cloud from the RGB-D pair to enrich scene understanding, and experiments on public datasets plus real captures from an iPad Pro show superior material segmentation performance.

CVPR LatinX Workshop

2024
SLiDE: Stereo-LiDAR Depth Estimation with Domain-Aware Attention

Abstract

Accurate depth completion is crucial for safety-critical applications such as autonomous navigation. This paper introduces a stereo-LiDAR fusion network with a domain-aware attention mechanism that jointly reasons in depth and disparity space instead of relying on only one representation. The method uses sparse but precise LiDAR measurements, propagates them according to disparity density, and predicts dense disparity and depth-residual maps that correct stereo triangulation errors and disparity-to-depth conversion inaccuracies. On the KITTI validation set, the proposed approach achieves state-of-the-art Mean Absolute Error with rapid convergence.

ColCACI

2025
UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields

Abstract

NeRF-based segmentation methods usually focus on semantics and rely only on RGB data, which limits material understanding. UnMix-NeRF integrates spectral unmixing into neural radiance fields to jointly perform hyperspectral novel view synthesis and unsupervised material segmentation. The model represents spectral reflectance with diffuse and specular components, learns a global dictionary of endmembers as pure material signatures, and estimates per-point abundances to capture their distribution across the scene. These learned spectral signatures also support unsupervised material clustering and material-aware scene editing, while experiments show stronger spectral reconstruction and material segmentation than prior methods.

ICCV

2025

Awards

Winner of SoccerNet Monocular Depth Estimation Challenge 2025

Abstract

The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understanding. This year's challenges span four vision-based tasks: (1) Team Ball Action Spotting, focused on detecting ball-related actions in football broadcasts and assigning actions to teams; (2) Monocular Depth Estimation, targeting the recovery of scene geometry from single-camera broadcast clips through relative depth estimation for each pixel; (3) Multi-View Foul Recognition, requiring the analysis of multiple synchronized camera views to classify fouls and their severity; and (4) Game State Reconstruction, aimed at localizing and identifying all players from a broadcast video to reconstruct the game state on a 2D top-view of the field. Across all tasks, participants were provided with large-scale annotated datasets, unified evaluation protocols, and strong baselines as starting points. This report presents the results of each challenge, highlights the top-performing solutions, and provides insights into the progress made by the community. The SoccerNet Challenges continue to serve as a driving force for reproducible, open research at the intersection of computer vision, artificial intelligence, and sports.

CVPR

2025
Back to Research
© 2026 CVAIL. All Rights Reserved
Policies and Terms
Powered by CVAIL Research Group