Hi Lingbot-Depth team,
First of all, thank you for releasing this great work. The results of Lingbot-Depth are very impressive, especially in challenging indoor scenes with transparent glass, glass doors, or reflective surfaces.
I noticed that the predicted depth around transparent glass regions looks surprisingly robust. This is a very difficult hard case for RGB-D depth completion, since the raw depth from sensors can be missing, noisy, or even penetrate through the glass and hit the background instead.
I would like to better understand what mainly contributes to this strong performance on transparent glass scenes:
- Is the improvement mainly due to the Lingbot-Vision backbone and its stronger visual/structural representation?
- Or did the training data include specific transparent-glass / glass-door / reflective-surface scenes?
- If special data was used, was it collected with dense ground-truth depth for glass surfaces, or was it handled through pseudo labels, synthetic data, or other supervision strategies?
- Are there any recommended practices for adapting Lingbot-Depth to transparent glass hard cases in custom indoor robot scenarios?
I am particularly interested in whether the model learns this capability mostly from the general-purpose vision backbone, or whether transparent/glass-specific data is still necessary for robust performance.
Thanks again for the excellent work!
Hi Lingbot-Depth team,
First of all, thank you for releasing this great work. The results of Lingbot-Depth are very impressive, especially in challenging indoor scenes with transparent glass, glass doors, or reflective surfaces.
I noticed that the predicted depth around transparent glass regions looks surprisingly robust. This is a very difficult hard case for RGB-D depth completion, since the raw depth from sensors can be missing, noisy, or even penetrate through the glass and hit the background instead.
I would like to better understand what mainly contributes to this strong performance on transparent glass scenes:
I am particularly interested in whether the model learns this capability mostly from the general-purpose vision backbone, or whether transparent/glass-specific data is still necessary for robust performance.
Thanks again for the excellent work!