Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
This paper improves monocular depth estimation models by repurposing image generation models, using a diffusion transformer architecture, to produce sharper and more detailed depth maps that generalize well to out-of-distribution inputs. Practitioners might care about this research because it could lead to better performance in applications such as scene reconstruction and computational photography.