RefCaptioner: Multi-Reference Image-Grounded Video Captioning
This paper introduces a new task called multi-reference image-grounded video captioning, where models must describe video content while referencing multiple images. Practitioners might care because this can improve the accuracy and faithfulness of video captions in real-world applications.