25 upvotes · 30 JUL 2026 · Tengfei Liu, Yang Shi, Yuran Wang et al.
This paper introduces a new task called multi-reference image-grounded video captioning, where models must describe video content while referencing multiple images. Practitioners might care because this can improve the accuracy and faithfulness of video captions in real-world applications.
11 upvotes · 31 JUL 2026 · Yingmao Miao, Pengfei Zhang, Xiaochen Lv et al.
This paper develops a new method to evaluate and verify image editing consistency across multiple references, addressing a challenge in reinforcement learning for multi-reference editing. Practitioners may care about this approach as it enables more accurate and reliable reinforcement learning for image editing tasks.