Retrieval can guide reinforcement learning to improve vision-language models for image captioning by helping identify and correct errors, outperforming standard supervised fine-tuning approaches.
This paper proposes Re³Cap, a method that uses retrieval-guided reasoning to improve image captioning with reinforcement learning. Instead of just fine-tuning models, it retrieves similar images and captions to help identify and fix errors (hallucinations and omissions) in generated descriptions, achieving better results than supervised fine-tuning without needing extra labeled data.