The task of learning to match and relate images with their corresponding text descriptions or captions.
Quality of vision, audio, and image understanding (distinct from modality support)