Finding similar scenes across different input types (e.g., finding a visual scene matching an audio description).