Training a model on multiple types of data (text and images) simultaneously to learn shared representations.