Training a model using aligned labels from different data sources or modalities (e.g., video and biomechanical data).