A machine learning approach that processes both audio and visual information together to better understand speech and communication.