The process of identifying and labeling different speakers in an audio or text transcript, answering the question 'who said what'.