Hallucination detectors trained on general text fail dramatically on specialized domains—use domain-matched models like PubMedBERT for biomedical content, and combine classification with uncertainty quantification for best results.
This paper builds a hallucination detector for LLMs using a three-part system: a fine-tuned DeBERTa classifier, uncertainty estimation via MC Dropout, and temperature calibration. Tested on HaluEval, it achieves 91.5% F1 on general tasks but struggles cross-domain (52% F1 on biomedical text), showing that domain-specific pre-training is essential for reliable hallucination detection.