Automated LLM auditing can be made dramatically more efficient by combining adaptive questioning strategies with logit-based output reweighting, finding harmful behaviors 2x more often than baseline methods without requiring model retraining.
BLOOM-WILT is an automated auditing system that efficiently finds rare problematic behaviors in deployed language models. It uses two key techniques: an auditor that learns better questioning strategies across conversations, and logit tilting that reweights the model's output distribution to surface behavior-relevant responses.