You can detect triggered backdoors in LLMs for free by monitoring an existing inference optimization (speculative decoding), without adding computation or making assumptions about trigger types.
SpecGuard detects hidden backdoors in large language models during inference by monitoring speculative decoding—a speed optimization technique. When a backdoor is triggered, the target model's behavior shifts while the draft model doesn't, causing token acceptance rates to change. This detection happens automatically without slowing down inference.