A weakness in how a model's internal safety mechanisms work that can be exploited through targeted attacks.