Standard semantic similarity metrics are unreliable for legal text simplification because they can't distinguish between preserving words and preserving legal meaning—a problem that requires new evaluation approaches beyond token overlap.
This paper exposes a critical flaw in how we measure whether simplified legal text preserves meaning. Current metrics fail because they conflate surface-level word overlap with actual legal meaning.