Applying RL optimization in a compressed learned representation rather than the full action space to improve efficiency and safety.