DAPO (Direct Alignment from Preference Optimization) — Glossary — ThinkLLM
DAPO (Direct Alignment from Preference Optimization)
training
A fine-tuning technique that aligns a model's behavior with human preferences by directly optimizing based on preference data rather than using reinforcement learning.