Open-source models can match much larger models at multi-step API calling by training on execution-verified data—showing that smart data synthesis matters more than model size for tool-use tasks.
This paper addresses a real-world problem: open-source AI models struggle when they need to call multiple government APIs in sequence to complete tasks. The authors create KOPA-Bench, a benchmark of 145 real Korean government API tasks, and introduce EDGE, a method that learns which API outputs can feed into other APIs by actually testing them live.