Software that executes a trained model on a device, handling the low-level operations needed to run inference efficiently.