Choose an evaluation workflow
The SWE-bench guide evaluates supplied patches. It doesn’t generate them or call a model. The DeepSWE and AutomationBench guides include agent execution and model requests as part of the evaluation.
Separate model inference from evaluation execution
The model provider generates responses. The agent harness uses those responses to select commands to run. CoreWeave Sandbox executes the commands, and the benchmark verifier determines whether the result satisfies the task. Model selection and inference hosting depend on your harness. The DeepSWE and AutomationBench guides use Serverless Inference as the example provider, which isn’t required to run evaluations on CoreWeave Sandbox. The AutomationBench recipe accepts another OpenAI-compatible model endpoint through its.env settings. The DeepSWE recipe includes provider-specific configuration, so changing providers there requires adapting its model connection and authentication.
Where the agent loop runs also depends on the integration. The DeepSWE recipe runs it on the machine that launches the evaluation and executes its shell actions in remote sandboxes. AutomationBench runs Harbor locally and its agent loop inside each trial sandbox. See Agents for integrations that run agent processes inside a sandbox or connect a managed agent service.