What this does
SWE-bench tests whether language models can fix real bugs in real repositories. Each task applies a patch, runs the test suite, and checks if the fix works. You can run SWE-bench locally with Docker, but your machine’s resources limit you. CoreWeave sandboxes let you run evaluations at scale on CoreWeave infrastructure. Spin up dozens or hundreds of sandboxes concurrently without managing any of it yourself. The script pulls pre-built images from Epoch AI’s registry on GHCR, so no local Docker build step is needed.Setup
Complete the following steps to install the required tools and configure authentication before running any evaluations.-
Python SDK와 평가 의존성을 설치하세요.
-
Clone the repo for the example evaluation scripts:
-
인증을 위해
WANDB_API_KEY를 W&B API 키로 설정하세요.[WANDB-API-KEY]를 본인의 키로 바꾸세요. -
examples/swebench/run_evaluation.py를 열고SandboxDefaults(...)호출에auth=AuthStrategy.WANDB를 추가한 다음,cwsandbox에서AuthStrategy를 임포트하세요. 이 스크립트는 W&B 인증을 자동으로 선택하지 않습니다.
Quick start
Before you scale up, start with a single-instance run to confirm your environment, credentials, and network access to GHCR all work. Thegold option uses the known-correct fix from the dataset:
Run in parallel
Run multiple instances at once:Evaluate model predictions
To test custom model output:Options
Adjust resources
The default is 2 CPUs and 4Gi memory per sandbox. Change this inrun_evaluation.py:
How it works
This section explains the moving parts behind the script so you can adapt it to your own evaluation workflows.Container images
Epoch AI hosts pre-built images on GHCR. Each instance has its own image with the repo checked out at the right commit, dependencies installed, and test environment ready. Image format:ghcr.io/epoch-research/swe-bench.eval.x86_64.{instance_id}:latest
Sandboxes pull these directly from GHCR. No local builds are needed.
Evaluation flow
Each instance goes through these steps:
Steps 2 to 5 and 8 run remotely. Steps 1, 6, and 7 run on your machine.
Parallel execution
The script usesThreadPoolExecutor to run instance workflows concurrently. Each thread drives one instance through its workflow. While one instance runs tests, another can apply its patch, another grade locally, another start up. The overlap is where the speed comes from.
Results come back as workflows finish through as_completed().
Cleanup
이 스크립트는 샌드박스Session으로 샌드박스를 추적합니다. 세션이 종료되면(정상 종료, 예외 발생 또는 Ctrl+C) 세션이 모든 샌드박스를 정리합니다.
Output
Results go to{output-dir}/{run-id}/{model-name}/{instance-id}/:
Report format
Troubleshooting
Patch application fails
If you seeAPPLY_PATCH_FAIL in logs, the patch is probably malformed, targets the wrong commit, or has whitespace issues. Run git apply --check locally to see what’s wrong. Make sure the instance ID matches the prediction.
Test timeouts
Some test suites take longer than 30 minutes. Model-generated code might also have infinite loops. If needed, increase--timeout, or check the test output to see where it’s stuck.
Image pull errors
If the container fails to start with an image pull error, either the instance ID doesn’t exist in Epoch AI’s registry, or a network issue prevents reachingghcr.io. Verify the instance ID is in the SWE-bench dataset.
See also
- Command execution guide: details about
exec()patterns. - Cleanup patterns: managing sandbox lifecycle.