> ## Documentation Index
> Fetch the complete documentation index at: https://docs.coreweave.com/llms.txt
> Use this file to discover all available pages before exploring further.

# SWE-bench

> CoreWeave のサンドボックスを使用して SWE-bench 評価を実行します。

Run SWE-bench evaluations in parallel using sandboxes on CoreWeave. This guide is for ML practitioners and researchers who want to evaluate language model code-fixing performance at scale. By the end, you have a working setup that runs SWE-bench evaluations concurrently across many CoreWeave sandboxes, with results written locally for analysis.

## What this does

[SWE-bench](https://www.swebench.com/) tests whether language models can fix real bugs in real repositories. Each task applies a patch, runs the test suite, and checks if the fix works.

You can run SWE-bench locally with Docker, but your machine's resources limit you. CoreWeave sandboxes let you run evaluations at scale on CoreWeave infrastructure. Spin up dozens or hundreds of sandboxes concurrently without managing any of it yourself. The script pulls pre-built images from Epoch AI's registry on GHCR, so no local Docker build step is needed.

## Setup

Complete the following steps to install the required tools and configure authentication before running any evaluations.

1. Python SDK と評価用の依存関係をインストールします。

   ```bash theme={"system"}
   uv pip install 'cwsandbox[wandb]' swebench datasets
   ```

1) Clone the repo for the example evaluation scripts:

   ```bash theme={"system"}
   git clone https://github.com/coreweave/cwsandbox-client.git
   cd cwsandbox-client
   ```

1. 認証に使用するため、`WANDB_API_KEY` に [W\&B APIキー](/ja/products/sandboxes/serverless/get-started) を設定します。`[WANDB-API-KEY]` はご自身のキーに置き換えてください。

   ```bash theme={"system"}
   export WANDB_API_KEY="[WANDB-API-KEY]"
   ```

2. `examples/swebench/run_evaluation.py` を開き、`SandboxDefaults(...)` の呼び出しに `auth=AuthStrategy.WANDB` を追加したうえで、`cwsandbox` から `AuthStrategy` をインポートします。このスクリプトは W\&B 認証を自動的には選択しないためです。

## Quick start

Before you scale up, start with a single-instance run to confirm your environment, credentials, and network access to GHCR all work. The `gold` option uses the known-correct fix from the dataset:

```bash theme={"system"}
uv run python examples/swebench/run_evaluation.py \
    --predictions-path gold \
    --instance-ids astropy__astropy-12907 \
    --run-id test
```

This run passes, confirming that your sandbox can pull the image, apply the patch, and run the test suite end to end.

### Run in parallel

Run multiple instances at once:

```bash theme={"system"}
uv run python examples/swebench/run_evaluation.py \
    --predictions-path gold \
    --instance-ids \
        astropy__astropy-12907 \
        django__django-11039 \
        django__django-11099 \
        django__django-11283 \
        matplotlib__matplotlib-23476 \
        scikit-learn__scikit-learn-13142 \
        sympy__sympy-13031 \
        sympy__sympy-13647 \
    --run-id parallel-test \
    --max-workers 8
```

This command spins up eight sandboxes and runs them concurrently. All pass because gold patches are the correct fixes.

### Evaluate model predictions

To test custom model output:

```bash theme={"system"}
uv run python examples/swebench/run_evaluation.py \
    --predictions-path predictions.json \
    --instance-ids django__django-11039 scikit-learn__scikit-learn-13142 \
    --run-id eval-run-1 \
    --max-workers 10
```

The predictions file maps instance IDs to patches:

```json theme={"system"}
[
  {
    "instance_id": "django__django-11039",
    "model_name_or_path": "gpt-4",
    "model_patch": "diff --git a/..."
  }
]
```

## Options

| Option | Default | Description |
| - | - | - |
| `--predictions-path` | Required | Path to predictions JSON, or `gold` for gold patches |
| `--instance-ids` | Required | Space-separated instance IDs |
| `--run-id` | Required | Identifier for this run |
| `--max-workers` | 10 | Max parallel sandboxes |
| `--timeout` | 1800 | Per-instance timeout in seconds (30 minutes) |
| `--output-dir` | `logs/swebench` | Where to write logs and reports |
| `--dataset` | `princeton-nlp/SWE-bench_Lite` | Hugging Face dataset name |
| `--force` | false | Re-run instances even if report.json exists |

### Adjust resources

The default is 2 CPUs and 4Gi memory per sandbox. Change this in `run_evaluation.py`:

```python theme={"system"}
defaults = SandboxDefaults(
    tags=(f"swebench-{run_id}",),
    resources={"cpu": "4", "memory": "8Gi"},
)
```

## How it works

This section explains the moving parts behind the script so you can adapt it to your own evaluation workflows.

### Container images

Epoch AI hosts pre-built images on GHCR. Each instance has its own image with the repo checked out at the right commit, dependencies installed, and test environment ready.

Image format: `ghcr.io/epoch-research/swe-bench.eval.x86_64.{instance_id}:latest`

Sandboxes pull these directly from GHCR. No local builds are needed.

### Evaluation flow

Each instance goes through these steps:

| ステップ | 実行場所 | 処理内容 |
| - | - | - |
| 1. データセットの読み込み | ローカル | Hugging Face からインスタンスのメタデータを取得します |
| 2. サンドボックスの作成 | サンドボックス | インスタンスのイメージを使用してサンドボックスを起動します |
| 3. パッチの書き込み | サンドボックス | パッチを `/tmp/patch.diff` に書き込みます |
| 4. パッチの適用 | サンドボックス | `git apply` を実行します (必要に応じて `patch` にフォールバックします) |
| 5. テストの実行 | サンドボックス | `/root/eval.sh` を実行します |
| 6. 結果の採点 | ローカル | `swebench.harness.grading` で出力を解析します |
| 7. レポートの書き出し | ローカル | 結果を `logs/swebench/` に保存します |
| 8. クリーンアップ | サンドボックス | サンドボックスを停止します |

Steps 2 to 5 and 8 run remotely. Steps 1, 6, and 7 run on your machine.

### Parallel execution

The script uses `ThreadPoolExecutor` to run instance workflows concurrently. Each thread drives one instance through its workflow. While one instance runs tests, another can apply its patch, another grade locally, another start up. The overlap is where the speed comes from.

Results come back as workflows finish through `as_completed()`.

### Cleanup

このスクリプトは、サンドボックスの `Session` を使用してサンドボックスをトラッキングします。セッションが終了すると (正常終了、例外発生、`Ctrl+C` のいずれの場合も) 、そのセッションがすべてのサンドボックスをクリーンアップします。

```python theme={"system"}
with cwsandbox.Session(defaults=defaults) as session:
    with ThreadPoolExecutor(max_workers=max_workers) as executor:
        # The session cleans up sandboxes created here when it exits
```

## Output

Results go to `{output-dir}/{run-id}/{model-name}/{instance-id}/`:

| File | Contents |
| - | - |
| `report.json` | Resolved status, sandbox ID, duration |
| `test_output.txt` | Full test output |
| `patch.diff` | The applied patch |

### Report format

```json theme={"system"}
{
  "instance_id": {
    "resolved": true,
    "tests_status": {
      "PASSED": ["test_foo", "test_bar"],
      "FAILED": []
    }
  },
  "sandbox_id": "sb-abc123",
  "duration_seconds": 45.2
}
```

The format is compatible with standard SWE-bench tooling.

## Troubleshooting

### Patch application fails

If you see `APPLY_PATCH_FAIL` in logs, the patch is probably malformed, targets the wrong commit, or has whitespace issues. Run `git apply --check` locally to see what's wrong. Make sure the instance ID matches the prediction.

### Test timeouts

Some test suites take longer than 30 minutes. Model-generated code might also have infinite loops. If needed, increase `--timeout`, or check the test output to see where it's stuck.

### Image pull errors

If the container fails to start with an image pull error, either the instance ID doesn't exist in Epoch AI's registry, or a network issue prevents reaching `ghcr.io`. Verify the instance ID is in the SWE-bench dataset.

## See also

* [Command execution guide](../../client/guides/execution): details about `exec()` patterns.
* [Cleanup patterns](../../client/guides/cleanup-patterns): managing sandbox lifecycle.
