Evaluation workflow
The AutomationBench recipe runs Harbor on your local machine. Each trial gets a CPU sandbox where AutomationBench installs its agent loop, simulated applications, and assertion-based evaluator. The model runs at your inference endpoint. No GPU or real application credentials are required in the sandbox. The recipe implements Harbor’sBaseAgent and BaseEnvironment interfaces and uses public CoreWeave Sandbox software development kit (SDK) calls. It runs the unchanged native auto-bench evaluator and doesn’t require an upstream AutomationBench Harbor adapter. Harbor, AutomationBench, and cwsandbox SDK versions are pinned in the recipe.
Prerequisites
Before you begin, prepare the following:uvand Git on your local machine. The recipe requires Python 3.13 or later. If it isn’t installed,uvdownloads a matching Python version.- A CoreWeave API access token with sandbox access.
- Enough concurrent sandbox quota for the selected concurrency and resource policies that permit 2 vCPUs and 4 GiB of memory per sandbox.
- An inference endpoint with tool calling, an API key, and sufficient inference quota. For the example, create a Serverless Inference API key.
- Network access from your machine to the sandbox API and package repositories, and from sandboxes to GitHub, Debian and Python package repositories, and your inference endpoint.
Install the recipe
Clone the recipes repository and install the pinned dependencies:Configure the model and credentials
Load credentials from your secret manager into the process environment. SetCWSANDBOX_API_KEY to your CoreWeave sandbox credential and MODEL_API_KEY to your inference credential. Don’t save keys in the .env file.
In the .env file, set only the model endpoint, model name, and optional project. For the Serverless Inference example, use these settings. Replace [WANDB-TEAM] and [WANDB-PROJECT] with names you can access:
MODEL_PROJECT sets the OpenAI-Project header for Serverless Inference usage tracking.
The model ID is an example, not a requirement. Choose a tool-calling model from the available models. For another compatible provider or a model you serve yourself, replace the model settings in the .env file. Replace [INFERENCE-HOST] and [MODEL-ID] with your provider’s values:
MODEL_API_KEY in the process environment and keep CWSANDBOX_API_KEY set. If the provider doesn’t use the project header, remove MODEL_PROJECT from the .env file and unset any inherited MODEL_PROJECT environment variable before running the command. The endpoint must be reachable from the sandbox. localhost refers to the sandbox itself, not your machine. Other API protocols require adapting the native evaluator options.
The recipe explicitly uses CoreWeave authentication for sandbox operations. The sandbox token stays on the host. Only model configuration and the inference key are injected into the sandbox at creation. The job configuration doesn’t contain credential values.
For broader credential-handling guidance, see environment variables and secrets. Protect downloaded logs and traces, which can contain task data and provider error messages.
Run a setup check
Evaluate one simple task with one sandbox:results-smoke/job-result.json file for completed and errored trial counts. An exit status of zero means no trials remain errored or missing. It doesn’t mean the model passed the task’s assertions. The native agent/automationbench.json export records passed, score, timing, and token usage. The verifier maps passed and score to Harbor rewards named pass and partial_credit. In job-result.json, the stats.evals entry reports the mean of each reward across trials under metrics. Trials without a reward count as zero. reward_stats lists the trials for each reward value.
The agent uses the api toolset, a maximum of 50 response steps, one task at a time per sandbox, and the provider’s default reasoning setting. Native whole-task autohealing is deactivated so Harbor owns trial retries. The model client can still retry individual requests internally.
Increase parallelism
After the setup check succeeds, start with a small subset:
To evaluate the scored set, select its task list and a concurrency your quotas support:
Retry capacity and inference-limit failures
The recipe configures Harbor’sRetryConfig to allow five retries after the original attempt, waiting 30, 60, 120, 240, and 300 seconds. The retry allowlist includes these exceptions:
SandboxResourceExhaustedErrorSandboxUnavailableErrorSandboxRequestTimeoutErrorEnvironmentStartTimeoutErrorApiRateLimitErrorApiUsageLimitError
ApiUsageLimitError by default. The recipe removes it from that exclusion set and adds it to the allowlist. The adapter maps recognized terminal inference quota and rate-limit messages from the native evaluator to the corresponding Harbor exceptions. This mapping depends on provider error text. Inspect logs when a provider reports a different format.
Persistent capacity or credit exhaustion still requires capacity or credits to become available. If errors remain after retries, the script exits nonzero and retains the final error. Ordinary assertion failures, refusals, malformed tool-call JSON, and agent execution timeouts don’t trigger the retry policy.
Harbor removes failed trial directories before retrying. The recipe’s end-of-trial hook archives each attempt first, including its result and available logs. Keep those attempts and disclose infrastructure recovery when reporting scores. Don’t silently substitute better retry outcomes for model failures.
Inspect results and slow trials
Each output directory contains the following artifacts:
Use native exports for task timing and token usage. The adapter doesn’t populate Harbor’s aggregate agent token or cost fields. Record the model, endpoint, benchmark revision, task list, reasoning settings, concurrency, and retry policy alongside results. If you stop a run, report only the exported tasks as completed evaluations and label the results as partial.
Model response time can dominate a task even with few steps or output tokens. Compare
model_time_s with tool_time_s and inspect the agent/eval.log file before attributing delays to sandbox resources. Provider rejection of malformed tool-call JSON is a model or protocol failure, not evidence of sandbox quota exhaustion. AutomationBench lists rollouts that stopped mid-turn in the export’s summary.aborted_tasks field. Their per-task errors arrays can be empty, so check that field and the logs as well as task errors.
Setup has an 11-minute timeout. Agent execution has a separate 45-minute timeout, and the native evaluation command inside it has a 40-minute timeout. Each sandbox has a 1-hour lifetime cap, which ends the sandbox even if a timeout hasn’t elapsed. Agent execution timeouts aren’t retried by this policy.
Stop and clean up
To stop a foreground run, press Ctrl+C and allow Harbor to exit. The job setsdelete=True for teardown. If the process is forcibly terminated or cleanup fails, stop the local runner before deleting remaining sandboxes so it can’t create replacements.
After the runner stops, use its output directory to verify cleanup:
results-smoke with the directory for the job you stopped. The script deletes only recorded sandbox IDs and prints their terminal states. A sandbox that’s already absent is reported as not_found. Investigate a nonzero exit, which means deletion or terminal-state verification failed. The configured lifetime cap also bounds sandbox lifetime if the host process disappears.
Validation scope
During recipe development, the native evaluator exported results for 730 of 800 public tasks across eight parallel CPU sandboxes withopenai/gpt-oss-120b on Serverless Inference. That run was stopped before completion, so its results aren’t a benchmark score. Harbor retry scheduling was verified separately: an injected SandboxResourceExhaustedError, a 30-second backoff, and one real task that completed, exported results, and cleaned up. The recipe replaces the initial experiment’s environment hooks with a BaseEnvironment implementation using public SDK calls. That refactor has offline coverage but hasn’t been run against the live service. A full dataset at high Harbor concurrency hasn’t been validated.
For offline tests and an optional controlled retry check, see the recipe README.