Skip to main content
Weights & Biases integrates with AI assistants in two complementary ways:
  • W&B Skills teach coding agents how to use Weights & Biases effectively in your code and analysis workflows.
  • The Weights & Biases MCP Server connects AI assistants to your Weights & Biases data and documentation so they can answer natural-language questions about your runs, traces, evaluations, and artifacts.
Use Skills when you want your coding agent to write or modify Weights & Biases-aware code. Use the MCP server when you want an AI assistant to query live Weights & Biases data or search Weights & Biases documentation. The two work well together: Skills provide workflow patterns, and MCP provides data access. Depending on the integration, Weights & Biases works with several major coding agents, IDEs, and chat assistants, including:
  • Claude Code
  • Codex
  • Cursor
  • Gemini CLI
  • Visual Studio Code (VS Code)
  • Mistral LeChat
  • Claude Desktop
For a full list of agents supported by W&B Skills, see the W&B Skills CLI documentation.

W&B Skills

W&B Skills are reusable instruction sets that teach coding agents how to use Weights & Biases effectively. Instead of manually guiding your agent through W&B APIs and best practices, install Skills so that the agent can work with experiment tracking, tracing, evaluations, and monitoring on its own.

Capabilities

Skills cover both the W&B Python SDK (training runs, metrics, artifacts, sweeps) and the Weave SDK (traces, evaluations, scorers). They include helper libraries, reference docs, and data analysis patterns so your agent can handle the following workflows.

Prerequisites

W&B Skills require the following:
  • Node.js for the npx command.
  • A Forge API Key. Create one at forge.coreweave.com/settings#apikeys and then set it as an environment variable. Replace [YOUR-API-KEY] with your API key:
  • Optional: Set your Weights & Biases project name as a WANDB_PROJECT environment variable. This lets your agent target the correct Weights & Biases project without you specifying it each time.

Install W&B Skills

Choose a global installation to make Skills available to all your projects, or a project-specific installation to scope Skills to a single project. To install W&B Skills globally for all your projects, use the --global flag:
To install Skills only for the current project, run the install command from your project directory without the --global flag:
Install Skills for specific agents using the --agent flag:
For a list of --agent and --skill options, see the Vercel Labs skills CLI documentation. After installation completes, your agent has access to W&B Skills and is ready to handle Weights & Biases-related tasks.

Use W&B Skills

Ask your agent to perform Weights & Biases-related tasks for your project. The following example prompts demonstrate some of the tasks your agent can do with W&B Skills:
  • “Log training metrics for my PyTorch model to Weights & Biases.”
  • “Analyze the loss curves for my last 10 runs and identify the best performing configuration.”
  • “Trace my LangChain agent and log the results to Weave.”
  • “Run an evaluation on my agent using the test dataset and summarize the results.”
  • “Find the failure modes in my last evaluation and classify them.”
  • “Compare the configs of run A and run B and show me the differences.”

W&B Skills usage tips

Skills performs better with specific queries than with broad, open-ended questions. The following table compares recommended prompts with prompts that are too vague.

Weights & Biases MCP Server

Model Context Protocol (MCP) is an open standard that lets AI agents call external tools. The Weights & Biases MCP Server gives your IDE, coding assistant, or chat agent direct access to your Weights & Biases data and documentation. With this access, your agent can answer questions about your runs, traces, evaluations, and artifacts without copy-paste. For details of what you can do with the server, see the Weights & Biases MCP Server capabilities section.

Deployment types

The Weights & Biases MCP Server is available in two deployment options. Use the hosted server for the fastest setup, or set up a local version if you need more isolation and flexibility. The local version requires your client to use a different URL to access the server.

Hosted server (recommended)

A Weights & Biases-managed MCP server that your client connects to over HTTP with your API Key. No installation, no local process to maintain.Use the hosted server

Local install

Run the MCP server on your own machine over STDIO or HTTP. Use when you need air-gapped operation, pinning to a specific release, custom server behavior, active server development, or support for a client that only speaks STDIO.Run the MCP server locally

Prerequisites

Before you configure any client, make sure you have the following in place:
  • Create an API Key at forge.coreweave.com/settings#apikeys.
  • Set the key as the WANDB_API_KEY environment variable, or pass it to your client as a bearer token.
  • For Dedicated Cloud, Self-Managed, and local installs against a non-default instance, set the WANDB_BASE_URL environment variable to your instance URL.
  • Weights & Biases pins the mcp SDK to 1.14.0, release 2024-11-05. Clients must connect using mcp SDK 1.14.x. For streamable HTTP support in W&B Dedicated Cloud, mcp SDK 1.14.x release 2025-03-26 or newer is required.

Use the hosted server

Weights & Biases runs a managed MCP server for every deployment type. You don’t need to install anything. Configure your client to connect over HTTP with an API Key in the Authorization header.

Connection URL

The URL depends on your type of Weights & Biases deployment: For Dedicated Cloud or Self-Managed, replace https://mcp.withwandb.com/mcp with https://[YOUR-INSTANCE]/mcp and keep everything else the same. The following client configurations use the Multi-tenant URL.
Register the Weights & Biases MCP server with Claude Code, replacing the bearer token with your API Key:
Add --scope user to configure Claude Code globally. Omit it to configure only the current project.Verify the connection by asking List my W&B entities. The agent should call list_entities_tool and return your username and any teams. If the connection fails, see Troubleshooting. For more information, see Claude Code’s MCP documentation.

Run the MCP server locally

A local install is an alternative to the hosted server, not the default for any deployment type. Use it when the hosted server doesn’t fit your setup. Common reasons to run locally:
  • Air-gapped or offline environments where your client can’t reach a hosted Weights & Biases endpoint.
  • Pinned version. The hosted server follows the main branch. A local install can pin to a specific release tag.
  • Custom server behavior such as changing tool descriptions, adding tools, or setting a non-default response token budget.
  • Active development on the server itself.
  • STDIO-only clients or clients that require a local process.
For Dedicated Cloud or Self-Managed users, prefer the hosted path. Only use a local install from wandb/wandb-mcp-server if the hosted server isn’t yet enabled on your instance or one of the preceding reasons applies. Set a WANDB_BASE_URL environment variable to your instance URL.

Local prerequisites

To run the server locally, make sure you have the following:
  • Python 3.11 or higher.
  • uv or pip.
  • An API Key, set as WANDB_API_KEY.
  • WANDB_BASE_URL set to your instance URL if you use Dedicated Cloud or Self-Managed.

Install the server

Choose an install method and run the following command to install the MCP server:

Configure your client

After installing the server, configure your client to launch it. Select your MCP client and then run the following configuration, replacing [YOUR-WANDB-API-KEY] with your API Key as necessary:
Register the local server with Claude Code. Add --scope user for a global configuration.

Run the server with HTTP transport

For web-based clients and for testing, run the server with HTTP transport:
To expose a local server to external clients, such as the OpenAI Responses API, use a tunnel:
Update your MCP client configuration to use the tunnel URL.

Environment variables

The following environment variables control authentication, instance routing, and server behavior for local installs. Set them in your client’s env block or export them in your shell. For the complete command-line reference and advanced options, see the wandb-mcp-server README.

Weights & Biases MCP Server capabilities

Use the MCP server to analyze experiments, debug traces, create reports, manage registry and artifacts, and answer questions from the Weights & Biases docs. The following example prompts demonstrate some of the tasks you can ask your agent to perform when it’s connected to the Weights & Biases MCP Server:
  • “Show me the top 5 runs by eval/accuracy in your-team/your-project.”
  • “How did the latency of my hiring agent’s predict traces evolve over the last month?”
  • “Generate a W&B report comparing decisions made by the hiring agent last week.”
  • “What versions of the production-model artifact exist, and what changed between v2 and v3?”
  • “How do I create a leaderboard in Weave?”

Available tools

The server offers several tools grouped by purpose. The following table lists each tool’s name, when the agent should use it, and a concrete prompt you can use to invoke that tool.
Tools that help you discover project and entity names, and inspect schemas.

Schema-first trace queries

For Weave trace queries, call infer_trace_schema_tool first to discover available fields, then call query_weave_traces_tool with a precise column list and detail_level: This pattern keeps token usage low for broad questions and lets the agent escalate to full only for the traces that matter.

Usage tips

The following sections describe practices and workflows that help you get better results from the Weights & Biases MCP Server. Start with the general practices, then read the section that matches your workload for more specific advice and multi-step tool chains.

General best practices

Follow these practices regardless of your use case:
  • Specify the entity and project. MCP tools need an explicit entity (your team or personal account) and project name. Include both in every question, for example “in your-team/your-project”.
  • Ask focused questions. Prefer “Which eval had the highest F1 score?” to “What is my best evaluation?”. Specific metrics and time ranges produce better tool calls.
  • Verify full retrieval. For broad questions such as “What are my best performing runs?”, ask the agent to confirm it retrieved all available runs rather than only the most recent ones.
  • Combine with W&B Skills. W&B Skills teach coding agents how to structure Weights & Biases workflows. Skills provide patterns and MCP provides data access, and the two work well together.

For trace-heavy workflows

Follow these practices when working with Weave traces:
  • Start with the schema. Call infer_trace_schema_tool before query_weave_traces_tool to provide the agent with the valid fields and filter values.
  • Pick the right detail_level. Use schema to browse, summary (the default) for analysis, and full only when drilling into a small number of specific traces.
  • Chain resolve_trace_roots_tool. After a child-trace query, pass the resulting trace_id list to resolve_trace_roots_tool to map each trace to its root session in one batched call.
  • Prefer summarize_evaluation_tool for evals. It aggregates the Evaluation.evaluate and predict_and_score hierarchy automatically. Only fall back to query_weave_traces_tool for raw trace data.
For an end-to-end workflow, see Triage failing LLM calls.

For run-heavy workflows

Follow these practices when working with W&B runs:
  • Probe before you query. Call probe_project_tool on an unfamiliar run-based project to discover metric keys, config keys, and tags before constructing GraphQL.
  • Use get_run_history_tool for time series. GraphQL doesn’t sample, so for loss curves and other time-series data get_run_history_tool is both faster and cheaper.
  • Let compare_runs_tool do the diff. It returns config and metric deltas with aligned history in a single call, avoiding manual comparison.
  • Run a health check first. When a training run looks wrong, call diagnose_run_tool before digging into history manually.
For end-to-end workflows, see Diagnose a bad training run and Summarize evals and compare model versions.

For Dedicated Cloud and Self-Managed

Follow these practices for non-multi-tenant deployments:
  • Prefer the hosted server on your instance at https://[YOUR-INSTANCE]/mcp. It exposes the same tools as the Multi-tenant server with no client-side WANDB_BASE_URL needed. Only fall back to a local install if the hosted server isn’t yet enabled.
  • When you do run locally against your instance, set WANDB_BASE_URL to your instance URL in the client’s env block. Without it, the server targets api.wandb.ai and the server returns no data.
  • Rate limits on Dedicated Cloud are separate from Multi-tenant. See Dedicated Cloud rate limits for defaults and how to request changes.

For local installs

Follow these practices when running the server on your own machine:
  • Prefer STDIO transport for desktop clients (Cursor, VS Code, Claude Code, Claude Desktop). Only switch to HTTP transport when a client explicitly requires it (for example, the OpenAI Responses API).
  • When tool calls fail silently, set MCP_SERVER_LOG_LEVEL=DEBUG in the client’s env block and recheck the client’s MCP logs.
  • If you install from GitHub (uvx --from git+https://github.com/wandb/wandb-mcp-server wandb_mcp_server), uvx pins to the default branch. Pin an explicit tag by appending @v0.3.2 to the Git URL when you need a stable version.
Most real questions need more than one tool. The following workflows show common multi-step tool chains you can ask your agent to perform.

Explore an unfamiliar project

To explore what has been logged to a project, chain these tools:
  1. list_entities_tool to find an entity or team.
  2. query_wandb_entity_projects to find the project.
  3. probe_project_tool for run-based projects, or infer_trace_schema_tool for Weave trace projects.
  4. A targeted query_wandb_tool or query_weave_traces_tool call using the discovered keys.

Triage failing LLM calls

To find bad traces and the sessions that produced them, chain these tools:
  1. query_weave_traces_tool with a filter on error or exception fields, and detail_level="summary".
  2. resolve_trace_roots_tool on the resulting trace_id list to map each failure to its root session.
  3. query_weave_traces_tool with detail_level="full" on a small number of specific roots to drill in.
  4. create_wandb_report_tool to document the findings.

Diagnose a bad training run

To run a health check on a suspicious training run, chain these tools:
  1. get_run_history_tool to pull the loss and validation curves.
  2. diagnose_run_tool for automated convergence, overfitting, and NaN checks.
  3. compare_runs_tool against a known-good baseline run.
  4. create_wandb_report_tool with line-plot panels to share the diagnosis.

Summarize evals and compare model versions

To find which model version performed best on an evaluation, chain these tools:
  1. summarize_evaluation_tool for per-scorer pass rates and error counts.
  2. list_artifact_versions_tool on the relevant model collection.
  3. compare_artifact_versions_tool between the candidate and current production version.
  4. log_analysis_to_wandb and create_wandb_report_tool to publish the comparison.

Troubleshooting

Use the following table to help you diagnose and resolve issues using the Weights & Biases MCP Server:
Last modified on September 30, 2026