Skip to main content
This page includes release notes for the Weave Python SDK (weave package). For the W&B Models Python SDK (wandb package), see W&B SDK releases. Release packages and per-commit history are on GitHub Releases for wandb/weave.
These entries summarize user-facing SDK and trace server changes, and omit details about Internal-only test, CI, and refactor work. For every merged change, see the GitHub release entry for that tag. If a tag has no GitHub release entry, compare it with the previous tag, for example v0.53.10...v0.53.11.
September 25, 2026

Changed

  • Conversation SDK spans identify the SDK through the wandb.sdk.name, wandb.sdk.version, and wandb.sdk.language OpenTelemetry resource attributes.

Fixed

  • Fixed weave.init() failing with AttributeError: 'OTLPSpanExporter' object has no attribute '_session' after the release of opentelemetry-exporter-otlp-proto-http 1.45.0. weave now requires a version earlier than 1.45.
  • Sibling agent spans that share a start time are ordered by the earliest end time.
September 24, 2026

Added

  • Agent span queries and stats accept insight_filters for filtering by intent category, failure category, failure severity, intent and failure clusters, and intent sentiment.
  • Trace server support for routing Playground chat completions to a Dedicated Inference connection.

Changed

  • The SDK uses the generated trace server client by default. To use the previous client, set WEAVE_USE_STAINLESS_SERVER=false.
  • Token totals that the provider reports, including totals of zero, are preserved.

Fixed

  • Fixed Google GenAI calls receiving traced Weave values instead of the underlying values.
  • Gemini content parts keep their original order.
  • Agent chat messages include the span’s error_type and status_message.
  • Illegal query arguments return 400 instead of 403.
September 10, 2026

Changed

  • The opt-in generated trace server client (WEAVE_USE_STAINLESS_SERVER=true) covers the remaining trace server routes and the calls_complete write path.
  • Updated model provider and cost data.

Fixed

  • When the startup server info request fails, weave.init() reports the underlying cause, such as a TLS certificate error, instead of a generic message.
  • With the generated trace server client, WEAVE_INSECURE_DISABLE_SSL and the SDK request timeout take effect, and parallel chunked table uploads turn on.
  • With the generated trace server client, permanent 4xx errors fail on the first response and name the lost calls instead of being retried.
September 3, 2026

Added

  • Federated identity support. When WANDB_IDENTITY_TOKEN_FILE points to a JWT, the SDK exchanges it for an access token and caches the token in the credentials file (WANDB_CREDENTIALS_FILE), as the wandb SDK does.

Changed

  • Updated model cost data.
August 27, 2026

Added

  • EvaluationLogger accepts trace_scores. Set trace_scores=False to record score values on Evaluation.predict_and_score without creating a separate call for each score. The default is True.

Changed

  • The trace server requires clickhouse-connect 1.6.

Fixed

  • Fixed ImportError when importing the trace server after installing weave[trace_server]. The extra now declares every package that the trace server imports.
  • Fixed unnamed evaluations failing to save when the dataset name reached the 128-character name limit, which gave the evaluation the same object ID as its dataset.
  • Evaluation summaries keep rows that produced no score.
  • ClickHouse stream failures return 502, and the trace server reads structured ClickHouse error codes.
August 13, 2026

Added

  • A @weave.op call links to the agent spans it produces, and a call records the agent span that invoked it.
  • record_error() on Conversation SDK Tool, LLM, SubAgent, and Turn spans records an exception without ending the span.
  • The Claude Agent SDK integration traces subagents.

Fixed

  • Fixed streaming chat completions for Playground and Inference closing before the stream is consumed.
  • Fixed Claude Agent SDK async-iterable queries not being traced per turn, and fixed turns being attributed to the wrong async query.
  • Fixed an UnboundLocalError in the MCP read_resource wrapper.
  • Fixed duplicate OpenAI Agents SDK telemetry.
  • Fixed the mapping of OpenAI cache write tokens.
  • Traces show the details of OpenTelemetry exception events.
  • Latency stats ignore unfinished calls, and root usage stats include OpenTelemetry spans.
  • Cost timestamps are generated in UTC.
  • Fixed new model costs not being added when the database schema was already current.
July 31, 2026

Added

  • WeaveClient.register_custom_runtime() registers or replaces a custom runtime, which is an OpenAI-compatible endpoint with a list of runtime IDs. weave.CustomRuntimeID sets a token limit for an individual ID.
  • Conversation SDK turns record output messages through Turn.output_messages and the output_messages argument of Turn.record().
  • The Claude Agent SDK integration traces image prompts.
  • Trace file storage on Azure supports workload identity through DefaultAzureCredential when no connection string or account key is configured.
  • Cost data for Claude Opus 5.

Changed

  • The Claude Agent SDK OpenTelemetry integration is built on the Conversation SDK.
  • The trace server logs and ignores unknown request fields instead of rejecting the request.

Fixed

  • Cost queries are scoped to the requested project.
  • The trace server validates field qualifiers in queries.
  • Custom provider names are preserved in model selectors.
July 31, 2026

Added

  • WeaveClient.add_cost() accepts cache read and cache creation token prices, and query_costs() returns them.
  • The completions API accepts reasoning_effort.
  • AgentDashboard object type for saving, loading, and forking agent dashboard configurations.
  • Feedback queries return total_count.
  • Agent spans stamped with evaluation metadata link to evaluation results.

Fixed

  • Fixed calls queries failing with 400 because of a collision between boolean and integer parameters.
  • OpenAI cached_tokens values in OpenTelemetry spans count as cache read tokens.
  • Traces of Serverless Inference calls made through the OpenAI SDK show the requested model.
  • Normalized cached input tokens from the Claude Agent SDK integration.
July 16, 2026

Added

  • WeaveClient.get_agent_span_feedback(), get_agent_turn_feedback(), and get_agent_conversation_feedback() return feedback queries for agent spans, turns, and conversations. Reactions added through them appear as thumbs in the agent view.
  • SubAgent.start_subagent() nests a sub-agent under another sub-agent.
  • Server-side sampling for agent spans. The existing call sampling settings, such as WEAVE_INGEST_SAMPLE_RATE, also apply to spans, and traces with weave.eval.* attributes are always kept.
  • The trace server stores base64 and data URL content from agent OpenTelemetry exports as content objects, as it does for SDK traces.
  • Object delete responses report the deleted versions.
  • Cost and provider data for new models, including GPT-5.6.

Changed

  • Turn and SubAgent factory methods are renamed to start_llm(), start_tool(), and start_subagent(). The old llm(), tool(), and subagent() methods still work but emit a DeprecationWarning.
  • Improved read and write performance for calls_complete storage.

Fixed

  • Direct openai calls are traced again when use_otel_v2 is on, which is the default. From 0.52.43 through 0.53.1, direct calls weren’t traced unless you called weave.integrations.patch_openai().
  • Fixed child spans detaching from their parent when they start after the parent’s with block exits, including from another thread.
  • Fixed missing Bedrock cache tokens in streaming Converse responses.
  • Fixed missing costs for Serverless Inference model calls.
  • Summary input_tokens include Anthropic prompt cache tokens.
  • Client-side digests are stable for dictionaries with non-string keys.
  • Support for google-adk 2.3.0 and later.
  • Fixed recursion errors when converting refs in deeply nested payloads.
  • Request validation failures return 400 instead of 500.
  • ClickHouse migrations are idempotent and recover automatically from partial runs.
July 2, 2026

Added

  • Conversation accepts agent_id, agent_description, and agent_version defaults, which each turn inherits. start_turn() accepts the same arguments.
  • Evaluation result summaries include a predict-only cost total.
  • The weave-instrument agent skill in the wandb/weave repository helps a coding agent add Weave tracing to a Python or TypeScript codebase.
  • Cost and provider data for new models, including Claude Sonnet 5.

Changed

  • add_event() on Conversation SDK spans is deprecated. Use set_attributes() instead.
  • The trace_server extra no longer installs ddtrace.
  • Attachment uploads run in parallel on all writes.

Fixed

  • Agent traces go to the active project when you call weave.init() again.
  • Completion error payloads keep the provider’s HTTP status code.
  • Negative JSON path array indices return 422 instead of 502.
July 2, 2026

Added

  • WeaveClient methods for reading agent data, including get_agents(), get_agent_versions(), get_agent_spans(), get_agent_turn(), get_agent_turns(), and get_agent_span_stats().
  • Turn.record() and SubAgent.record() set many fields in one call, and log_turn() accepts agent_id, agent_description, and agent_version.
  • start_conversation(), log_turn(), and log_conversation() accept custom attributes, which are stamped on every span of the conversation.
  • Turn and sub-agent spans record system instructions.
  • OpenTelemetry-native exporter for the OpenAI Realtime integration, selected by default.
  • Evaluation results support match-any and match-all filtering across evaluations, and filtering on inputs.
  • Evaluation result summaries include a predict-only token total.
  • Agent span and stats APIs include query-time costs.

Changed

  • The Session SDK is renamed the Conversation SDK. Use Conversation, weave.start_conversation(), weave.end_conversation(), weave.get_current_conversation(), and weave.log_conversation(), or the weave.conversation module. The session names still work but emit a DeprecationWarning, and session_id and session_name arguments map to conversation_id and conversation_name.
  • Calls from declarative evaluations carry evaluation metadata, matching EvaluationLogger.
  • Improved query performance for evaluation results and call stats.

Fixed

  • Fixed the DSPy, GEPA, and LangChain integrations publishing a new object version on every run because repr() output included memory addresses.
  • weave installs alongside google-adk.
  • Fixed dictify getting stuck in a reference cycle.
  • LLMAsAJudgeScorer no longer includes op methods in its published payload.
  • Fixed model ID detection for Bedrock inference profile ARNs with the invoke API.
  • Clearer error when you don’t have permission to access a project.
  • The trace server returns 400 instead of 500 for bad query parameters and malformed annotation specs, and invalid-field errors list the allowed fields.
June 15, 2026

Added

  • set_attributes() and add_event() on Session SDK Turn, LLM, Tool, and SubAgent spans.
  • Integrations record integration metadata on the calls and spans they create.
  • allow_unsafe_custom_obj_decode setting, True by default. Server-side evaluation workers turn it off, so they decode data-only types such as images and audio but refuse ops and other code-bearing objects.
  • Cost and provider data for new models, including Claude Fable 5.

Changed

  • OpenTelemetry-based integrations are on by default (use_otel_v2, WEAVE_USE_OTEL_V2). In this mode, direct openai calls aren’t traced automatically. To trace them, call weave.integrations.patch_openai() or set WEAVE_USE_OTEL_V2=false.
  • Publishing an object under a name that an object of a different type already uses fails with a 400 error.
  • import weave is faster because jsonschema is imported only when needed.
  • Removed the SQLite trace server.

Fixed

  • Fixed creating monitors by op name, which the documentation examples use.
  • WEAVE_INSECURE_DISABLE_SSL takes effect when you set it after import weave.
  • Fixed flush() stalling on unpaired eager calls when calls_complete is on.
  • Fixed PaginatedIterator instances sharing one page cache.
  • Streaming calls keep caller-set started_at and ended_at values.
  • Fixed Bedrock streaming of tool-use and reasoning deltas.
  • Fixed WAV detection for in-memory audio when libmagic returns application/octet-stream.
  • Fixed scorer auto_summarize for mixed BaseModel and WeaveDict results.
  • OpenTelemetry spans with an all-zero parent_span_id are treated as root spans.
  • The trace server returns 4xx errors instead of 500 for unsupported or unselectable fields.
June 2, 2026

Added

  • WeaveClient methods for managing annotation queues, including create_annotation_queue(), add_calls_to_annotation_queue(), and get_annotation_queue().
  • Preliminary Google ADK integration that traces agent spans automatically through OpenTelemetry.
  • OpenTelemetry-based integrations for the OpenAI Agents SDK and the Claude Agent SDK.
  • Agent span usage includes cache creation and cache read input tokens, and reasoning output tokens.
  • Trace server rescore endpoint that reruns scorers against an existing evaluation run and creates a new evaluation run.
  • Cost and provider data for new models, including Claude Opus 4.8.

Changed

  • The use_calls_complete setting defaults to True. The SDK pairs each call’s start and end and writes the call as one record.
  • The Session SDK respects WEAVE_DISABLED, WEAVE_REDACT_PII, WEAVE_CAPTURE_CLIENT_INFO, and WEAVE_CAPTURE_SYSTEM_INFO.

Fixed

  • str(ref) returns the Weave URI, so leaderboard columns built with str(ref) match evaluation results.
  • Fixed LLM.reasoning appearing in output messages when include_content=False.
  • Project trace storage size includes OpenTelemetry span data.
  • Feedback created_at timestamps are read as UTC.
  • The trace server validates image URLs before fetching them for image completions.
  • Azure file storage no longer overwrites existing blobs.
May 28, 2026

Added

  • Spans that GenAI integrations emit inside an evaluation’s predict_and_score call link to that evaluation automatically.
  • Type-cast inference for dynamic JSON filters also applies to feedback queries.

Changed

  • Every object publish writes an explicit latest alias. Republishing existing content makes it the latest version, and when you delete the latest version, latest falls back to the next surviving version.
  • The MCP integration is renamed FastMCP. weave.integrations.patch_mcp() is now patch_fastmcp(), the weave.integrations.mcp module is now weave.integrations.fastmcp, and the mcp extra is now fastmcp.
  • The OTLP exporter respects WEAVE_INSECURE_DISABLE_SSL.

Fixed

  • Fixed import weave failing with a protobuf TypeError when wandb was installed before weave. weave now requires the opentelemetry packages at version 1.28.0 or later.
  • Saved views persist expand_columns.
  • Failed file uploads are no longer cached, so a later upload of the same file is retried.
May 14, 2026
This release is yanked on PyPI because of a protobuf-related issue. Use 0.52.41 or later, which fixes import weave failing when wandb is installed first.

Added

  • Call filters infer type casts for dynamic JSON fields under inputs, output, attributes, and summary from the comparison value, so many numeric and boolean comparisons no longer need $convert.
  • Trace server support for feedback on agent spans, turns, and conversations.

Changed

  • Reduced per-op overhead when tracing async functions.
  • The ClickHouse trace server uses asynchronous inserts by default.

Fixed

  • Fixed $not filters around $contains, $eq, or $in on large fields returning too few calls.
  • Fixed a potential memory leak in the streaming accumulator.
May 11, 2026
This release is yanked on PyPI because of a protobuf-related issue. Use 0.52.41 or later, which fixes import weave failing when wandb is installed first.

Added

  • The Session SDK emits spans that follow the OpenTelemetry GenAI semantic conventions, and adds weave.start_tool(), weave.start_subagent(), and MediaAttachment.
  • Trace server storage and queries for GenAI agent spans.
  • Cost and provider data for new models, including Grok 4.3.

Changed

  • weave depends on opentelemetry-api, opentelemetry-sdk, and opentelemetry-exporter-otlp-proto-http.
  • weave.init() is faster because it requests server info only once.
  • Setting WANDB_ERROR_REPORTING=false turns off the Weave SDK’s Sentry error reporting.

Fixed

  • EvaluationLogger sends its root call’s start immediately, so the evaluation appears in the UI before log_summary() runs.
  • Fixed flushing returning before chained background work finished.
  • Fixed handling of newer OpenAI Agents SDK span types.
  • Fixed handling of invalid UTF-8 surrogates in trace data.
May 5, 2026

Added

  • weave.link_prompt_to_registry() and the matching WeaveClient method link a published prompt or object version to a registry.
  • Session SDK API surface for manually instrumenting agents: Session, Turn, LLM, Tool, and SubAgent, plus functions such as weave.start_session() and weave.start_turn(). In this release the API is a stub that doesn’t emit spans.
  • Preliminary GEPA integration that logs traces to Weave automatically, installed with the gepa extra.
  • The /eval_results API supports server-side sorting, filtering, and pagination, and resolves dataset-backed evaluation inputs.
  • Cost and provider data for new models, including GPT-5.5.

Changed

  • wandb is an optional dependency again.
  • ClickHouse migrations on replicated clusters use Atomic databases with ReplicatedMergeTree tables and ON CLUSTER statements instead of the Replicated database engine.

Fixed

  • The SDK retries object reads that fail with 404 because of replica lag.
  • Fixed OpenAI Agents SDK responses.create calls not linking to agent traces.
  • Fixed error handling when unwrapping raw OpenAI responses.
April 17, 2026

Added

  • OpenTelemetry integrations follow the latest semantic conventions.
  • Cost calculations and provider integrations include cache token usage.

Changed

  • Improved performance for some call queries by optionally skipping a nested subquery.
  • ClassifierMonitor is exported from the top-level weave package.
  • The calls API accepts an optional query parameter for richer filtering.

Fixed

  • Fixed the evaluation results API SQL for some filter combinations.
  • Fixed a bug where the TypeScript SDK could drop the URL scheme from WANDB_BASE_URL.
  • Fixed a bug where calls with multiple feedback rows could appear multiple times in list results.
  • Fixed ISO-8601 timestamp handling for ClickHouse-backed queries.
  • Fixed distributed ClickHouse mutations that incorrectly appended a _local suffix to table names.
  • Fixed missing spans when tracing OpenAI Agents SDK flows.
  • Fixed Google GenAI tracing when responses include non-text parts.
  • Fixed path sanitization for in-memory trace file artifacts and provider hostname handling in HTTP clients.
April 1, 2026

Added

  • Call queries can resolve human-readable usernames.
  • ref.get() works without explicitly initializing the client in some flows.
  • APIs and storage paths for text-based evaluation results, including improved handling of large evaluation sets.

Changed

  • LiteLLM is pinned for compatibility with bundled integrations.
  • Moonshot is available as a model provider.
  • Underlying write-ahead logging infrastructure improves durability of trace writes.

Fixed

  • Fixed a bug where abandoning a streaming generator could surface GeneratorExit incorrectly.
  • Fixed deserialization when stored objects include extra metadata fields.
  • Fixed feedback filters when a call has multiple feedback rows.
  • Fixed SQLite-backed call cost tracking and invalid trace_id handling in batch upserts.
  • Improved ClickHouse migration behavior, including retries on transient errors and clearer migrator exit status.
  • Fixed a bug where evaluation runs with more than 1000 calls could fail to load for prediction and scoring.
March 19, 2026

Added

  • Write-ahead log support for trace ingestion.
  • Feedback statistics queries for analytics over feedback data.

Changed

  • Ref.uri can be read as a property (ref.uri) without calling ref.uri().
  • You can configure the maximum size of debounced scoring history via environment variables.

Fixed

  • Fixed ClickHouse casting for negative numeric filter values.
  • Fixed replicated database engine errors during on-cluster migrations.
  • Fixed DelegatingTraceServerMixin not forwarding some ServiceInterface methods.
  • Fixed retries when calling the W&B API during weave.init().
  • Fixed a bug where EvaluationLogger could crash when WEAVE_DISABLED is set.
  • Fixed RefJSONEncoder edge cases and classmethod instantiation for subclasses.
March 12, 2026

Added

  • Python SDK methods and HTTP models for object tags and aliases.
  • Instrumentation for the OpenAI Realtime API, including tool calls plus optional audio, text, and voice capture (see the TypeScript examples in the upstream release notes).

Fixed

  • Fixed accumulation of Anthropic streaming completions in traces.
  • Fixed authenticated PUT requests in RemoteHTTPTraceServer.
  • Fixed cost query handling for calls_complete projections.
March 10, 2026

Added

  • Trace server support for tags and aliases on stored objects.
  • Claude Agents tracing integration.
  • Timestamps and time-to-first-token metrics for Realtime sessions.
  • Monitors can use merged scorers.

Changed

  • Improved OpenTelemetry performance with a cross-request operation reference cache.

Fixed

  • Fixed multiple issues in cost query construction, including escaping of internal fields and distributed ClickHouse setups.
  • Fixed LangChain integration handling for Pydantic v2 Run objects.
  • Fixed edge cases in prediction and scorer resolvers when inputs or metadata are missing.
  • Fixed Vertex AI text accumulation and thread visibility in calls_complete queries.
March 10, 2026

Added

  • Score backfill API for recomputing stored scores.
  • Gemini request tracking in the TypeScript SDK.

Changed

  • Sharded distributed calls tables by trace_id or project_id for better query performance.
  • Simplified calls_complete query plans for lower latency.

Fixed

  • Fixed a bug where Gemini media could fail to render in the Weave UI.
  • OpenTelemetry batch inserts now use async ClickHouse inserts.
  • Fixed NO_PROXY handling in HTTP clients.
  • Fixed entity versus team naming in some error messages.
February 27, 2026

Added

  • Usage APIs expose metadata for unfinished calls.
  • Optional python-magic integration for richer MIME detection.
  • Realtime threads participate in usage summaries.
  • Schema support for tags and aliases on Weave objects.
  • Structured eval_results query API for evaluation tables.

Changed

  • Improved indexing by storing timestamps as strings in ClickHouse.
  • Reduced duplicate work during batched file creation on call upsert.

Fixed

  • Fixed a bug where generators did not respect configured sampling rates.
  • Fixed buffering so streams flush immediately when a call ends.
  • Fixed sort-query index errors and deterministic JSON serialization for digest computation.
  • Fixed OpenTelemetry display names, make_safe_name handling, and LangChain serialization for Pydantic models.
February 14, 2026

Added

  • OpenTelemetry resource attributes can carry W&B run and project variables.
  • The ORM supports $lt and $lte comparisons.
  • OpenTelemetry projects can write directly into the calls_complete table.
  • General availability improvements for OpenAI Realtime tracing.

Changed

  • Improved performance for table scans and call statistics queries.

Fixed

  • Fixed filtering calls by thread ID, including cases where filters were incorrectly optimized away.
  • Fixed idempotent annotation queue state updates.
  • Fixed a bug where weave.finish() did not always flush pending client data.
  • Fixed iterator typing for PaginatedIterator, Dataset.select metadata preservation, and large trace size queries that could run out of memory.
February 3, 2026

Added

  • Usage statistics APIs plus /trace/usage and /calls/usage endpoints for aggregated usage.
  • Optional performance mode flag for high-throughput deployments.
  • Anthropic structured-parse patching.
  • Saved views support a column_order field.

Changed

  • Improved PREWHERE optimization for distributed cluster queries.

Fixed

  • Fixed a bug where failed ClickHouse inserts could leak buffered rows.
  • Fixed a bug where importing IPython at import time could slow cold starts.
  • Fixed synchronous mutation migrations, summary filtering on calls_complete, and Google GenAI token overcounting.
  • Fixed OpenTelemetry spans with monitors and duplicate upload handling for Google Cloud Storage.
January 20, 2026

Fixed

  • Removed redundant HTTP response capture in some integrations.
  • Google GenAI tracing now records system instructions.
  • Google GenAI tracing now records thinking tokens separately from completion tokens.
January 15, 2026

Added

  • TypeScript helper APIs for working with prompts in the Node SDK.
  • Trace server safely autoconverts Base64 payloads where appropriate.
  • Leaderboard schema updates for upcoming comparison features.
  • redact_pii_exclude_fields setting to fine-tune PII redaction.
  • Audio inputs in LLMAsAJudgeScorer and richer op metadata (kinds and colors) for integrations.

Fixed

  • Fixed invalid characters blocking op creation in edge cases.
  • Fixed HTTP and HTTPS proxy handling for the HTTPX client.
  • Fixed nested tracing when wrapping generators.
January 8, 2026

Added

  • Parsing helpers for Logfire Pydantic AI instrumentation inputs and outputs.

Fixed

  • Fixed non-deterministic ordering for large-table evaluations.
  • Fixed cost queries that omitted input and output tokens.
  • Fixed distributed replicated ClickHouse tables and guarded Kafka flush behavior behind configuration.
January 8, 2026

Added

  • Prompts and template variables on LLMStructuredCompletionModels persist through completion calls.

Fixed

  • Fixed ClickHouse import compatibility issues in some environments.
November 26, 2025

Added

  • Completions streaming APIs accept prompts and template variables.
  • ObjectRef.from_uri reconstructs objects from Weave URIs.
  • OpenAI Responses API tracing records x-request-id headers.
  • Bedrock Agents integration coverage.
  • TypeScript SDK withAttributes helper for span metadata.

Fixed

  • Fixed a memory leak in the OpenAI Agents tracing processor.
  • Fixed a bug where refs could still publish when Weave was disabled.
  • Fixed iterator behavior when migrating HTTP client code from requests to httpx.
For releases older than 0.52.20, see GitHub Releases for wandb/weave. Those entries use the default GitHub compare format (pull request links and commit prefixes). This docs page is updated for newer versions with customer-facing wording.
Last modified on September 29, 2026