Skip to main content

3 posts tagged with "http"

HTTP(s) data connector related topics and usage

View All Tags

Spice v2.4.0-rc.1 (Oct 8, 2026)

ยท 64 min read
Sergei Grebnov
Member of Technical Staff at Spice AI

Spice v2.4.0-rc.1 is now available! ๐Ÿ”ฅ

Spice v2.4.0-rc.1 is the first release candidate for v2.4.0. It adds performance improvements, S3 event-driven ingestion, SQL results-cache warmup, and adaptive HTTP rate controls. The release also upgrades to DataFusion v55, Ballista v55, Arrow v59, Vortex v0.86, Iceberg v0.11, and Turso v0.81.

Highlights in v2.4.0-rc.1 include:

What's New in v2.4.0-rc.1โ€‹

Performance & Query Engineโ€‹

This release upgrades Apache DataFusion to the v55.2.0 dependency line and Apache Arrow to v59.3.0. It also upgrades Vortex to v0.86.1, Apache Iceberg to the v0.11.0 fork, and Apache Ballista to v55.

Apache DataFusion v55โ€‹

The DataFusion v55 release adds the following improvements:

  • Sort pushdown and TopK pruning: Parquet scans reevaluate each unread row group as the threshold for ORDER BY ... LIMIT tightens. They skip groups that cannot contribute to the result. TopK pruning also supports multiple sort columns.
  • Join planning: The optimizer converts eligible inner joins to semi joins and removes redundant sides of outer joins. It also orders filter predicates by estimated cost.
  • Aggregation and expressions: Multi-column GROUP BY uses column-oriented storage for all supported key types, such as fixed-size binary UUIDs. More string functions preserve dictionary encoding, and IN lists use specialized paths for small integer types.
  • Parquet reads: Scans skip nested fields that the declared schema does not contain. They also skip page-index reads when a file has no page index.
  • Spill handling: Sorts bound the number of streams in a merge. If memory is insufficient, they spill the largest stream again in smaller batches.
  • SQL diagnostics and functions: EXPLAIN accepts PostgreSQL-style options and FORMAT pgjson. New array functions cover element-wise addition, subtraction, scaling, sums, and averages.

Spice carries these changes through its query plans and preserves statistics across plan wrappers. Cayenne keeps Vortex scans below 10 MiB unsplit to avoid repeated footer reads. See #14612.

Vortex v0.79.0 to v0.86.1โ€‹

Spice v2.3.2 used the Vortex v0.79.0 fork. This upgrade covers the full upstream range from v0.79.0 through v0.86.1, not only the v0.86 changes:

  • Types and arithmetic: Vortex adds native Map arrays, Arrow map conversion, and operations and compression for maps. It also adds union arrays and decimal addition, subtraction, multiplication, and division.
  • Row selection: Piecewise-sequence indices represent contiguous selections without an expanded index for every row. Specialized paths handle chunked arrays, fixed-size lists, variable-length lists, and binary values.
  • Scans and expressions: Layout scans gain a physical plan and expression optimization. Filters can pass through scalar functions with multiple arguments. Filters on wide lists restrict child elements to the selected range, and constant masks can resolve from metadata.
  • Compression: Binary arrays support FSST compression with variable-length offsets. OnPair becomes a stable encoding for reads and gains a storage-backed dictionary. Vortex can convert run-end arrays of lists and decimals.
  • File metadata: Files can store custom metadata. Readers cache decoded type descriptions, and file-format editions define the supported encodings and types.
  • Memory and execution: Builders append nested values in batches, preserve list views, and propagate buffer allocators. Expression rewrites retain unchanged nodes, and row functions support batch execution.
  • Correctness and validation: NULL handling changes cover dictionary predicates, BETWEEN bounds, and empty arrays. Readers add validation for footer offsets, compression metadata, and array indices. Other corrections cover nested scalar hashes, decimal operations, and variable-length binary selections above 4 GiB.

See the complete Vortex v0.79.0 to v0.86.1 changelog for every upstream change. Changes to standalone Vortex bindings and GPU execution do not imply new Spice features.

Apache Iceberg v0.11.0โ€‹

Spice updates its Iceberg reader and catalog integrations from v0.10.1 to the v0.11.0 fork for DataFusion 55 and Arrow 59. The upstream v0.11 changes extend reads and catalog compatibility:

  • Iceberg v3 reads: The reader applies deletion vectors from Puffin files and carries row identifiers and sequence numbers through scans.
  • Metadata-only scans: Metadata-only projections do not require data-column reads. Manifest reads reuse partition types, and positional-delete processing buffers runs instead of allocating a key per row.
  • REST catalogs: Clients negotiate server-advertised endpoints and use session-scoped OAuth2 authentication.

The upgrade also aligns Iceberg storage with OpenDAL v0.58. See #14771.

Apache Ballista v55โ€‹

Spice.ai Enterprise feature. See the Enterprise documentation.

The Ballista v55 upgrade adds virtual-core resource accounting for distributed tasks. A protocol-version handshake detects incompatible schedulers and executors. Cancellation identifies tasks by their task IDs, and task-state records use an append-only model.

Schedulers and executors must run the same version. Upgrade all cluster components together.

S3 Event-Driven Ingestionโ€‹

S3 listing datasets can use refresh_mode: changes with S3 event notifications delivered through SQS:

datasets:
- from: s3://my-bucket/events/
name: events
params:
file_format: parquet
s3_region: us-east-1
s3_auth: iam_role
s3_changes_queue_url: ${ secrets:events_queue_url }
acceleration:
enabled: true
engine: cayenne
mode: file
refresh_mode: changes

New object notifications append the object's rows. A periodic listing backfill covers missed or expired notifications. Each dataset needs its own queue and permissions to read and delete SQS messages, alongside its S3 read/list permissions.

This is object ingestion: removal notifications are ignored by default. Set s3_on_object_removed: rebuild to rebuild the entire prefix when an object is removed. An overwrite of an already applied object key is not ingested again; use new object keys for incoming data. See #14121.

SQL Results-Cache Warmup and Shared Fetchesโ€‹

The SQL results cache can persist query plan shapes and replay them after a dataset's first full or append refresh:

runtime:
caching:
sql_results:
enabled: true
warmup: on_first_refresh

For example, run these queries against an accelerated orders dataset with warmup enabled:

SELECT id, status FROM orders WHERE id = 1;
SELECT id, status FROM orders WHERE id = 2;

Spice records one query shape because only the equality-filter value differs. After a restart and the dataset's first full or append refresh, warmup reruns that shape with distinct id values from the refreshed dataset, filling the cache before the dataset becomes ready. Keep the local .spice/data directory across restarts, or configure runtime.state.location, to retain recorded query shapes.

Warmup replays up to ten distinct recorded query shapes and tries up to 1,024 distinct filter-value combinations per shape, stopping when the cache is full. A dataset stays not ready until warmup completes. Later refreshes do not repeat warmup. The feature requires the default plan-based cache key; cache_key_type: sql is incompatible. See #14178.

Concurrent cache-miss fetches for the same request now share a source fetch. SQL results caching also includes changes for tables updated during query execution and for stale results served while revalidation runs. HTTP dataset caching supports RFC 5861 stale-if-error handling. See #14142, #14710, #14708, and #14134.

SQL, search, and embedding caches now use Spice's sharded cache backend. Existing engine: moka and engine: pingora values are accepted for configuration compatibility but no longer select a backend.

Adaptive HTTP Rate Controlsโ€‹

HTTP rate controls now adapt admission to upstream failures while staying within configured request limits. rate_control_acquire_timeout bounds how long a request waits for capacity and defaults to the connector's client timeout. rate_control_failure_threshold and rate_control_window control the response to upstream failures.

In Spice.ai Enterprise, instances that share a runtime.state.location coordinate per-second and per-minute limits through shared state; OSS instances keep these limits in memory independently. Concurrency limits remain local to each instance. Components sharing an upstream origin must use matching rate-control settings. See HTTP rate-control documentation and #14143.

ORC Files and Object Metadata Queriesโ€‹

Listing connectors support file_format: orc. Object-store listing also uses predicates on metadata columns, including _last_modified, to narrow eligible objects. Queries that select only partition or metadata columns can use those values without reading file contents. See #14075, #14265, #14303, and #14116.

Hugging Face Datasetsโ€‹

The new Hugging Face data connector queries and accelerates datasets from the Hugging Face Hub. It supports Parquet, CSV, TSV, JSON, and ORC files. Public datasets need no credentials:

datasets:
- from: hf://datasets/stanfordnlp/imdb/plain_text/
name: imdb
acceleration:
enabled: true

The location format is hf://datasets/<owner>/<dataset>[@<revision>][/<path>]. A path can select a file, a folder, or a glob. A revision can name a branch, a tag, or a commit. Use @~parquet to select the Hub's automatic Parquet conversion.

Set hf_token for private or gated datasets. Set hf_endpoint for a Hub mirror or proxy. Each scan reads one commit. Refreshes follow the selected branch, but the dataset keeps its registered schema until reload. See #14877.

Automatic Primary-Key Handlingโ€‹

Cayenne keeps one row per primary_key without requiring an on_conflict policy. When a dataset sets time_column, the row with the newest time wins; without it, the last arrival wins. This applies to full and append refreshes as well as writes. Existing explicit conflict policies remain accepted during the deprecation period.

datasets:
- from: s3://my-bucket/orders/
name: orders
time_column: updated_at
params:
file_format: parquet
acceleration:
enabled: true
engine: cayenne
mode: file
primary_key: id

For a read-write dataset whose writes should stay in its acceleration, set acceleration.write_mode: acceleration. Its source need not support writes. This mode cannot be combined with a dataset that refreshes by changes.

Cayenne secondary indexes also support dynamic join filters, and their write handling covers inserts, updates, deletes, and refreshes. See #14726, #14282, and #14593.

PostgreSQL and MySQL replication, and MongoDB change streams, rejected Cayenne datasets that omitted on_conflict. Their validation still required an explicit upsert policy. These sources now accept Cayenne datasets with primary_key alone. See #14880.

For file-mode Cayenne datasets, append refreshes failed with a configuration that combined primary_key, time_column, and retention_sql. The refresh selected a version-resolution path that did not support retention. The refresh now resolves each key's newest version before Cayenne applies retention. See #14878.

Cayenne maintained aggregates now share one compact index of per-key contributions across views. Rebuilds capture concurrent writes and apply them after the scan. The aggregate budget uses 10% of a bounded query pool without the former 512 MiB cap. An unbounded pool retains the 512 MiB budget. See #14762.

Read Datasets from Published Snapshotsโ€‹

Spice.ai Enterprise feature. See the Enterprise documentation.

A dataset can now read published acceleration snapshots directly, without configuring the original source connector:

datasets:
- from: s3://my-bucket/spice/snapshots/orders/
name: orders
params:
file_format: snapshot
s3_region: us-east-1

Spice reads the snapshot metadata to select the engine, restores the published data, and checks for newer snapshots. The dataset is read-only. Snapshot-mode readers can also use S3 notifications delivered through SQS, with periodic checks retained for missed notifications. Each reader process needs its own queue.

This release includes changes to snapshot retries, slow-connection bootstrap, publication metadata, and coordination between snapshot archiving and Cayenne maintenance. Cayenne datasets with a datalake tier cannot create acceleration snapshots. See snapshot documentation, #14529, and #14335.

Connector and Protocol Updatesโ€‹

  • Connector status: ADBC, Databricks Spark Connect and SQL Warehouse, FlightSQL, Glue, HTTP/HTTPS, Iceberg, Localpod, and MongoDB are now Stable data connectors.
  • MCP: support for specification version 2026-07-28, alongside the earlier protocol era. See #14043.
  • GitHub: nested GraphQL pagination and rate-limit pacing updates; the default concurrency limit is now four. See #14179 and #14431.
  • GitHub nested pages: Scans could return incomplete reviews or comments because pagination accepted a short page as complete. The connector now rejects incomplete connections and repeated cursors. It retries a failed nested page without another fetch of the outer page. Datasets with the same token also share the REST quota. See #14862.
  • Iceberg REST: Clients could read an empty dataset because the catalog synthesized metadata without snapshots. The catalog now returns the source metadata for Iceberg datasets that Spice reads unchanged. Other datasets return 400 BadRequestException. Clients need their own storage credentials. See #14588 and Breaking Changes.

Other Fixesโ€‹

The release includes fixes in the following areas; the linked PRs provide details of the changes:

  • Cayenne queries and writes: NULL-aware NOT IN, maintained aggregates, dynamic filters, partition-filter forwarding, memory-mode DML and retention, and primary-key handling across CDC checkpoints. See #14429, #14761, #14370, #14047, and #14344.
  • Startup and reloads: retry datasets with unavailable sources, serve existing accelerations during source outages, and invalidate cached plans and results on catalog replacement or dataset unload. See #14623, #14624, #13914, and #14365.
  • Federation: local evaluation of casts and functions whose source semantics differ, plus filter pushdown changes for DynamoDB, Cosmos DB, and MongoDB. See #14484, #14601, and #14419.
  • HTTP and GraphQL: response-status handling during refresh, retry-budget handling, non-JSON gateway responses, and URL redaction in HTTP errors. See #13538, #14313, #14781, and #14490.
  • Search: deletion of obsolete Elasticsearch chunks, non-finite embedding handling, and source-scan coordination during full-text refresh. See #13960, #13902, and #14663.
  • Models and tools: tool-call-only assistant turns, required tool choices, streaming tool-use completion, model-load diagnostics, and propagation of the API-key principal into MCP tool calls. See #14232, #14460, #14548, and #14828.
  • CDC shutdown and reconnects: source-position recording before accelerations close, and MySQL shared-stream reconnect handling. See #14702 and #14751.
  • SQL weekdays: date_part('dow') aligns with EXTRACT(dow), with Sunday represented as zero. See #14796.
  • Cayenne schema statistics: Decimal bounds could retain an old scale after schema evolution because maintenance published statistics from the previous schema. Cayenne now rejects statistics from an obsolete schema and keeps row counts conservative. See #14856.
  • Vector search: vector_search planning failed after the DataFusion 55 upgrade because a second optimization pass tried to reorder a join with a dynamic filter. The planner now preserves that join's input order. See #14857.

Default accelerator: Datasets and views that enable acceleration without engine now use Cayenne. Explicit engine settings keep their behavior. Storage still defaults to memory. Set mode: file for persistent acceleration. See Breaking Changes for migration guidance.

Dependency Updatesโ€‹

Dependency / ComponentVersion
DataFusionv55.2.0
Apache Arrowv59.3.0
Vortexv0.86.1
Apache Icebergv0.11.0
Apache Ballistav55.0.0
Tursov0.8.1
ADBCv0.24
Rust toolchainv1.98.1

Contributorsโ€‹

Breaking Changesโ€‹

Cayenne is the default accelerator on supported platforms. A dataset or view that omits acceleration.engine switches from Arrow to Cayenne. To retain Arrow, set engine: arrow explicitly before upgrading. Windows keeps Arrow as its default. Persistent datasets should continue to name their engine and use mode: file.

on_conflict is deprecated and scheduled for removal in v3.0. Cayenne automatically keeps one row per primary key, choosing the newest time_column value when configured, or the last arrival otherwise. Existing explicit policies remain supported during the deprecation period. Review those policies before removing them, especially drop or policies that reject conflicting rows.

on_conflict no longer routes writes to the acceleration. For read-write datasets whose writes should stay in the acceleration, use:

acceleration:
enabled: true
engine: cayenne
write_mode: acceleration

This setting cannot be used with refresh_mode: changes, including a connector's default change-stream mode. The default write_through and write_back modes require a writable source.

Cache engine selection is retired. engine: moka and engine: pingora remain accepted but are ignored. Remove the field and use caching_policy to select eviction behavior.

GitHub connector default concurrency is four. Review explicit concurrency settings if your deployment relied on the previous default.

HTTP rate-control waits are bounded by default. Requests waiting for rate-control capacity now time out after the connector's client timeout. Set rate_control_acquire_timeout to a suitable duration, or 0 to retain the previous unbounded wait behavior.

Cayenne acceleration snapshots are unavailable for datalake-tier datasets. Review snapshot settings on datasets using cayenne_datalake_location; this release disables snapshotting that configuration.

Iceberg REST no longer synthesizes metadata for unsupported datasets. GET /v1/namespaces/{namespace}/tables/{table} returns 400 BadRequestException for accelerated datasets, views, and other datasets that Spice does not read unchanged from Iceberg. If a client used this endpoint for schema discovery, use SQL DESCRIBE or information_schema.columns instead. Query these datasets through /v1/sql or Arrow Flight SQL. For eligible Iceberg datasets, clients read the source metadata and need their own storage access. See the Get a table API.

Cookbook Updatesโ€‹

The Spice Cookbook provides recipes to help you get started with Spice.

Upgradingโ€‹

To upgrade to v2.4.0-rc.1 once the release artifacts are available, use one of the following methods:

CLI:

spice upgrade v2.4.0-rc.1

Docker:

Pull the spiceai/spiceai:2.4.0-rc.1 image:

docker pull spiceai/spiceai:2.4.0-rc.1

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.4.0-rc.1

AWS Marketplace:

Spice is available in the AWS Marketplace. Marketplace availability follows its published versions.

What's Changedโ€‹

Changelogโ€‹

  • fix(runtime): discard cached logical plans when a hot reload replaces a catalog (fixes #13910) by @claudespice in #13914
  • fix(acceleration): let a schema repair correct a checkpoint without resetting the freshness clock (fixes #13817) by @claudespice in #13894
  • fix(search): filter a chunked Elasticsearch delete on a field that can match the key (fixes #13714) by @claudespice in #13926
  • fix(search): classify a partially non-finite embedding as unindexable on every backend (fixes #13872) by @claudespice in #13902
  • fix: stabilize GitHub tests and bound GraphQL registration (fixes #13762) by @lukekim in #13939
  • fix(postgres): decode versioned JSONB binary replication values by @phillipleblanc in #13962
  • docs: require a reviewed Enhancement before any user-facing surface changes by @lukekim in #13970
  • ci: upgrade spiceio setup action to v0.9.0 by @lukekim in #13971
  • fix(postgres): preserve microseconds in timestamp writeback by @phillipleblanc in #13963
  • feat(hash-index): verify the bloom filter's block index with Verus by @lukekim in #13777
  • build(deps-dev): bump js-yaml by @dependabot in #13989
  • docs: release notes for v2.3.0 by @bjchambers in #13999
  • fix(ci): drop the dangling substrait-compliance submodule pointer by @bjchambers in #14002
  • fix(cayenne): release the keyset bytes an abandoned PK checkout accounted (fixes #13668) by @grokspice in #13925
  • fix(caching): keep a declared key from disabling eviction and stranding stale rows (fixes #13976) by @bjchambers in #13992
  • Add Substrait compliance harness (IBM TPC-H Mode A + FlightSQL Mode B stub) by @lukekim in #13879
  • ci: skip DynamoDB TPC-H benches in OSS testoperator dispatch by @phillipleblanc in #14016
  • docs: update security support and roadmap after v2.3.0 by @phillipleblanc in #14024
  • chore: post v2.3.0 release housekeeping by @bjchambers in #13969
  • test(adbc): guard BigQuery corpus offline and in release gate by @phillipleblanc in #14017
  • fix(duckdb): deny the regexp built-ins DuckDB cannot answer faithfully (fixes #13809) by @claudespice in #13871
  • fix: Update tpch benchmark snapshots for federated/adbc[bigquery].yaml by @app/github-actions in #13984
  • fix(search): drop the chunks a shortened row no longer produces from a chunked index (refs #13717) by @claudespice in #13960
  • Reduce Cayenne allocations during primary-key validation and filtering by @lukekim in #14009
  • test(forks): guard seven fork patches that had no repo-side test by @krinart in #13996
  • fix(deps): bump arrow-rs to correctly-rounded Decimalโ†’Float cast (closes #13978) by @Jeadie in #14012
  • perf(vortex): defer projection setup on filtered scans until the filter resolves by @bjchambers in #14035
  • endgame: include spiceai/skills versioned release by @lukekim in #14031
  • fix(caching): partition doomed entries at the survivor cutoff so eviction converges (closes #13994) by @Jeadie in #14021
  • fix(deps): bump arrow-rs fork pin for Decimal->Float rounding fix by @Jeadie in #14049
  • Fix subqueries with use_source acceleration by @phillipleblanc in #14022
  • fix(cayenne): make DELETE, UPDATE and INSERT work on a mode: memory acceleration (fixes #12008) by @bjchambers in #14047
  • fix: clarify OpenDAL S3 retry warnings by @lukekim in #14040
  • test(s3): run the parquet-overwrite fixtures on RustFS by @bjchambers in #14067
  • feat(cayenne): materialize multi-reference CTEs on the query path by @lukekim in #13918
  • perf(vortex): answer a constant IN list by probing a set, and falsify it by interval by @bjchambers in #14061
  • perf(vortex): skip a scan split whose zones cannot satisfy the filter by @peasee in #14064
  • fix(arrow): report an exact row count from the indexed point-lookup scan by @krinart in #13972
  • fix(cayenne): apply sort_columns with refresh_mode: full by @peasee in #14063
  • fix(smb): pad an empty CREATE buffer so Samba lists the share root (fixes #13293) by @grokspice in #14050
  • perf(cache): key the logical-plan cache on SQL text, not parameter values by @bjchambers in #14069
  • fix: harden HuggingFace E2E chat against slow Metal generation by @lukekim in #14072
  • fix(cayenne): move accelerator filesystem I/O off Tokio workers by @lukekim in #14073
  • test(chbench): enable CTE materialization and IVM on mysql/postgres adaptive HTAP by @lukekim in #14070
  • ci: run Substrait Mode A TPC-H on pull requests and the merge queue by @lukekim in #14071
  • feat(mcp): support MCP specification 2026-07-28 (dual-era) by @lukekim in #14043
  • fix(ci): call a linker that died of a signal an infrastructure failure, not a check failure (fixes #13614) by @grokspice in #14044
  • fix: Provide temporary directory in docker images by @Jeadie in #14089
  • fix: restore OSS installer, CLI and test workflow coverage by @phillipleblanc in #14025
  • Delete v2.2.0.md by @Jeadie in #14095
  • docs(release): add v2.3.1 release notes by @phillipleblanc in #14094
  • fix(test): allow DELETE in the CORS allow-methods assertion by @claudespice in #14098
  • fix(ci): stop install-protoc unzipping into a shared ~/.local by @lukekim in #14097
  • feat(connectors): add ORC listing format via in-repo FileFormat by @lukekim in #14075
  • fix(cayenne): run snapshot bootstrap check before opening the metastore by @Jeadie in #14093
  • feat(cache): verify the results-cache namespace prefix with Verus by @lukekim in #14074
  • feat(cloud-connect): add a GetDatasets command that answers the /v1/datasets document (refs #13369) by @grokspice in #14051
  • Suppress Cayenne startup logs when no Cayenne dataset is configured by @Jeadie in #14042
  • fix(cache): re-bind parameter values when revalidating a stale result (fixes #14099) by @bjchambers in #14100
  • fix(cluster): support distributed HTTP scans by @phillipleblanc in #14108
  • fix(bigquery): keep ILIKE evaluation local by @phillipleblanc in #14110
  • docs: update security support for v2.3.1 by @phillipleblanc in #14117
  • feat(cayenne): reuse ScanView until write, lag only for read-only CDC by @lukekim in #14055
  • fix(deps): remediate open Dependabot alerts by @phillipleblanc in #14111
  • fix(testoperator): validate results in every scale factor 1 TPC-H, TPC-DS and ClickBench benchmark by @lukekim in #14119
  • fix(cayenne): reject ambiguous metastore paths by @phillipleblanc in #14130
  • perf: serve results-cache hits where the request arrives and cut per-hit overhead by @lukekim in #14103
  • test(runtime): record query previews in the management export test by @lukekim in #14155
  • ci: upgrade spiceio setup action to v0.11.0 by @lukekim in #14152
  • feat(caching): Make caching_stale_if_error RFC-5861 compliant (with stale-if-error header) by @Jeadie in #14134
  • perf(runtime-table): defer cache-eviction key extraction to entries a delete actually names by @Jeadie in #14138
  • fix(runtime): report the acceleration.ready_state deprecation once per component (fixes #13749) by @claudespice in #14006
  • fix(runtime): write the inferred Arrow sort order under the prefixed key its validation accepts (fixes #14023) by @claudespice in #14032
  • fix(connectors): Fix JSON/Orca files using metadata columns by @Jeadie in #14115
  • feat(cayenne): build secondary indexes from indexes in file and memory mode by @phillipleblanc in #14149
  • fix(cayenne): round-trip decimal, binary, and time stats and drop them on scale change by @lukekim in #14139
  • feat(cayenne): cluster warm and datalake tiers, and write full refreshes as key-range files by @lukekim in #14124
  • fix(cayenne): compile the cold-tier pruning test and backtick a doc literal by @lukekim in #14175
  • fix(cayenne): make the crates own targets lint and compile by @phillipleblanc in #14200
  • test(forks): guard five more fork patches, and drop a row that is not fork state by @krinart in #14015
  • perf(cache): promote encoded SQL results to raw after the second decode by @lukekim in #14199
  • fix(turso): build a dictionary column directly so a dictionary over a list, map or boolean value reads back (fixes #13033) by @grokspice in #14181
  • fix(ci): probe the macOS toolchain before reaching for brew in the release builds by @grokspice in #14203
  • fix(ci): skip Metal kernel precompilation in the macOS release build by @grokspice in #14204
  • fix(github): paginate nested GraphQL connections and pace to GitHub's rate limits by @lukekim in #14179
  • fix(vortex): stop an IN list holding a NULL from panicking the scan by @krinart in #14163
  • bench(cayenne): use std::hint::black_box in the clustering bench by @lukekim in #14129
  • fix(runtime-table): stop rebuilding SessionContext on every cache fetch by @Jeadie in #14141
  • test(chbench): cluster order_line, oorder and customer on the adaptive HTAP arms by @lukekim in #14192
  • fix(runtime): count a first load as still loading in the Dataset load summary (fixes #13974) by @claudespice in #14020
  • fix(duckdb): push regexp_count down again at a rendering that counts as the kernel does (fixes #13870) by @claudespice in #14153
  • build: lint and test the sign-off under the same profile as the merge queue by @lukekim in #14180
  • ci: require the Verus proofs in the merge queue as one check by @lukekim in #14189
  • ci: run the longest macOS jobs on their own runner pool by @lukekim in #14229
  • fix(cayenne): build the DELETE sink inside the execution-time write lock (fixes #13828) by @claudespice in #14218
  • build(deps): bump the github-actions-dependencies group across 1 directory with 7 updates by @dependabot in #14231
  • fix(cluster): decide what a Flight message carries by its IPC header, not its body length (refs #13737) by @claudespice in #14212
  • ci: run Mode A TPC-H on merge queue and trunk/release push only by @lukekim in #14247
  • chore(deps): bump spiceai/duckdb-rs to 76655d2f by @lukekim in #14246
  • ci: stop exporting empty AWS and DuckLake endpoints to the schema test by @phillipleblanc in #14194
  • fix(runtime): reload a localpod dataset when the dataset it reads through is reloaded (fixes #3288) by @claudespice in #14208
  • build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14255
  • fix(ci): resolve Homebrew prefix when brew is the spice flock wrapper by @lukekim in #14210
  • chore(deps): raise the datafusion-table-providers pin to include the NUMERIC result-column fix by @phillipleblanc in #14254
  • ci: align the DuckLake bootstrap with the embedded DuckDB, wait for Databricks startup, and stop dispatching legs that cannot pass by @phillipleblanc in #14250
  • fix(runtime): install the Spice function deny-list on the PostgreSQL catalog connector (refs #13664) by @claudespice in #14225
  • perf(cayenne): share inline-cache view entries by Arc instead of cloning them per scan by @krinart in #14191
  • build(deps): bump aws-actions/configure-aws-credentials by @dependabot in #14256
  • ci: stop dispatching the indexed turso TPC-H SF1 tests by @phillipleblanc in #14252
  • fix(catalog): keep the tables registered under an existing schema by @phillipleblanc in #14193
  • feat(cache): Spice sharded cache as the sole LruCache engine by @lukekim in #14206
  • ci: lint GitHub Actions definitions with actionlint, and fix the 73 findings it surfaced by @grokspice in #14223
  • perf(cache): tighten the Raw SQL results-cache serve path by @lukekim in #14205
  • ci: stop triggering the CUDA build on pull requests by @lukekim in #14267
  • ci: run CodeQL on pull requests and the merge queue by @lukekim in #14269
  • Release 2.3.2 release notes by @krinart in #14271
  • chore: make AGENTS.md the canonical agent instructions by @lukekim in #14237
  • ci: run CodeQL Analyze on spiceai-dev-runners by @lukekim in #14281
  • feat: TypeSafe Jev System One evaluation provider by @lukekim in #14215
  • fix(cayenne): tag file statistics bounds as the column's Arrow type (fixes #14280) by @phillipleblanc in #14283
  • ci: install spiceio when the runner has no gh (refs #14233) by @lukekim in #14288
  • test(forks): 11 repo guards by @krinart in #14261
  • docs: add Spice.ai in Action manuscript and companion labs by @lukekim in #13965
  • test(duckdb): add an integration test for the index CTE materialization by @sgrebnov in #13885
  • fix(cayenne): coalesce inline writes into one batch by @sgrebnov in #14279
  • release: Update SECURITY.md and endgame template after 2.3.2 by @peasee in #14293
  • fix(cayenne): count each Arrow allocation once in the inline-cache gauge by @krinart in #14272
  • feat(s3): SQS event-driven changes for refresh_mode: changes by @lukekim in #14121
  • feat(caching): single-flight coalesce concurrent cache-miss fetches by @Jeadie in #14142
  • fix(caching): make caching_stale_if_error detect transient HTTP failures on real schemas by @krinart in #14161
  • Use Cayenne secondary indexes for dynamic join filters by @phillipleblanc in #14282
  • fix(ci): compare DynamoDB sets without relying on element order by @bjchambers in #14328
  • docs(release): Remove QA analytics step from endgame by @peasee in #14329
  • ci: run CodeQL on the merge queue and trunk, not on pull requests by @lukekim in #14290
  • docs(cayenne): reposition the reference, add query serving, re-audit against trunk by @lukekim in #14338
  • docs(endgame): update versioned docs release steps by @ewgenius in #14277
  • test(forks): guard the ballista per-task file-scan restriction by @krinart in #14292
  • fix(cli): surface the full error chain for spice chat connection failures by @krinart in #14289
  • fix(ci): degrade the incomplete-sign-off handler when the runner has no gh (refs #14234) by @claudespice in #14304
  • fix(runtime-table): serialize a direct write against acceleration snapshot creation (fixes #13548) by @claudespice in #14310
  • fix(duckdb): screen regexp_like and regexp_replace as regexp_count is screened (fixes #14148) by @claudespice in #14321
  • fix(http): honor retry budget without an extra origin request by @phillipleblanc in #14313
  • Fix SchemaCastScanExec's schema conversion in fn partition_statistics by @Jeadie in #14258
  • fix(cluster): recognise a Flight keepalive by its empty envelope, not by what its header declares (fixes #13737) by @claudespice in #14327
  • perf(http): defer zero-TTL acceleration lookup until origin failure by @phillipleblanc in #14302
  • fix(cayenne): keep a key visible when it is re-inserted over a stale-insert tombstone by @sgrebnov in #14312
  • perf(cayenne): read only key and filter columns in filtered key deletes by @sgrebnov in #14374
  • fix(cayenne): let small protected-snapshot merges run during a long large-tier merge by @sgrebnov in #14296
  • perf(cayenne): serve primary-key lookups on a freshly loaded table in ~1 ms by @lukekim in #14314
  • fix(cayenne): skip min/max statistics for nested columns (fixes #14368) by @sgrebnov in #14392
  • test(caching): cover SchemaCastScanExec statistics projection by name, retype, and SQL filter by @Jeadie in #14146
  • fix(turso): keep a quantified comparison out of Turso SQL (fixes #14041) by @grokspice in #14393
  • fix(udfs): declare the local_embed dev-dependency the embed tests need (fixes #13092) by @claudespice in #14377
  • perf(cayenne): batch metastore manifest rewrites, upsert in place, and run every write on one writer connection by @lukekim in #14369
  • fix(runtime-table): ignore zero-row batches in stale fallback by @phillipleblanc in #14331
  • Prune object-store file listing by _last_modified predicates by @Jeadie in #14265
  • Reading only partition or metadata columns needlessly scans all file contents by @Jeadie in #14116
  • fix(duckdb): keep a concat over a binary operand out of the federated plan (fixes #13915) by @claudespice in #14333
  • fix(ci): keep the sign-off attribution inside GitHub's 140-character status cap (fixes #14076) by @grokspice in #14390
  • fix(ci): expose a present-but-unlinked cc tool on macOS runners instead of routing it through brew install (fixes #13479) by @grokspice in #14387
  • ci: re-measure integration.yml's job bounds after the archive consolidation (fixes #13429) by @grokspice in #14388
  • fix(ci): start DuckLake's local MinIO from an image that is still published by @grokspice in #14399
  • test(runtime-table): make the metric-scraping refresh tests pass under cargo test by @claudespice in #14381
  • fix(federation): keep a correlated subquery predicate above a join of two sources (refs #8220) by @claudespice in #14372
  • fix(caching): snapshot staleness before the origin fetch for stale_if_error by @Jeadie in #14263
  • perf(snapshots): skip unchanged snapshot metadata with a conditional GET by @sgrebnov in #14409
  • fix(search): prune the rest of a key group from an Elasticsearch chunked index (refs #13717) by @claudespice in #14320
  • fix(bench): derive MySQL's empty-field NULL handling from the column type (refs #13152) by @claudespice in #14345
  • fix(graphql): debit a LIMIT by the rows a page returned, not the declared page size (fixes #14308) by @claudespice in #14353
  • fix(ci): fit retention_oom's retry budget inside its workflow step, and guard the coupling (fixes #13512) by @grokspice in #14389
  • docs: say plainly what Spice is, refresh the README for v2.3, and promote connector statuses by @lukekim in #14410
  • fix(ci): pin, checksum and retry the oha download in the E2E graceful-shutdown jobs by @grokspice in #14418
  • fix(snapshots): resolve snapshot entries relative to the metadata location (#14425) by @sgrebnov in #14426
  • Lower the GitHub connector default concurrency limit to 4 by @lukekim in #14431
  • build(deps): bump nvidia/cuda in the docker-dependencies group by @dependabot in #14439
  • build(deps): bump the aws-sdk group with 3 updates by @dependabot in #14440
  • build(deps): bump the github-actions-dependencies group across 1 directory with 5 updates by @dependabot in #14441
  • test(cayenne): bound the refused-build guard by builds, not by the host's speed by @grokspice in #14424
  • fix(cache): don't report an invalidation cancelled by runtime shutdown as a failure by @grokspice in #14417
  • ci: run remote sign-off on the spiceai-macos pool by @lukekim in #14449
  • ci: run CodeQL Analyze on spiceai-macos by @lukekim in #14442
  • fix(runtime): keep the built-in date_part so both weekday spellings agree (fixes #13920) by @claudespice in #14154
  • fix(cayenne): include the in-memory CDC tier when an overwrite or a retention pass covers the whole table by @lukekim in #14428
  • fix: keep NOT IN null-aware through the join reorder and the Cayenne sort-merge rewrite by @lukekim in #14429
  • fix(cayenne): discard a compaction whose snapshot an overwrite replaced mid-pass by @lukekim in #14432
  • fix(cayenne): stop maintained views, Vortex IN lists and dynamic-filter sharing from returning wrong rows by @lukekim in #14427
  • fix(cayenne): hide spilled rows on CDC upsert fallback by @bjchambers in #14416
  • fix(ci): skip the integration, ADBC, chDB and E2E gate jobs on pull requests instead of passing them (fixes #13841) by @grokspice in #14454
  • fix(cayenne): judge a filtered key delete's captured sources by the index captured with them (refs #13913) by @claudespice in #14455
  • fix(ci): clear the macos-15 image's openssl@1.1 symlink before installing MySQL by @lukekim in #14489
  • ci: produce CodeQL SARIF in the Analyze job on spiceai-macos by @lukekim in #14500
  • fix(cayenne): correctness fixes for the goal-driven adaptive controller, with a closed-loop simulation harness by @lukekim in #14443
  • fix(cdc): keep the newest source commit timestamp on a coalesced change batch by @lukekim in #14463
  • fix(cayenne): draw a sequence for a current-snapshot append so per-key OCC can order it (fixes #13685) by @claudespice in #14360
  • fix(http): name a configured model's load failure instead of reporting it not found (fixes #13303) by @claudespice in #14395
  • fix(duckdb): keep inferred source indexes off change-stream accelerations so upserts commit under concurrent reads (refs #13929) by @claudespice in #14396
  • fix(duckdb): keep a text cast over a binary operand out of the federated plan (fixes #14355) by @claudespice in #14448
  • fix(llms): honor tool_choice required and allowed_tools on mistral.rs-hosted models instead of panicking (fixes #14230) by @claudespice in #14460
  • test(runtime): run the load-error counter test in its own process (fixes #13085) by @claudespice in #14462
  • fix(runtime): stop a replaced dataset configuration's load from registering over the new one (fixes #1458) by @claudespice in #14367
  • fix(install): stop asking for sudo on a first install into a fresh HOME (fixes #14445) by @claudespice in #14495
  • perf(cayenne): serve primary-key point lookups in half the time by @lukekim in #14433
  • fix(deps): bump DataFusion for upstream fixes to wrong results from filter pushdown, simplification and planning by @lukekim in #14430
  • fix(runtime): size every internal DataFusion session from the CPU budget by @bjchambers in #14412
  • fix(runtime): serve nested, zoned and half-float columns from the Iceberg catalog API (fixes #4815) by @claudespice in #14480
  • fix(acceleration): keep cached results when a snapshot refresh finds no newer snapshot by @sgrebnov in #14497
  • fix(runtime): sleep cron tests to the next boundary, not one that already fired (fixes #13759) by @claudespice in #14474
  • fix(ci): derive every lint-rust guard make runs, whatever its separator or recipe layout (fixes #13783) by @claudespice in #14475
  • fix(cayenne): stop the small-file compaction of a position-mode PK table from deadlocking on its own write lock (fixes #14420) by @claudespice in #14481
  • fix(runtime-table): report refresh bytes for the rows each batch holds by @Jeadie in #14469
  • fix(spark,databricks): keep Spice-only functions out of SQL sent to Spark Connect and Databricks SQL Warehouse (refs #13664) by @claudespice in #14498
  • fix(vortex): size a cached footer by what it retains, not its serialized bytes (fixes #12917) by @claudespice in #14502
  • fix(cayenne,telemetry): Register accelerated sink dataset immediately from existing acceleration by @peasee in #13955
  • Update spicepod.yml by @Jeadie in #14533
  • fix(cayenne): keep a rewrite's count inexact when it retains a late protected snapshot (fixes #14383) by @claudespice in #14385
  • fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14408
  • build(rust): upgrade toolchain to 1.98.1 by @lukekim in #14560
  • fix(postgres): release a shared slot's hold on a table its publication cannot drop (fixes #13032) by @claudespice in #14527
  • fix(cayenne): size the build side before sort-merging a non-Cayenne outer join by @krinart in #14520
  • Update DF Upgrade template by @krinart in #14526
  • fix(ci): wait for the refresh to invalidate the results cache, not a fixed 3s by @grokspice in #14598
  • fix(deps): move the DataFusion pin past the four unparser fixes, and guard each of them (fixes #13022) by @grokspice in #14570
  • fix(dynamodb, cosmosdb, mongodb): push filters down only where the source evaluates them as SQL does by @lukekim in #14419
  • feat(snapshots): serve a dataset from published snapshots with from: s3://โ€ฆ and file_format: snapshot by @lukekim in #14529
  • perf(cayenne): keep the per-shard PK index across checkpoint flushes and back off futile bakes; fix q17 and two wrong-results races (fixes #14235) by @lukekim in #14555
  • test(chbench): lower the MySQL adaptive pods' CDC linger to 500 ms by @lukekim in #14615
  • fix(graphql): show the parse-failure bytes in JSON decode previews by @Jeadie in #14530
  • Defer source-first cache fallback planning by @phillipleblanc in #14556
  • perf(cayenne): order whole-table rewrites per scan partition, then merge by @bjchambers in #14488
  • fix(ci): tolerate the parquet-rename race's third DuckDB error in the E2E log scan by @grokspice in #14541
  • fix(graphql): retry an inferred credential refusal instead of failing the refresh by @Jeadie in #14539
  • fix(cayenne): serve a widened table's files from their persisted statistics (refs #13829) by @claudespice in #14316
  • fix(acceleration): checkpoint DuckDB's write-ahead log before a snapshot copies its file (fixes #13912) by @claudespice in #14325
  • fix(ci): let E2E cleanup run when the job failed before creating its working directory by @grokspice in #14553
  • Replace unmaintained backoff with a workspace crate by @phillipleblanc in #14621
  • test: read Expect stand-ins with the installed shell by @phillipleblanc in #14634
  • fix(models): apply a tool_choice that forces a call to one round of the tool-use loop (fixes #14459) by @claudespice in #14635
  • fix(runtime): don't end a tool-use stream on a stale tool_calls finish (fixes #13309) by @claudespice in #14548
  • fix(smb): name listed locations under the share so directory datasets load (fixes #14060) by @claudespice in #14549
  • fix(models): name a configured model's load failure in ai() and streaming nsql (fixes #14394) by @claudespice in #14632
  • test(chbench): remove the mysql-cayenne[file]-adaptive-split HTAP arm by @sgrebnov in #14620
  • fix(cayenne): record sharded CDC keys in the table-wide PK index by @lukekim in #14603
  • fix(sqlite): keep TRY_CAST and every cast SQLite answers differently out of the federated plan (fixes #14398) by @claudespice in #14496
  • ci: move remaining Ubuntu 22.04 references to 24.04 by @lukekim in #14677
  • fix(snapshots): retry a failed snapshot attempt with backoff by @sgrebnov in #14676
  • fix(snapshots): large snapshots no longer fail bootstrap on slow connection by @sgrebnov in #14571
  • fix(refresh): start a refresh's source scan only when the sink reads it, so the full-text index keeps its rows (fixes #14619) by @claudespice in #14663
  • test(cayenne): compare suite answers cell by cell against SQLite and chDB by @lukekim in #14473
  • Evaluate any chat model via /v1/evaluate by @lukekim in #14568
  • ci: quarantine the management API integration schedule until its dev OAuth client is reactivated (refs #12376) by @grokspice in #14684
  • ci: stop scheduled Testoperator Ballista benchmarks by @phillipleblanc in #14682
  • fix(cayenne): stop snapshotting datasets with a datalake tier, whose restored copies deleted each other's files by @lukekim in #14583
  • fix(cayenne): count each Arrow allocation once for the mem-tier limit and checkpoint write sizing by @krinart in #14675
  • fix(ci): classify a sign-off that reached its step budget as signalled, so an expiry publishes no verdict (fixes #13843) by @grokspice in #14685
  • fix(runtime): retry a dataset whose source is unreachable at startup (fixes #14609) by @bjchambers in #14623
  • fix(cayenne): archive only the snapshot directories a dataset snapshot references (fixes #14605) by @sgrebnov in #14627
  • ci: route trunk-push macOS builds to the standard pool by @phillipleblanc in #14645
  • ci: install strip for the retention OOM regression test by @phillipleblanc in #14701
  • fix(smb): serve every share on a host from the one store registered for it (fixes #14550) by @grokspice in #14689
  • fix(runtime): discard cached plans and results when a dataset is unloaded (refs #14251) by @claudespice in #14365
  • feat(snapshots): reload snapshot-mode datasets from S3 event notifications (SQS) by @lukekim in #14335
  • fix(cayenne): report runs per size tier when protected-snapshot compaction declines (fixes #13622) by @claudespice in #14476
  • fix(cayenne): re-run a post-write compaction pass a concurrent append asked for (fixes #13906) by @claudespice in #14479
  • fix(http): keep the request URL out of HTTP connector errors (fixes #13534) by @claudespice in #14490
  • fix(runtime): load every localpod dataset that reads from one parent at startup (fixes #13087) by @claudespice in #14532
  • fix(build): share one cargo metadata helper across the lint guards, so a broken cargo is never a violation (fixes #13121) by @claudespice in #14534
  • fix(duckdb): push regexp_replace down again for a one-digit group reference, and refresh the Postgres ClickBench q35 plan (refs #13966) by @claudespice in #14552
  • fix(duckdb): keep a cast into binary out of the federated plan (refs #14397) by @claudespice in #14633
  • fix(runtime): keep finished async query jobs finished, and delete expired ones by @lukekim in #14585
  • fix(federation): keep DataFusion's cast built-ins out of every federated plan (fixes #14444) by @claudespice in #14484
  • test(cayenne): cover a swapped join filter through the sort-merge rewrite end to end (refs #14235) by @claudespice in #14551
  • fix(runtime): decide every Spark/built-in function collision by name, and refuse an undecided one (fixes #14361) by @grokspice in #14400
  • fix(ci): run every bin target's unit tests in the sign-off gate, and give spice connect its own --cloud-region refusal (fixes #13426) by @grokspice in #14406
  • test(data_components): compile the federation unparser guards in every scoped test run (fixes #13625) by @grokspice in #14447
  • fix(runtime-component): skip an inferred Cayenne index on a floating-point column instead of failing the load (fixes #14590) by @grokspice in #14599
  • fix(cache): serve stale cached results during frequent table updates (stale_while_revalidate_ttl) by @sgrebnov in #14708
  • Update Turso to 0.8.1 and retry write conflicts a metastore statement raises by @lukekim in #14680
  • fix(snapshots): stop a snapshot dataset's load when a reload replaces it by @lukekim in #14673
  • fix(cayenne): hand non-partition filters to every partition scan (fixes #12959) by @claudespice in #14370
  • fix(postgres-accel): resolve secret references in the sidecar connection parameters (fixes #13296) by @claudespice in #14544
  • fix(cayenne): drain the in-memory CDC tier before a full rewrite scans it (fixes #14450) by @claudespice in #14486
  • feat(caching): warm SQL results cache on first refresh from persisted plan shapes by @lukekim in #14178
  • fix(ci): retry only the container startup in the zero-retry MySQL CDC tests by @grokspice in #14727
  • feat(key-index): immutable secondary index runs over compound Arrow keys by @bjchambers in #14592
  • fix(federation): keep a fractional-to-integer cast out of the plans pushed to DuckDB, PostgreSQL, MySQL and BigQuery (fixes #14482) by @grokspice in #14601
  • fix(testoperator): give accelerated bench configs a 300s ready_wait and name unready datasets on timeout (fixes #13973) by @claudespice in #14728
  • fix(federation): keep arrow_typeof and its plan-introspection siblings local on every backend (fixes #14334) by @grokspice in #14695
  • test(bench): refresh the tpch_q16 explain snapshots for the null-aware NOT IN plan (fixes #13977) by @claudespice in #14731
  • ci: build trunk-push macOS legs on spiceai-macos-large again, keeping one-at-a-time coalescing by @grokspice in #14735
  • ci(verus): derive the verified crates from cargo metadata; pin vstd once by @bjchambers in #14709
  • docs: raise the test and evidence bar: differential first, exact assertions, performance always measured by @lukekim in #14723
  • test(vortex): wait for a dropped segment cache to be freed before asserting it is gone (fixes #13295) by @claudespice in #14470
  • ci(integration): report each integration part's own result in its required check by @grokspice in #14743
  • fix(cayenne): back off futile bakes under a violated query goal too by @lukekim in #14736
  • feat(object_store_occ): add transactional WAL and MVCC snapshots by @lukekim in #14732
  • fix(snapshots): bootstrap readers from late snapshot publications by @phillipleblanc in #14468
  • fix(cayenne): apply retention_sql to a mode: memory acceleration (fixes #14045) by @claudespice in #14344
  • fix(duckdb, sqlite): keep the first copy of a key a write repeats under on_conflict: drop (fixes #14629) by @claudespice in #14748
  • fix(cache): release a removed dataset's memory when the plan cache discards its plans (fixes #14251) by @claudespice in #14760
  • build(deps): bump the aws-sdk group with 2 updates by @dependabot in #14764
  • fix(llms): explain an Anthropic model's refusal of a sampling control instead of passing the bare 400 through (fixes #13564) by @grokspice in #14700
  • build(deps-dev): bump dompurify by @dependabot in #14613
  • fix(snapshots): allow cayenne_file_path and cayenne_metadata_dir on file_format: snapshot datasets (fixes #14696) by @sgrebnov in #14698
  • test(snapshots): make snapshot integration tests robust under parallel CI runs by @sgrebnov in #14765
  • feat(cayenne): keep the secondary index current across every write by @bjchambers in #14593
  • fix: Update tpch benchmark snapshots for accelerated/file[parquet]-duckdb[memory].yaml by @app/github-actions in #14674
  • Isolate Docker integration fixtures across concurrent processes by @phillipleblanc in #14639
  • ci: run Substrait Mode A TPC-H at SF 1 on spiceai-macos in every merge-queue entry by @lukekim in #14522
  • ci: stop repeating scheduled benchmarks: one source per accelerator, hosted sources weekly, a release commit once by @lukekim in #14681
  • build(deps): bump nvidia/cuda by @dependabot in #14763
  • fix(http): stop a non-2xx response body replacing an accelerated table's rows (fixes #13515) by @claudespice in #13538
  • fix(chat-api): answer a tool-call-only assistant turn instead of panicking (fixes #13207) by @grokspice in #14232
  • fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) by @claudespice in #14670
  • fix(spiceai, duckdb, sharepoint): name the Spicepod key a missing-parameter error asks for (refs #14446) by @claudespice in #14769
  • Add table-bound ChangeSink ownership and backends by @phillipleblanc in #14704
  • fix(ci): serialize rustup installs on shared macOS runners by @lukekim in #14717
  • fix(cayenne): keep the per-shard PK index when the table-wide index is discarded by @lukekim in #14604
  • Upgrade to DataFusion 55.1 and Arrow 59.3 by @krinart in #14612
  • Update openapi.json by @app/github-actions in #14742
  • build(deps): bump rustls from 0.23.40 to 0.23.45 by @dependabot in #14773
  • Upgrade to iceberg-rust to 0.11.0 and DF 55.2 by @sgrebnov in #14771
  • fix(cayenne): write an append into an empty unkeyed table as one load by @phillipleblanc in #14784
  • fix(mysql_replication): don't re-send a live member's delivered commits after a reconnect by @lukekim in #14751
  • ci(codeql): don't fail the SARIF upload when the merge queue already deleted its ref by @grokspice in #14793
  • fix(acceleration): accept time_format iso8601 when the time_column is already a timestamp by @lukekim in #14777
  • fix(cache): store SQL results that go stale mid-query by @lukekim in #14710
  • fix(cayenne): serve filtered and global maintained aggregates by @lukekim in #14761
  • perf(cayenne): build the checkpoint's tombstone union after releasing the capture locks by @lukekim in #14759
  • ci: keep no artifacts or caches from PR and merge-queue checks, and run gating macOS jobs on spiceai-macos-large by @lukekim in #14804
  • fix(cayenne): keep maintenance from deleting files a snapshot is archiving by @sgrebnov in #14789
  • fix(snapshots): never overwrite shared snapshot metadata, and publish more than once to a file:// location by @lukekim in #14582
  • perf(cdc): build deferred change rows in the reader while the apply runs by @lukekim in #14746
  • fix: Update tpch benchmark snapshots for accelerated/on_zero_results/file[parquet]-cayenne[file]-on_zero_results.yaml by @app/github-actions in #14786
  • test(testoperator): add a cold-start time-to-ready regression test by @phillipleblanc in #14785
  • test(testoperator): make append tests more robust by @sgrebnov in #14817
  • fix(runtime,data_components): wait for change-data-capture sources to record their positions before shutdown closes the accelerations (fixes #14523) by @grokspice in #14702
  • fix(graphql): detect non-JSON responses and treat gateway errors as retryable by @lukekim in #14781
  • fix(runtime): require keys for s3_auth: key, honor gs:// state location params, and leave newer rate-control state alone by @lukekim in #14584
  • build(deps): bump hickory-resolver from 0.26.1 to 0.26.2 by @dependabot in #14774
  • fix(listing): skip zero-byte objects (S3 folder markers) in format-selected listings by @sgrebnov in #14822
  • Avoid unnecessary repartitioning in indexed dynamic-filter joins by @bjchambers in #14799
  • feat(cayenne): keep one row per primary key automatically and deprecate on_conflict by @bjchambers in #14726
  • fix(object-store): stop reading a response body at its first error by @phillipleblanc in #14831
  • fix(ci): use byte-order collation in the benchmark Postgres container (refs #14815) by @sgrebnov in #14836
  • fix(listing): keep a partition predicate as a residual filter on the _location fast path by @grokspice in #14790
  • fix(ci): search only the unixodbc keg, not all of Homebrew's lib, in the macOS release builds by @grokspice in #14846
  • fix(federation): keep date, timestamp and interval values local on SQLite reached through ADBC or ODBC (fixes #14753) by @claudespice in #14840
  • chore: pin the fork at spiceai/datafusion#249 and guard the DataFusion fixes backported from 54 by @krinart in #14800
  • Revert "fix(sqlite, duckdb): keep a decimal AVG, and on SQLite a decimal SUM, out of the federated plan (fixes #14492) (#14670)" by @krinart in #14825
  • perf(runtime): skip EnsureRequirements while planning a point lookup by @lukekim in #14807
  • fix(testoperator): report CH-benCH queries with no rows on either side as vacuous, not as matches by @lukekim in #14810
  • feat(acceleration): make Cayenne the default acceleration engine by @phillipleblanc in #14837
  • fix(cayenne): keep a join's LIMIT when the oversized-join rewrite makes it a sort-merge join by @lukekim in #14805
  • test(bigquery): exclude subtrees an empty join build side never runs from the corpus job count (fixes #14848) by @sgrebnov in #14849
  • fix(cayenne): fail contradictory write settings once, and order NULL times below every time by @bjchambers in #14847
  • fix(runtime): serve an existing acceleration while its source is unavailable (fixes #14610) by @bjchambers in #14624
  • Prune object-store file listing by metadata columns (#14264) by @Jeadie in #14303
  • fix(vortex): fold partition values into the file-pruning predicate; test null-equal joins under mode:file by @lukekim in #14797
  • feat(rate-control): adaptive, bounded and cluster-coordinated HTTP rate controls by @Jeadie in #14143
  • fix(sql): align date_part('dow') with EXTRACT(dow) Sunday=0 by @lukekim in #14796
  • fix(runtime): carry API-key principal into MCP tools/call by @lukekim in #14828
  • perf(cayenne): hold maintained-aggregate retraction state in a compact shared index by @lukekim in #14762
  • fix(deps): keep a hash join's order once it has a dynamic filter, so vector_search plans again by @Jeadie in #14857
  • fix(cayenne): fence schema statistics and control maintenance tests by @bjchambers in #14856
  • fix(github): retry nested GraphQL pages in place and fail incomplete nested connections by @sgrebnov in #14862
  • fix(cayenne): load an append dataset that has retention_sql and orders versions by time by @phillipleblanc in #14878
  • fix(cdc): Cayenne replication needs only a primary key, not on_conflict by @phillipleblanc in #14880
  • feat(connector-huggingface): Hugging Face datasets data connector (hf://datasets/...) by @lukekim in #14877
  • fix(iceberg): Iceberg REST clients read the real table or get an error, never an empty one by @lukekim in #14588

Full Changelog: https://github.com/spiceai/spiceai/compare/v2.3.2...v2.4.0-rc.1

Spice v2.3.2 (Sep 22, 2026)

ยท 12 min read
William Croxson
Member of Technical Staff at Spice AI

Spice v2.3.2 is now available! โšก

Spice v2.3.2 makes point lookups and repeated queries faster with Spice Cayenne. SQL results-cache improvements apply to every cached query, regardless of its data source or accelerator. Across two benchmark rounds, the time for a small cache hit fell by 50-52% over HTTP and 44-48% over Flight SQL. For small cache hits, CPU time per server request fell by 58-61% over HTTP and 49-51% over Flight SQL.

Highlights in v2.3.2 include:

  • Indexed Point Lookups โ€” Cayenne now honors indexes, so a lookup reads the matching rows instead of scanning
  • Cayenne Data Layout โ€” cayenne_cluster_by groups related rows together across storage tiers
  • Faster Queries After a Full Refresh โ€” refreshed tables are written so that filtered queries read fewer files
  • Faster Cached Responses โ€” these improvements apply to every query in the SQL results cache, regardless of its data source or accelerator. For small results, cache-hit time fell by 44-52%
  • Faster Repeat Queries โ€” Cayenne reuses its prepared view of a table until the data changes
  • Localpod Fix โ€” a localpod dataset no longer serves data its source has replaced
  • Catalog Fix โ€” tables no longer go missing when datasets load at the same time
  • Safer Cayenne Configuration โ€” an ambiguous metastore configuration is reported instead of appearing as lost data
  • Distributed HTTP Queries โ€” async distributed queries can read HTTP datasets

What's New in v2.3.2โ€‹

Faster Point Lookups with Cayenne Indexesโ€‹

Every other accelerator already accepted indexes, and Cayenne logged that it ignored them. Cayenne now builds an index for each entry, in both mode: file and mode: memory:

acceleration:
engine: cayenne
mode: file # or memory
indexes:
'(TenantId, ServiceId)': enabled

A query that pins every column of an index to a value, such as WHERE TenantId = 7 AND ServiceId = 'a', now reads the matching rows directly. Without an index, Cayenne can only skip files whose minimum and maximum values rule the key out, which rarely helps when related rows are spread across the table โ€” a lookup on a 5.4M-row test dataset had to open about half its files for nearly every key.

Indexes only reduce what a query reads, so results are identical either way. A query that uses a range, an IN list, an OR, or a cast on the indexed column reads the table as before. Index definitions are not stored with the table, so adding or removing one takes effect the next time the dataset loads. Floating-point columns cannot be indexed and are reported at load time.

EXPLAIN shows whether a query used an index, and the cayenne_lookup_index_probe_total metric counts lookups by outcome.

Control How Cayenne Lays Out Dataโ€‹

cayenne_cluster_by stores rows with similar values near each other, so a filtered query reads fewer files. It applies to every storage tier, replacing cayenne_datalake_clustering_columns, which affected only the coldest tier.

acceleration:
engine: cayenne
params:
cayenne_cluster_by: 'tenant_id, event_time'

CREATE TABLE ... CLUSTER BY (column, ...) is also supported, and several cases it previously rejected โ€” a single column in parentheses, and names whose capitalization differs from the column definition โ€” now work. A column name that does not exist is reported when the dataset loads rather than failing later.

Faster Queries After a Full Refreshโ€‹

A full refresh previously spread each key across every file it wrote, so a filtered query had to open all of them even when it wanted a single row. A refreshed table is now written so that each file holds a distinct range of the data, and queries filtering on that range read only the files that can match.

This needs no configuration and applies to any dataset whose accelerated table is replaced by a refresh. Datasets that already set cayenne_sort_columns or cayenne_cluster_by keep their existing layout, and the very first load is unchanged.

Faster Cached Query Responsesโ€‹

Before this release, the runtime planned a query before it checked the SQL results cache. It also repeated other work that a cache hit does not need. The runtime now checks the cache first and returns cached answers directly. These improvements apply to every cached query, regardless of its data source or accelerator.

Across rounds of benchmarking queries that returned a seven-row GROUP BY result, cache-hit time over HTTP fell by 50-52%. Over Flight SQL, cache-hit time fell by 44-48%. In one HTTP round, cache-hit time fell from 63.8 ยตs to 30.8 ยตs. In one round over Flight SQL, it fell from 55.3 ยตs to 30.9 ยตs.

With task history enabled, the server used 125 ยตs of CPU time for a small HTTP hit before the change. It used 49 ยตs after the change. For a small result over Flight SQL, the server used 330 ยตs before the change. It used 168 ยตs after the change. These changes reduced CPU time by 61% and 49%, respectively.

For a wide result over HTTP, the server used 241 ยตs of CPU time before the change. It used 116 ยตs after the change. This is a 52% reduction. Wide results did not reduce the runtime's cache-hit time in those runs. Some measurements were slower.

Raw cache entries now share each stored batch with streaming JSON HTTP and Flight SQL responses. Buffered HTTP formats โ€” including CSV, plain text, vnd.* envelopes, and JSON with union columns โ€” still clone batches while encoding.

A stream benchmark with eight 20-column batches fell from 1.8722 ยตs to 226.23 ns. This is an 8.3x speedup. One 200-column batch fell from 2.0190 ยตs to 131.94 ns. This is a 15x speedup. These figures measure the construction and full consumption of the stream. They do not measure query latency from start to finish.

For compressed entries, the first read keeps the entry compressed. The second read promotes it to raw data when the raw size fits the cache. Later reads avoid decompression. In a cache benchmark, a later read took 17.1-17.8 ยตs across three payload sizes. A decode took 39.1-154.8 ยตs. Cached answers, expiry times, and cache settings are unchanged.

Faster Repeat Queries on Cayenneโ€‹

Cayenne prepares a view of a table's current contents before it can answer a query. It now reuses that preparation until the data actually changes, instead of rebuilding it, so repeated queries against a table that is not being written return faster and stop querying the metastore entirely once warm.

Datasets fed by continuous change data capture โ€” a cdc: or debezium: source, or refresh_mode: changes โ€” reuse the preparation for up to one second so a burst of incoming changes can share it. Datasets that accept writes always see their own writes immediately. No configuration is required.

Better Pruning for Decimal, Binary, and Time Columnsโ€‹

Cayenne did not record minimum and maximum values for decimal, binary, time, and 16-bit float columns, so queries filtering on them could not skip files and had to read more data than necessary. Those columns now carry the same statistics as every other type, and clustering on a decimal column works as intended.

After a change that widens a decimal column's scale, Cayenne discards the affected statistics and rebuilds them, so queries read a little more until that completes. Results are unaffected.

Localpod Datasets Stay in Sync with Their Sourceโ€‹

A localpod dataset reads through another dataset in the same Spicepod. When the source dataset was reloaded โ€” because its configuration changed โ€” the localpod dataset kept reading the replaced copy, so it answered with data the source no longer had, and both copies kept refreshing.

A localpod dataset now reloads whenever the dataset it reads through does, including through several levels of chaining, and its cached results are cleared at the same time. This also works when the source is named with its full path, such as localpod:spice.public.parent. Fixes #3288.

Catalog Tables No Longer Go Missingโ€‹

When two datasets in the same schema loaded at the same time, one could be silently discarded, and every later query against it failed with Table not found. Both datasets reported that they had loaded successfully, and which one went missing varied between restarts. Datasets that share a schema now always both register.

Cayenne Detects an Ambiguous Metastore at Startupโ€‹

Cayenne datasets that each set a different cayenne_file_path, without a shared cayenne_metadata_dir, could open the wrong metadata directory after a restart. Cayenne then started up empty even though the data was still on disk, which looks like a total loss of accelerated data.

This configuration is now rejected at startup, naming each dataset and path involved and linking to the documentation. If a Spicepod uses several Cayenne data paths, set the same cayenne_metadata_dir on each dataset before upgrading.

Distributed Queries over HTTP Datasetsโ€‹

An async distributed query submitted to /v1/queries failed if it read an unaccelerated HTTP dataset. These queries now run, with credentials, headers, and pagination behaving as they do for a non-distributed query. Fixes #14104.

Other Fixesโ€‹

  • Partitioned datasets: a full refresh of a partitioned dataset only replaced the partitions the new data reached, so rows deleted at the source stayed queryable, and a refresh that returned no rows changed nothing. Every partition is now replaced.
  • Iceberg write-through: a write to a partitioned dataset could deadlock against a refresh running at the same time, leaving both waiting.
  • BigQuery: a case-insensitive LIKE could fail the query or return the wrong rows, because BigQuery has no ILIKE. Spice now evaluates it locally. Ordinary LIKE is unchanged.
  • PostgreSQL: timestamps written back to PostgreSQL lost everything below the second. Microseconds are now preserved.

Dependency Updatesโ€‹

No crate versions changed in this release. DataFusion remains at v54.1.0, Arrow remains at v58.3.0, and Vortex remains at v0.79.0.

Spice updates two fork revisions: datafusion-table-providers for the BigQuery and PostgreSQL fixes above, and duckdb-rs so the bundled DuckDB builds against the macOS 27 SDK.

Contributorsโ€‹

Breaking Changesโ€‹

cayenne_datalake_clustering_columns is replaced by cayenne_cluster_by. Rename the parameter before upgrading. The new one groups data on every storage tier, not only the coldest:

acceleration:
engine: cayenne
params:
cayenne_cluster_by: 'tenant_id, event_time'

A dataset that sets both cayenne_sort_columns and a cluster key is now rejected when it loads. Remove cayenne_sort_columns to keep the cluster key.

Cayenne datasets using several data paths must share a metastore. If your Cayenne datasets set different cayenne_file_path values, set the same cayenne_metadata_dir on each one before upgrading. Spice now refuses to start on this configuration instead of risking an empty-looking acceleration:

acceleration:
engine: cayenne
params:
cayenne_file_path: /mnt/a/cayenne
cayenne_metadata_dir: /mnt/shared/cayenne-metadata

Cookbook Updatesโ€‹

No new cookbook recipes.

The Spice Cookbook includes more than 104 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v2.3.2, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:2.3.2 image:

docker pull spiceai/spiceai:2.3.2

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.3.2

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

  • fix(postgres): preserve microseconds in timestamp writeback by @phillipleblanc in #13963
  • feat(cayenne): reuse ScanView until write, lag only for read-only CDC by @lukekim in #14055
  • perf: serve results-cache hits where the request arrives and cut per-hit overhead by @lukekim in #14103
  • fix(cluster): support distributed HTTP scans by @phillipleblanc in #14108
  • fix(bigquery): keep ILIKE evaluation local by @phillipleblanc in #14110
  • feat(cayenne): cluster warm and datalake tiers, and write full refreshes as key-range files by @lukekim in #14124
  • fix(cayenne): reject ambiguous metastore paths by @phillipleblanc in #14130
  • fix(cayenne): round-trip decimal, binary, and time stats and drop them on scale change by @lukekim in #14139
  • feat(cayenne): build secondary indexes from indexes in file and memory mode by @phillipleblanc in #14149
  • fix(catalog): keep the tables registered under an existing schema by @phillipleblanc in #14193
  • ci: stop exporting empty AWS and DuckLake endpoints to the schema test by @phillipleblanc in #14194
  • perf(cache): promote encoded SQL results to raw after the second decode by @lukekim in #14199
  • perf(cache): tighten the Raw SQL results-cache serve path by @lukekim in #14205
  • fix(runtime): reload a localpod dataset when the dataset it reads through is reloaded (fixes #3288) by @claudespice in #14208

Full Changelog: https://github.com/spiceai/spiceai/compare/v2.3.1...v2.3.2

Spice v2.0-stable (Jun 5, 2026)

ยท 94 min read
Phillip LeBlanc
Co-Founder and CTO of Spice AI

53 releases since Spice 1.0-stable, Spice.ai OSS has reached the 2.0-stable milestone! ๐ŸŽ‰

Spice v2.0.0 is the next major release of Spice and a major milestone in the project's development, advancing Spice from a single-node engine into a distributed data and query platform built for enterprise AI agents. These agents need low-latency, governed access to data spread across many production systems, and because they generate their own queries autonomously, that access has to be sandboxed, observable, and able to absorb occasional heavy analytical queries without overwhelming the underlying systems. The release is headlined by multi-node distributed query, now generally available โ€” multi-active, highly-available, and object-store-native, built on Apache Ballista โ€” distributing both query execution and ingestion across executors with data-local routing and per-executor statistics for distributed join planning. Alongside it, the Spice Cayenne data accelerator is generally available, built on the Vortex compressed columnar format, with a high-throughput CDC write path, MERGE INTO, SQL-defined partitioning, inline writes, a dedicated compaction runtime, and write-path statistics for distributed join sizing. The engine also moves to DataFusion v52 with sort pushdown, a rewritten merge join, and dynamic filters, and the Spice CLI is rewritten in Rust as a single self-contained binary.

v2.0 also expands real-time and write-path capabilities across the platform: native CDC from MongoDB Change Streams and PostgreSQL WAL logical replication, durable Kafka CDC offsets, DML write-back for PostgreSQL, Snowflake, DynamoDB, Arrow, and DuckLake, DDL and MERGE INTO for Iceberg catalogs, mutual TLS across server endpoints and outbound connectors, HashiCorp Vault and Azure Key Vault secret stores, user-defined functions, hybrid search with Elasticsearch and DuckDB HNSW vector indexes, provider-aware LLM prompt caching, and the Responses API across all model providers.

Highlights in v2.0.0 include:โ€‹

  • Spice Cayenne (GA) โ€” generally available on the Vortex compressed columnar format, with WAL-staged writes, inline low-latency writes, fast-path CDC deletes, merge-on-read position deletes, composite & SQL-defined partitioning, MERGE INTO, dedicated compaction runtime, and join-sizing statistics maintained on the write path
  • Multi-Active HA Distributed Query (GA) โ€” multi-node distributed query built on Apache Ballista, with object-store-native clustering, dynamic cluster sizing, distributed ingestion, data-local query routing, per-executor table statistics for distributed join planning, and async queries via /v1/queries
  • Mutual TLS (mTLS) โ€” public mTLS for HTTP and Flight, TLS cert hot-reload, and mTLS client certificates for FlightSQL and Spice.ai connectors
  • Enterprise Authentication & Authorization โ€” OIDC bearer-token verification and Cedar-based authorization policy with per-principal row- and column-level filtering
  • New Secret Stores โ€” HashiCorp Vault and Azure Key Vault
  • CDC Sources โ€” native MongoDB Change Streams, PostgreSQL WAL logical replication, and durable Kafka CDC offsets โ€” no Debezium or Kafka middleware required
  • DML & DDL โ€” INSERT/UPDATE/DELETE write-back for PostgreSQL, Snowflake, DynamoDB, and Arrow; CREATE TABLE/DROP TABLE and MERGE INTO for Iceberg catalogs
  • User-Defined Functions โ€” SQL UDFs in spicepods, remote UDFs over HTTP, and optional geospatial ST_* UDFs
  • On-Demand Dataset Loading & Unified Query Cancellation โ€” faster startup and end-to-end cancellation across HTTP, Flight, FlightSQL, and MCP
  • Dynamic HTTP Connector โ€” OAuth2 refresh tokens, pagination, dynamic headers, subquery-driven parameters, and rate-control state persisted across restarts
  • Storage-Profile Accelerator Tuning & refresh_mode: snapshot โ€” storage-aware acceleration defaults and point-in-time snapshot acceleration
  • Search & Vectors โ€” Elasticsearch data connector with native hybrid search, DuckDB HNSW vector engine with a statically linked VSS extension, multi-vector MaxSim embeddings, and a rerank() UDTF
  • AI & LLM โ€” provider-aware prompt caching, Responses API across all providers, MCP Streamable HTTP transport, and a searchable LLM tool registry
  • New Data Connectors โ€” Elasticsearch (Alpha), GCS (Alpha), Azure Cosmos DB (Alpha), Git (RC), ADBC, DuckLake (Beta), and catalog connectors for PostgreSQL, MySQL, MSSQL, and Snowflake
  • Rust CLI โ€” single-binary spice CLI with spice query async REPL, shell completions, and --output=json
  • Dependency upgrades including DataFusion v52.5, DuckDB v1.5.3, Arrow v57.2, iceberg-rust v0.9.1, Turso v0.6.1, and Vortex v0.69

Spice v2.0 includes several breaking changes. Review the breaking changes section before upgrading.

Distribution Changesโ€‹

AI/ML support including local LLM/ML model and hosted LLM inference is now included in the default Spice build and image. The separate models build variant has been removed.

With models now included by default, the data-only distribution (without AI/ML support) is only published in nightly builds. Official production-ready data-only distributions are available exclusively through Spice Cloud and the Enterprise release.

A new Network Attached Storage (NAS) distribution with built-in SMB and NFS data connector support is also available in nightly builds and with Spice.ai Enterprise.

Distribution / VariantOpen SourceSpice CloudEnterprise
Defaultโœ…โœ…โœ…
DataNightly onlyโœ…โœ…
NAS (SMB + NFS)Nightly onlyโŒโœ…
Metal (macOS)โœ…โœ…โœ…
CUDA (Linux)Nightly onlyโœ…โœ…
Allocator variantsNightly onlyโœ…โœ…
ODBC connectorLocal build onlyโœ…โœ…

Native Windows builds are no longer provided; use WSL for local development. For more details, see the Distributions documentation.

What's New in v2.0.0โ€‹

Spice Cayenne Reaches General Availabilityโ€‹

The Spice Cayenne data accelerator is generally available in v2.0, with a major focus across the release candidates on write-path throughput, correctness, and distributed operation.

Write path & ingest:

  • Staged Append Writes: WAL-based staged append writes prevent partial writes and data loss on stream errors โ€” batches commit atomically.
  • Inline Writes: Small writes are serialized as Arrow IPC and committed directly into the Cayenne metastore, bypassing the staged Vortex write path for low-latency ingest. Inline upserts atomically rewrite existing inline rows, inline data stays query-visible via an in-memory union scan, and rows are checkpointed to Vortex when thresholds are reached. Inline writes now also proceed with pending deletions in flight, and inline flush caps scale with available memory and storage class.
  • Fast-Path CDC Deletes: DELETE statements whose filters identify primary keys directly โ€” including composite keys expressed as (k1, k2) IN ((...), (...)) โ€” skip the table scan entirely.
  • Merge-On-Read Position Deletes: Primary-key upsert tables use position deletes with memory-pool accounting, avoiding full-table rewrites on update-heavy workloads.
  • Resident Upsert Keysets: CDC upsert primary-key keysets stay resident between batches, avoiding per-batch full-table rebuilds.
  • CDC Sub-Batch Efficiency: Interleaved upsert/delete workloads produce fewer sub-batch splits, with last-write-wins deduplication applied within batches.
  • Dedicated Compaction Runtime: Background compaction runs on a dedicated thread pool with CDC pipelining and protected snapshots, isolating compaction work from query and ingest paths.

Query & planning:

  • Join Filter Propagation: Filters propagate across equi-join keys, with range fallback for large join filters and IN-list rewrites.
  • Write-Path Join-Sizing Statistics: Cayenne maintains live row counts and HyperLogLog-based distinct-value estimates on the write path, so distributed JoinSelection can correctly size joins without rescans.
  • Scan-Result Cache: A new scan-result cache accelerates hot reads, with parallel Vortex partition writes and lock-free deletion caches with bloom-prefiltered probes.

SQL & catalog:

  • MERGE INTO: Upsert-style MERGE INTO for Cayenne catalog tables, distributed across executors in cluster mode.
  • PARTITION BY in SQL: Define partitioning directly in CREATE TABLE ... PARTITION BY (...); metadata is persisted in the catalog and survives restarts.
  • Composite Partitioning: partition_by: [col1, col2] with hierarchical path-like keys.
  • File-Based Retention Deletes: Time-based retention uses file-level deletes for both position-based and primary-key tables.

Correctness: Synchronized partition commits, correct NULL-sentinel handling for nullable partition expressions, tombstoned inline-checkpointed rows on upsert (preventing duplicate primary keys), and live reads through expired protected snapshots.

Multi-Active HA Distributed Query (GA)โ€‹

Spice.ai Enterprise feature. See High Availability.

Distributed Query is generally available. Built on Apache Ballista, it distributes query execution across multiple active executor nodes with no single point of failure, reading directly from object storage rather than relying on a central cluster.

Distributed query supports two execution modes:

  • Synchronous: Queries for accelerated datasets are distributed across executors and results stream back in real-time โ€” best for interactive, latency-sensitive queries.
  • Asynchronous: Queries submitted via the HTTP /v1/queries API materialize results to object storage for later retrieval โ€” best for long-running analytical and batch workloads.

Key capabilities:

  • Dynamic Cluster Sizing: The planner adjusts parallelism to the number of active executors as nodes join or leave.
  • Distributed Ingestion: Ingestion for partitioned accelerated tables is distributed across executors, with partition-aware write-through splitting scheduler-side Flight DoPut writes to the responsible executors.
  • Data-Local Query Routing: Cayenne catalog queries route to the executors holding the relevant partitions.
  • Per-Executor Table Statistics: Executors report table statistics โ€” including NDV-aware estimates โ€” so distributed JoinSelection can size joins correctly, fixing out-of-memory conditions on large semi-joins.
  • Readiness & Failure Detection: /v1/ready gates on a configurable executor quorum for safe rolling deployments; scheduler readiness additionally waits for executor partition loads; executor heartbeat timeout reduced from 180s to 30s.
  • Distributed DML & DDL: UPDATE/DELETE forwarding to all executors, executor DDL sync for late joiners, and distributed MERGE INTO.
  • Cluster Observability: New cluster metrics (including scheduler_active_executors_count), distributed runtime.task_history replication, and a Grafana dashboard.
  • Ballista S3 Shuffle: Async queries with runtime.params.shuffle_location: s3://... complete reliably with executor-environment-derived S3 clients.

Security: Mutual TLS, Secret Stores, and Hardeningโ€‹

Several capabilities in this section are Spice.ai Enterprise features. See Enterprise Security.

Mutual TLS across the platform:

  • Public mTLS for HTTP and Flight: client_auth_mode: request (optional, for migration windows) or required (strict) client-certificate verification.
  • TLS Cert Hot-Reload: The runtime reloads TLS certificates on SIGHUP for zero-downtime rotation.
  • Outbound mTLS Client Certificates: FlightSQL and Spice.ai data connectors present client certificates to upstream services; the spice sql REPL supports mTLS client auth.
runtime:
tls:
enabled: true
certificate_file: /etc/spice/tls/server.crt
key_file: /etc/spice/tls/server.key
client_auth_mode: required
client_auth_ca_file: /etc/spice/tls/client-ca.crt

Authentication & Authorization (Spice.ai Enterprise):

  • OIDC Authentication: Validate OIDC bearer tokens (JWTs) issued by enterprise identity providers โ€” Microsoft Entra ID, Okta, Auth0, AWS Cognito, and Google โ€” for secure access to runtime endpoints, standalone or combined with API keys.
  • Principal-Based Policy Enforcement: Fine-grained, Cedar-based authorization policy configured under runtime.authorization governs allow/deny access across datasets, models, tools, and endpoints. Combined with identity SQL functions (current_principal(), current_principal_email(), current_principal_groups()), policies enforce per-principal row-level filtering and column masking.

New Secret Stores: HashiCorp Vault (KV v1/v2; token, approle, kubernetes, and jwt auth with automatic lease renewal) and Azure Key Vault (service principal, managed identity, workload identity, Azure CLI, or auto-detect; sovereign cloud support).

Hardening:

  • Read-only API Key Enforcement on the Flight DoGet path and async query endpoints.
  • Per-Principal Cache Namespacing: SQL, search, and caching-accelerator caches are namespaced per authenticated principal so cached results never cross identity boundaries.
  • API Key Timing Leak & Remote-UDF SSRF: Closed a timing-based position-disclosure leak in API key comparison and blocked SSRF via remote UDF endpoints.
  • Snowflake Function Deny-List: A function deny-list is enforced in Snowflake federation pushdown, and Snowflake account identifiers and auth configuration are validated at startup.
  • MCP allowed_hosts: MCP servers can be restricted to an explicit allowlist of upstream hosts.

Change Data Capture (CDC) Sourcesโ€‹

See Change Data Capture (CDC) for an overview of CDC in Spice.

  • MongoDB Change Streams: MongoDB datasets with refresh_mode: changes stream changes natively into any local accelerator โ€” no Debezium or Kafka required.
  • PostgreSQL Native Replication (WAL): PostgreSQL datasets stream INSERT/UPDATE/DELETE directly from logical replication using pgoutput decoding, with automatic per-replica slot management, an initial REPEATABLE READ bootstrap snapshot, and durable LSN acknowledgement.
  • Kafka CDC Offset Persistence: Kafka CDC offsets persist in sidecar tables for durable, resumable streams across restarts and failovers.
  • Pipelined CDC Ingestion: Source reads overlap with batch apply, with envelope coalescing and improved nullability propagation.
  • Debezium Schema Evolution: Schema changes in Debezium-sourced datasets no longer break dataset initialization on reload.
datasets:
- from: postgres:my_table
name: my_table
params:
pg_host: localhost
pg_db: mydb
acceleration:
enabled: true
engine: duckdb
refresh_mode: changes

DML, DDL, and Write-Backโ€‹

Spice v2.0 turns more connectors and catalogs into full read/write tables:

  • PostgreSQL DML: INSERT, UPDATE, and DELETE write-back on PostgreSQL datasets, with foreign-key metadata exposed via the PostgreSQL catalog connector.
  • Snowflake DML: INSERT, UPDATE, and DELETE write-back on Snowflake datasets.
  • DynamoDB DML: INSERT, UPDATE, and DELETE for DynamoDB, complementing read and CDC streaming.
  • Arrow Primary Key Upserts: Native update-or-insert semantics for in-memory Arrow-accelerated tables.
  • DDL for Iceberg: CREATE TABLE and DROP TABLE via FlightSQL and /v1/sql for Iceberg, with catalog.access: read_write_create.
  • DuckLake INSERT: DuckLake catalog tables with read_write access support INSERT.

SQL & User-Defined Functionsโ€‹

See the SQL Reference for the full SQL surface area.

  • User-Defined Functions: Define reusable SQL UDFs as first-class spicepod components, or invoke remote functions over HTTP (Spice.ai Enterprise), plus table user functions.
  • Spatial SQL UDFs: Optional geospatial ST_* UDFs for geometry workloads.
  • JSON UDTFs: flatten_json, json_tree, and flatten_json_properties table-valued functions for JSON transformation and schema decomposition (with options such as expand_maps). See JSON Functions and Operators.
  • PostgreSQL Metadata UDFs: Dataset and column descriptions are exposed via PostgreSQL-compatible UDFs (obj_description, col_description), so BI tools and psql surface Spice metadata.
  • FlightSQL Substrait Plans: CommandStatementSubstraitPlan support for clients submitting Substrait-encoded plans.
  • SQL REPL Expanded View: Toggle \x for a vertical key-value layout on wide result sets.
  • Prepared statement, federation, and unparsing fixes across the engine, including keeping correlated subqueries out of JOIN ON conditions for Spice Cloud federation and correct EXISTS/NOT EXISTS subquery handling in the federation analyzer.

Runtime Featuresโ€‹

  • On-Demand Dataset Loading: Datasets can be deferred โ€” registered with a declared schema at startup (columns[].type, columns[].nullable) and fully resolved on first reference, reducing startup time and memory for large spicepods.
  • Unified Query Cancellation: HTTP, Flight, FlightSQL, MCP, and internal execution paths honour a unified cancellation signal โ€” disconnects, REPL Ctrl-C, and cancelled HTTP requests cancel the query end-to-end.
  • Storage-Profile Accelerator Tuning: acceleration.storage_profile (auto, local_ssd, ebs, tmpfs) applies storage-aware defaults across DuckDB, SQLite, Turso, and Cayenne file-mode accelerators; auto detects the backing storage.
  • refresh_mode: snapshot (Spice.ai Enterprise): Point-in-time snapshot acceleration with SQLite/Turso WAL flushing and Cayenne metastore slice integration, now reporting accurate readiness when no snapshot exists yet.
  • Structured Component Errors: /v1/datasets?status=true and /v1/models?status=true return structured error objects (category, type, code) and human-readable error_message fields; the CLI shows an ERROR column.
  • Actionable Config Errors: Parameter typos, missing secret references, and unknown engine names produce specific, actionable errors with suggestions.

Spicepod v2โ€‹

Spicepods now support version: v2, the default for spice init, while v1 spicepods continue to work with automatic migration of deprecated fields.

VersionStatus
v2Default. Used by spice init.
v1Supported. Deprecated fields auto-migrate.
v1beta1Removed. No longer accepted.
v1 (deprecated)v2 (preferred)Notes
runtime.results_cacheruntime.caching.sql_resultsAll fields migrate automatically. cache_max_size โ†’ max_size.
runtime.memory_limitruntime.query.memory_limitAuto-migrated. query.memory_limit takes priority if both set.
runtime.temp_directoryruntime.query.temp_directoryAuto-migrated. query.temp_directory takes priority if both set.
dataset.invalid_type_actiondataset.unsupported_type_actionAuto-migrated. v2 adds a new string variant.

New v2 fields include runtime.ready_state, runtime.query.spill_compression, runtime.caching.sql_results.stale_while_revalidate_ttl, runtime.caching.sql_results.encoding, scheduler partition-assignment configuration, and catalog.access: read_write_create.

Data Connectors & Catalogsโ€‹

New connectors:

  • Elasticsearch (Alpha, Spice.ai Enterprise): Query Elasticsearch indexes as SQL tables with native hybrid search โ€” vector_search() kNN, text_search() BM25, and rrf() fusion โ€” plus Elasticsearch as a backing vector engine, direct FTS engine configuration, and index lifecycle controls.
  • GCS (Alpha): Federated queries against Google Cloud Storage, with Iceberg table support.
  • Azure Cosmos DB (Alpha): Read-only NoSQL / Core SQL API connector with cross-partition scans and schema inference.
  • Git (RC): HTTPS/SSH auth, Git LFS support, and per-repo connection resilience.
  • ADBC: Data connector and catalog with full query federation, BigQuery support, and schema/table discovery.
  • DuckLake (Beta): Lakehouse-style data management with DuckDB as the metadata catalog and object storage for data โ€” ACID transactions, time travel, and schema evolution on Parquet.
  • Self-Hosted Spice Connector: Connect Spice to another self-hosted Spice runtime as a federated source.

New catalog connectors for PostgreSQL, MySQL, MSSQL, and Snowflake, using native metadata catalogs for schema and table discovery. Unity Catalog compatibility extends to OSS Unity Catalog deployments, and DDL-defined catalogs can expose and query views.

HTTP connector: OAuth2 refresh-token authentication, query-parameter and no-limit pagination, dynamic request headers parameterised from query predicates, subquery-driven request parameters for fan-out queries, response metadata as queryable columns, map-to-array conversion, shared and persistent rate-control state across restarts and replicas, no caching of transient 429/5xx errors, and a correctly populated fetched_at column.

JSON ingestion: Single-object documents, JSONL, BOM-prefixed input, Socrata SODA responses, format auto-detection, and RFC 6901 json_pointer extraction of nested payloads.

Databricks: Resilience controls, Unity Catalog-aware permission prechecks with structured advisory errors, Classic SQL Warehouse foreign-table compatibility, connect_timeout/client_timeout parameters, a Databricks SQL dialect for federation, and Delta Lake column mapping (Name and Id modes).

Other connector improvements: MongoDB SRV support; MySQL mysql_zero_date_behavior; Snowflake OBJECT, MAP, GEOGRAPHY, GEOMETRY, VECTOR, and TIMESTAMP_LTZ types plus key-pair auth; ClickHouse Date32; S3 s3_url_style for path-style addressing and faster Parquet reads; GraphQL custom auth headers; Oracle and MSSQL sort/limit pushdown; GitHub GraphQL resilience; and improved Kafka reliability.

AI & LLMโ€‹

  • Provider-Aware Prompt Caching: LLM calls automatically use provider-side prompt caching (e.g., Anthropic, OpenAI) for system prompts and tool descriptions, reducing latency and cost.
  • Responses API Across All Providers: The Responses API works with every configured model provider, including streaming response.output_text.delta events and Authorization: Bearer header support.
  • Multi-Vector Embeddings with MaxSim: List-of-string columns produce one embedding per element with MaxSim/mean/sum scoring for ColBERT-style late-interaction retrieval, plus a _match column identifying the best-matching element.
  • rerank() UDTF: Reorder results from vector_search, text_search, or rrf using any registered chat model as a reranker, with automatic query propagation and pushdown support.
  • Searchable LLM Tool Registry: Agents discover tools via semantic search instead of enumerating every tool in the system prompt.
  • MCP Improvements: Streamable HTTP transport (/v1/mcp) on rmcp v1.5.0, native auth for streamable HTTP tools (mcp_auth_token, mcp_headers), external MCP server tool calls traced in task history, and configurable allowed_hosts.
  • Per-Model Rate-Limited AI UDF Execution for controlling concurrent AI function invocations.

Search & Vectorsโ€‹

  • DuckDB Vector Engine: vector_engine: duckdb uses DuckDB's HNSW index for fast approximate nearest-neighbor search without an external vector store. In v2.0.0, the DuckDB VSS extension is statically linked into the bundled DuckDB, so HNSW vector search works out-of-the-box on clean machines with no extension download. HNSW indexes are preserved across data refresh, and cosine_distance pushes down via array_cosine_distance.
  • Hybrid Search: Combine kNN vector search and BM25 full-text search with reciprocal rank fusion (rrf()), backed by Tantivy, Elasticsearch, or DuckDB.
  • Full-Text Search Performance: Significantly faster Tantivy ingestion with rollback-on-error, and search metadata is correctly preserved on indexing and in Vortex physical schema calculation.
  • Embedding Validation: row_id columns are validated during dataset initialization.

Cachingโ€‹

Improvements across Caching:

  • Stale-While-Revalidate: runtime.caching.sql_results.stale_while_revalidate_ttl serves stale results while revalidating in the background.
  • Cache Encoding: Optional compression (e.g., zstd) for SQL results cache entries.
  • Retention Policies for cached query results, and improved CDC-driven cache invalidation (including view plan invalidation on updates).
  • Idle Cache Maintenance: Periodic maintenance drains invalidation predicates on idle caches, fixing unbounded memory growth in rarely-read caches.

Performance & Query Engineโ€‹

Apache DataFusion is upgraded to v52.5 over the course of the release cycle, bringing:

  • Sort Pushdown to Scans: ~30x faster top-K queries on pre-sorted data; Parquet scans reverse row-group order for DESC on ASC-sorted files.
  • Rewritten Sort-Merge Join: Up to three orders of magnitude faster in pathological cases (e.g., TPC-H Q21: minutes โ†’ milliseconds).
  • Dynamic Filters: MIN/MAX aggregates and hash-join build sides prune files, row groups, and rows during execution.
  • Faster CASE Expressions, statistics caching, and prefix-aware list-files caching for faster planning.
  • TableProvider DELETE/UPDATE hooks and the RelationPlanner API for extensible SQL planning.
  • Strict Overflow Handling: try_cast_to errors on overflow instead of silently producing NULLs.

Additional engine work: default query memory limit raised from 70% to 90% with GreedyMemoryPool, partial aggregation optimization for FlightSQLExec, improved partitioned query planning, and metastore transaction support to prevent concurrent conflicts.

Rust CLIโ€‹

The Spice CLI is completely rewritten from Go to Rust โ€” a single spice binary built from the same codebase as spiced, with full feature parity across 27+ commands.

  • spice query: Interactive REPL for async queries with multi-line SQL, progress indication, and cancellation.
  • spice dataset configure: Non-interactive flag-based configuration (--from, --description, --param KEY=VALUE, --set) alongside interactive prompts.
  • spice completions: Shell completion script generation.
  • --output=json: Machine-readable output for scripting; spice login --output adds env, json, and keychain modes.
  • spice init writes a yaml-language-server schema directive for IDE completions.

Observabilityโ€‹

  • OpenTelemetry: Exporter fixes, authenticated metrics export, configurable metric name prefix (runtime.telemetry.metric_prefix), delta temporality by default, and OTLP resource attributes via runtime.telemetry.properties.
  • Query Metrics: The query_executions metric gains a datasets dimension for per-dataset query attribution.
  • Ingestion Metrics: rows_written, bytes_written, and dataset_acceleration_size_bytes for acceleration refresh and Flight DoPut/ADBC ingestion, and EXPLAIN ANALYZE metrics in FlightSQLExec.
  • Task History: Distributed task history in cluster mode and tracing for external MCP server tool calls.

Notable Bug Fixesโ€‹

  • localpod synchronization: localpod child datasets correctly track parent refreshes when the parent uses the in-memory Arrow accelerator.
  • Spice Cloud federation: Correlated subqueries are kept out of JOIN ON conditions, fixing rejected federated queries.
  • refresh_mode: snapshot: No longer reports Ready with empty data when no snapshot exists.
  • Search metadata: Field and schema metadata preserved on search indexing and in Vortex physical schema calculation.
  • HTTP connector: fetched_at column is correctly populated.
  • Connector correctness: DynamoDB Streams transient-error retries and typed-NULL DML handling; ScyllaDB physical filter pushdown disabled to fix incorrect results; MSSQL TOP N pushdown; DuckDB DELETE/UPDATE on full and caching refresh modes; Turso checked arithmetic for timestamp conversions; ODBC queries no longer silently return 0 rows on failure; Flight GetFlightInfo/DoGet schema parity.

Dependency Updatesโ€‹

Dependency / ComponentVersion
DataFusionv52.5
Ballistav52
Arrow (arrow-rs)v57.2
DuckDBv1.5.3 (with statically linked VSS)
iceberg-rustv0.9.1
Turso (libsql)v0.6.1
Vortexv0.69.0
delta_kernelv0.18.2
rmcp (MCP)v1.5.0
mistral.rsv0.8.x (candle v0.10.1)
ADBC Corev0.23
Rust toolchainv1.94.1

Contributorsโ€‹

Breaking Changesโ€‹

  • Models included by default: The separate models build variant has been removed. Local LLM inference is always included in the default build and image.

  • Windows native builds removed: Use WSL for local development.

  • Spicepod version defaults to v2: spice init creates version: v2 spicepods. v1 remains supported with auto-migration; v1beta1 is no longer accepted.

  • Flattened runtime.scheduler configuration: The nested runtime.scheduler.partition_management block is flattened and renamed:

    # Before
    runtime:
    scheduler:
    partition_management:
    interval: 30s
    max_assignments_per_cycle: 16
    discovery_timeout: 10s

    # After
    runtime:
    scheduler:
    partition_assignment_interval: 30s
    max_assignments_per_interval: 16
    partition_discovery_timeout: 10s
  • S3 metadata columns renamed: location, last_modified, size โ†’ _location, _last_modified, _size.

  • Default query memory limit changed: Increased from 70% to 90%.

  • Metric renames: accelerated_refresh metrics renamed to acceleration_refresh; last_refresh_time gauge renamed to include the milliseconds unit.

  • DuckDB parameter rename: partitioned_write_flush_threshold โ†’ partitioned_write_flush_threshold_rows.

  • /v1/search API: Always returns an array in matches, even for single results.

  • /v1/evals API removed.

  • Perplexity model provider removed.

  • x.ai model endpoint: x.ai models exclusively use the /v1/responses endpoint.

Upgrade Guide from v1.xโ€‹

Most v1 spicepods continue to work on v2.0 โ€” v1 remains supported and deprecated fields auto-migrate at load time โ€” so many deployments can upgrade by updating the binary or image alone. The steps below cover the breaking changes that may require manual action. Review each before upgrading a production deployment.

1. Build, image, and platform changesโ€‹

  • Models are now included by default. The separate models build variant (and the corresponding -models image tags) has been removed; local LLM inference is always included in the default build and image. If your deployment pinned a models build or -models-tagged image, switch to the default build/image.
  • Native Windows builds are removed. Use WSL for local Windows development.

spice init now creates version: v2 spicepods. v1 spicepods remain supported with automatic migration, but v1beta1 is no longer accepted. To move to v2, set version: v2 and update the following fields โ€” each auto-migrates from v1, but updating now clears the deprecation:

v1 (deprecated)v2 (preferred)
runtime.results_cacheruntime.caching.sql_results (cache_max_size โ†’ max_size)
runtime.memory_limitruntime.query.memory_limit
runtime.temp_directoryruntime.query.temp_directory
dataset.invalid_type_actiondataset.unsupported_type_action

3. Update changed configurationโ€‹

  • DuckDB parameter rename: partitioned_write_flush_threshold โ†’ partitioned_write_flush_threshold_rows.
  • Default query memory limit raised from 70% to 90%. If you relied on the previous default to leave headroom for other processes on the host, set it explicitly via runtime.query.memory_limit.

4. Update queries and API clientsโ€‹

  • S3 metadata columns renamed: location, last_modified, size โ†’ _location, _last_modified, _size. Update any queries that reference these columns.
  • /v1/search always returns an array in matches, even for a single result. Update clients that assumed a scalar value.
  • /v1/evals API removed. Remove integrations that depend on it.

5. Update model providersโ€‹

  • Perplexity model provider removed. Re-point affected models to another provider.
  • x.ai models use the /v1/responses endpoint exclusively. Ensure x.ai integrations target the Responses API.

6. Update observabilityโ€‹

  • Metric renames: accelerated_refresh โ†’ acceleration_refresh, and the last_refresh_time gauge is renamed to include the milliseconds unit. Update dashboards and alerts that reference these metric names.

After updating, restart the runtime and verify datasets and models report ready via /v1/datasets?status=true and /v1/models?status=true (the CLI shows a Ready/ERROR column).

Cookbook Updatesโ€‹

New Spice Cookbook recipes added during the v2.0 release cycle:

The Spice Cookbook includes more than 100 recipes to help you get started with Spice quickly and easily.

Upgradingโ€‹

To upgrade to v2.0.0, use one of the following methods:

CLI:

spice upgrade

Homebrew:

brew upgrade spiceai/spiceai/spice

Docker:

Pull the spiceai/spiceai:2.0.0 image:

docker pull spiceai/spiceai:2.0.0

For available tags, see DockerHub.

Helm:

helm repo update
helm upgrade spiceai spiceai/spiceai --version 2.0.0

AWS Marketplace:

Spice is available in the AWS Marketplace.

What's Changedโ€‹

Changelogโ€‹

  • Add TPC-DS integration tests with S3 source and PostgreSQL acceleration by @phillipleblanc in #9006
  • fix(tests): fix flaky/slow/failing unit tests by @phillipleblanc in #9009
  • fix: Update benchmark snapshots for DF51 upgrade by @app/github-actions in #9008
  • fix: add feature gate to rrf TEST_EMBEDDING_MODEL by @phillipleblanc in #9017
  • fix: features check by @phillipleblanc in #9014
  • fix: Enable Cayenne acceleration snapshots by @lukekim in #9020
  • URL table support by @lukekim in #9018
  • ScyllaDB key filter by @lukekim in #8997
  • fix: Schema mismatch when using column projection with HTTP caching by @phillipleblanc in #9021
  • Add more tests for HTTP caching with columns selection by @sgrebnov in #9025
  • HTTP cache snapshots: default to time_interval and fix snapshots_creation_policy: on_change by @sgrebnov in #9026
  • Fix duplicate snapshot creation on startup by @sgrebnov in #9029
  • Add ScyllaDB and SMB to the README table by @krinart in #9034
  • Remove waiting for runtime to be ready before creating snapshot by @krinart in #9033
  • Fix snapshot on_change policy to skip when no writes occurred by @sgrebnov in #9028
  • Release notes for release release/1.11.0-rc.2 by @krinart in #9016
  • ci: use arduino/setup-protoc for official protobuf compiler by @phillipleblanc in #9036
  • ci: install unzip on aarch64 runner for arduino/setup-protoc by @phillipleblanc in #9038
  • fix: don't fail release if upload to minio fails by @phillipleblanc in #9039
  • Add missing protoc step to setup-cc action by @krinart in #9041
  • fix: Update Search integration test snapshots by @app/github-actions in #9013
  • Fix formula_1 and codebase_community in bird-bench by @Jeadie in #9000
  • Cayenne S3 Express One Zone improvements by @lukekim in #9015
  • Add zlib1g-dev to CI by @lukekim in #9052
  • Improve validation and logging for hash indexes by @lukekim in #9047
  • Upgrade Vortex with CASE-WHEN by @lukekim in #9051
  • x.ai models now exclusively use /v1/responses endpoint by @lukekim in #9400
  • Improvements for snapshot schema comparison by @krinart in #9401
  • v2.0 breaking changes by @lukekim in #9233
  • Create PartitionManagementTask for scheduler to update accelerated table partition assignments by @Jeadie in #9378
  • refactor(Cayenne): route all write orchestration through CayenneDataSink by @sgrebnov in #9402
  • Refactor benchmark to use QueryExecutor trait by @Jeadie in #9418
  • feat: Add spidapter build and release workflow by @peasee in #9427
  • Testoperator: add support for api-key when connecting to external spice instance by @sgrebnov in #9421
  • Initial implementation of Ducklake catalog & data connectors by @lukekim in #9083
  • Require aws_lc_rs since jsonwebtoken upgrade by @Jeadie in #9426
  • feat: Add spidapter tool by @peasee in #9425
  • Add release notes for 1.11.2 patch release by @sgrebnov in #9430
  • feat(spidapter): integrate system-adapter-protocol with SCP provisioning by @phillipleblanc in #9434
  • Add DuckLake TPCH E2E workflow and federated Spicepod configuration by @lukekim in #9431
  • fix(spidapter): use Flight handshake auth instead of x-api-key header by @phillipleblanc in #9435
  • [spidapter] Keep only what sparks joy by @Jeadie in #9439
  • Refactor binary operator balancing by @Jeadie in #9424
  • feat: Add Iceberg DDL support (CREATE TABLE / DROP TABLE) for default catalog override by @phillipleblanc in #9440
  • Fix Flight SQL schema consistency: expand view types and verify field names by @sgrebnov in #9438
  • Update spidapter for new system-adapter-protocol by @sgrebnov in #9442
  • docs: fix typos and syntax errors in style guide and error handling docs by @cluster2600 in #9445
  • Add acceleration refresh ingestion metrics (rows_written, bytes_written) by @phillipleblanc in #9461
  • Refactor(Cayenne): Replace CatalogError and string based errors with Snafu errors by @sgrebnov in #9403
  • Replace deprecated claude-3-5-haiku-latest with claude-haiku-4-5 by @Jeadie in #9492
  • Fix #9481: Preserve schema in results cache for empty query results by @phillipleblanc in #9485
  • Fix partition by serializing by @Jeadie in #9474
  • query: reconcile execution stream nullability with logical plan schema by @phillipleblanc in #9486
  • initial spice-cloud-client crate and spice cloud metrics --app <app-name>. by @Jeadie in #9480
  • feat: Return dataset error message in datasets API by @peasee in #9487
  • Spicebench by @lukekim in #9447
  • build(deps): consolidate dependabot dependency updates by @phillipleblanc in #9504
  • fix(cluster): route non-partitioned accelerated tables in distributed mode by @phillipleblanc in #9508
  • Enable core scalar UDFs in refresh SQL by @sgrebnov in #9502
  • Fix metrics in Spidapter again by @Jeadie in #9497
  • fix(cluster): tolerate Completed->status propagation race in distributed query handle by @phillipleblanc in #9510
  • feat: Support distributed ingestion in cayenne catalog by @peasee in #9506
  • Fix Cayenne duplicate primary keys after DELETE + UPSERT CDC sequences by @krinart in #9494
  • fix(cluster): rewrite table scans inside subqueries for distributed execution by @phillipleblanc in #9518
  • fix: Set catalog mode to readwritecreate in spidapter by @peasee in #9519
  • Upgrade AWS SDK crates & set APN user-agent in AWS SDK credential bridge by @lukekim in #8328
  • feat(runtime): add runtime ready_state on_registration semantics by @lukekim in #9522
  • fix: Add spidapter post-setup retries by @peasee in #9526
  • Make partition discovery more robust and make initialization non-blocking by @sgrebnov in #9499
  • Make lint-rust-fix support targeted packages and features by @Jeadie in #9511
  • Handle new Cloud SCP API by @Jeadie in #9532
  • Refactor and simplify streaming benchmarks by @krinart in #9405
  • fix: ensure spidapter only increments attempts on failures by @peasee in #9534
  • feat: Support specifying app resources in spidapter by @peasee in #9536
  • test(runtime): Spice Cayenne DDL integration test by @lukekim in #9535
  • fix: Handle schema evolution mismatch errors during data refresh by @lukekim in #9527
  • fix: resolve clippy lint warnings by @phillipleblanc in #9547
  • pr-builds --tag <TAG> for build_and_release.yml by @Jeadie in #9507
  • Add --output flag to spice login with env/json/keychain modes by @Jeadie in #9541
  • Don't use 'PartitionedTableScanRewrite' in async distributed query by @Jeadie in #9548
  • feat(spidapter): add local backend mode with single executor by @phillipleblanc in #9531
  • support chat template in HF by @Jeadie in #9543
  • fix(cayenne): stream PK retention deletes and run OOM regression in CI by @phillipleblanc in #9533
  • cayenne: Staged append writes to prevent partial writes and data loss on stream error by @sgrebnov in #9491
  • AcceleratedTable::scan use FederatedTable::scan when ClusterRole::Scheduler by @Jeadie in #9550
  • Upgrade to delta-kernel-rs v0.18.2 by @lukekim in #9528
  • Run cayenne tests as part of PR CI by @sgrebnov in #9554
  • Upgrade to DataFusion v52.2.0 by @lukekim in #9419
  • Remove Snapshot Compaction + Add snapshot existence check by @krinart in #9523
  • Update dependencies by @lukekim in #9566
  • fix: Update benchmark snapshots by @app/github-actions in #9565
  • fix: Compare Cayenne table configuration on startup by @peasee in #9529
  • Make Refresh::refresh_sql more robust to alterations over time. by @Jeadie in #9549
  • fix: Update datafusion-table-providers dependency to latest revision by @lukekim in #9574
  • Unset AWS_ENDPOINT_URL when empty by @krinart in #9575
  • fix: allow BytesProcessedExec repartitioning for unordered input by @lukekim in #9540
  • Sanitize DataFusion errors by @lukekim in #9530
  • Add conditional logging for partition assignments by @Jeadie in #9577
  • use 'properly early exit on SIGTERM' by @Jeadie in #9573
  • Update datafusion to 52.2.0 by @phillipleblanc in #9582
  • Ensure we query one and only one partition per request by @Jeadie in #9416
  • feat: Add support for Spicepod version v2 by @lukekim in #9583
  • [SpiceDQ] Improve error messages; Avoid race condition on allocate_initial_partitions. by @Jeadie in #9579
  • Update ballista dependencies to latest 52.0.0 revision by @lukekim in #9581
  • Fix Databricks spark_connect mode always disabled by @phillipleblanc in #9586
  • Support partitioning in Arrow accelerator by @Jeadie in #9571
  • Fix spice query CLI response deserialization by @phillipleblanc in #9588
  • fix: Update benchmark snapshots by @app/github-actions in #9584
  • fix: Share RuntimeEnv across Cayenne read/write/delete paths for targeted list_files_cache invalidation by @sgrebnov in #9589
  • feat: Add file:// state_location support for async queries scheduler by @phillipleblanc in #9590
  • Update endgame links by @krinart in #9598
  • ci: fix E2E CLI upgrade test to use latest release for spiced download by @phillipleblanc in #9613
  • fix(DF): Lazily initialize BatchCoalescer in RepartitionExec to avoid schema type mismatch by @sgrebnov in #9623
  • feat: Implement catalog connectors for various databases by @lukekim in #9509
  • Refactor and clean up code across multiple crates by @lukekim in #9620
  • fix: Improve error handling for distributed mode and state_location configuration by @lukekim in #9611
  • Properly install postgres in install-postgres action by @krinart in #9629
  • fix: Use Python venv for schema validation in CI by @phillipleblanc in #9637
  • Update spicepod.schema.json by @app/github-actions in #9640
  • Update testoperator dispatch to use release/2.0 branch by @phillipleblanc in #9641
  • fix: Align CUDA asset names in Dockerfile and install tests with build output by @phillipleblanc in #9639
  • Fix expect test scripts in E2E Installation AI test by @sgrebnov in #9643
  • testoperator for partitioned arrow accelerator by @Jeadie in #9635
  • Remove default 1s refresh_check_interval from spidapter for hive datasets by @phillipleblanc in #9645
  • Fix scheduler panic and cancel race condition by @phillipleblanc in #9644
  • Align Spice.ai connector parameter names across catalog/data connectors by @lukekim in #9632
  • docs: update distribution details and add NAS support in release notes by @lukekim in #9650
  • Enable postgres-accel in CI builds for benchmarks by @sgrebnov in #9649
  • perf: Cache Turso metastore connection across operations by @penberg in #9646
  • Add 'scheduler_state_location' to spidapter by @Jeadie in #9655
  • Implement Cayenne S3 Express multi-zone live test with data validation by @lukekim in #9631
  • chore(spidapter): bump default memory limit from 8Gi to 32Gi by @phillipleblanc in #9661
  • perf: Use prepare_cached() in Turso and SQLite metastore backends by @penberg in #9662
  • Improve CDC cache invalidation by @krinart in #9651
  • Refactor Cayenne IDs to use UUIDv7 strings by @lukekim in #9667
  • fix: add liveness check for dead executors in partition routing by @Jeadie in #9657
  • fix(s3): Fix metadata column schema mismatches in projected queries by @sgrebnov in #9664
  • s3_metadata_columns tests: include test for location outside table prefix by @sgrebnov in #9676
  • docs: Update DuckDB, GCS, Git connector and Cayenne documentation by @lukekim in #9671
  • Add s3_url_style support for S3 connector URL addressing by @phillipleblanc in #9642
  • Consolidate E2E workflows and require WSL for Windows runtime by @lukekim in #9660
  • Upgrade to Rust v1.93.1 by @lukekim in #9669
  • Security fixes and improvements by @lukekim in #9666
  • feat(flight): add DoPut rows/bytes written metrics for DoPut ETL ingestion tracking by @phillipleblanc in #9663
  • Skip caching http error response + add response_headers by @krinart in #9670
  • refactor: Remove v1/evals functionality by @Jeadie in #9420
  • Make a test harness for Distributed Spice integration tests by @Jeadie in #9615
  • Enable on_zero_results: use_source for views by @krinart in #9699
  • fix(spidapter): Lower memory limit, passthrough AWS secrets, override flight URL by @peasee in #9704
  • Show an error on a shared acceleration file with snapshots enabled by @krinart in #9698
  • Fixes for anthropic by @Jeadie in #9707
  • Use max_partitions_per_executor in allocate_initial_partitions by @Jeadie in #9659
  • [SpiceDQ] Accelerations must have partition key by @Jeadie in #9711
  • Upgrade to Turso v0.5 by @lukekim in #9628
  • feat: Rename metadata columns to _location, _last_modified, _size by @phillipleblanc in #9712
  • fix: bump datafusion-ballista to fix BatchCoalescer schema mismatch panic by @phillipleblanc in #9716
  • fix: Ensure Cayenne respects target file size by @peasee in #9730
  • refactor: Make DDL preprocessing generic from Iceberg DDL processing by @peasee in #9731
  • [SpiceDQ] Distribute query of Cayenne Catalog to executors with data by @Jeadie in #9727
  • Properly set primary_keys/on_conflict for Cayenne tables by @krinart in #9739
  • Add executor resource and replica support to cloud app config by @ewgenius in #9734
  • feat: Support PARTITION BY in Cayenne Catalog table creation by @peasee in #9741
  • Update datafusion and related packages to version 52.3.0 by @lukekim in #9708
  • Route FlightSQL statement updates through QueryBuilder by @phillipleblanc in #9754
  • JSON file format improvements by @lukekim in #9743
  • [SpiceDQ] Partition Cayenne catalogs writes through to executors by @Jeadie in #9737
  • Update to DF v52.3.0 versions of datafusion & datafusion-tableproviders by @lukekim in #9756
  • Make S3 metadata column handling more robust by @sgrebnov in #9762
  • Fetch API keys from dedicated endpoint instead of apps response by @phillipleblanc in #9767
  • Update arrow-rs, datafusion-federation, and datafusion-table-providers dependencies by @phillipleblanc in #9769
  • Chunk metastore batch inserts to respect SQLite parameter limits by @phillipleblanc in #9770
  • Improve JSON SODA support by @lukekim in #9795
  • Add ADBC Data Connector by @lukekim in #9723
  • docs: Release Cayenne as RC by @peasee in #9766
  • cli[feat]: cloud mode to use region-specific endpoints by @lukekim in #9803
  • Include updated JSON formats in HTTPS connector by @lukekim in #9800
  • Flight DoPut: Partition-aware write-through forwarding by @Jeadie in #9759
  • Pass through authentication to ADBC connector by @lukekim in #9801
  • Move scheduler_state_location from adapter metadata to env var by @phillipleblanc in #9802
  • Fix Cayenne DoPut upsert returning stale data after 3+ writes by @phillipleblanc in #9806
  • Fix JSON column projection producing schema mismatch by @sgrebnov in #9811
  • Fix http connector by @krinart in #9818
  • Fix ADBC Connector build and test by @lukekim in #9813
  • Support update & delete DML for distributed cayenne catalog by @Jeadie in #9805
  • Set allow_http param when S3 endpoint uses http scheme by @phillipleblanc in #9834
  • fix: Cayenne Catalog DDL requires a connected executor in distributed mode by @Jeadie in #9838
  • fix: Add conditional put support for file:// scheduler state location by @Jeadie in #9842
  • fix: Require the DDL primary key contain the partition key by @Jeadie in #9844
  • fix: Databricks SQL Warehouse schema retrieval with INLINE disposition and async retry by @lukekim in #9846
  • Filter pushdown improvements for SqlTable by @lukekim in #9852
  • feat: add iam_role_source parameter for AWS credential configuration by @lukekim in #9854
  • Fix ODBC queries silently returning 0 rows on query failure by @lukekim in #9864
  • feat(adbc): Add ADBC catalog connector with schema/table discovery by @lukekim in #9865
  • Make Turso SQL unparsing more robust and fix date comparisons by @lukekim in #9871
  • Fix Flight/FlightSQL filter precedence and mutable query consistency by @lukekim in #9876
  • Partial Aggregation optimisation for FlightSQLExec by @lukekim in #9882
  • fix: v1/responses API preserves client instructions when system_prompt is set by @Jeadie in #9884
  • feat: emit scheduler_active_executors_count and use it in spidapter by @Jeadie in #9885
  • feat: Add custom auth header support for GraphQL connector by @krinart in #9899
  • Add --endpoint flag to spice run with scheme-based routing by @lukekim in #9903
  • When executor connects, send DDL for existing tables by @Jeadie in #9904
  • fix: Improve ADBC driver shutdown handling and error classification by @lukekim in #9905
  • fix: require all executors to succeed for distributed DML (DELETE/UPDATE) forwarding by @Jeadie in #9908
  • fix(cayenne catalog): fix catalog refresh race condition causing duplicate primary keys by @Jeadie in #9909
  • Remove Perplexity support by @Jeadie in #9910
  • Fix refresh_sql support for debezium constraints by @krinart in #9912
  • Implement DML for DynamoDBTableProvider by @lukekim in #9915
  • chore: Update iceberg-rust fork to v0.9 by @lukekim in #9917
  • Run physical optimizer on FallbackOnZeroResultsScanExec fallback plan by @sgrebnov in #9927
  • Improve Databricks error message when dataset has no columns by @sgrebnov in #9928
  • Delta Lake: fix data skipping for >= timestamp predicates by @sgrebnov in #9932
  • fix: Ensure distributed Cayenne DML inserts are forwarded to executors by @Jeadie in #9948
  • Add full query federation support for ADBC data connector by @lukekim in #9953
  • Make time_format deserialization case-insensitive by @claudespice in #9955
  • Hash ADBC join-pushdown context to prevent credential leaks in EXPLAIN plans by @lukekim in #9956
  • fix: Normalize Arrow Dictionary types for DuckDB and SQLite acceleration by @sgrebnov in #9959
  • ADBC BigQuery: Improve BigQuery dialect date/time and interval SQL generation by @lukekim in #9967
  • Make BigQueryDialect more robust and add BigQuery TPC-H benchmark support by @lukekim in #9969
  • fix: Show proper unauthorized error instead of misleading runtime unavailable by @lukekim in #9972
  • fix: Enforce target_chunk_size as hard maximum in chunking by @lukekim in #9973
  • Add caching retention by @krinart in #9984
  • fix: improve Databricks schema error detection and messages by @lukekim in #9987
  • fix: Set default S3 region for opendal operator and fix cayenne nextest by @phillipleblanc in #9995
  • fix(PostgreSQL): fix schema discovery for PostgreSQL partitioned tables by @sgrebnov in #9997
  • fix: Defer cache size check until after encoding for compressed results by @krinart in #10001
  • fix: Rewrite numeric BETWEEN to CAST(AS REAL) for Turso by @lukekim in #10003
  • fix: Handle integer time columns in append refresh for all accelerators by @sgrebnov in #10004
  • fix: preserve s3a:// scheme when building OpenDalStorageFactory with custom endpoint by @phillipleblanc in #10006
  • Fix ISO8601 time_format with Vortex/Cayenne append refresh by @sgrebnov in #10009
  • fix: Address data correctness bugs found in audit by @sgrebnov in #10015
  • fix(federation): fix SQL unparsing for Inexact filter pushdown with alias by @lukekim in #10017
  • Improve GitHub connector ref handling and resilience by @lukekim in #10023
  • feat: Add spice completions command for shell completion generation by @lukekim in #10024
  • fix: Fix data correctness bugs in DynamoDB decimal conversion and GraphQL pagination by @sgrebnov in #10054
  • Implement RefreshDataset for distributed control stream by @Jeadie in #10055
  • perf: Improve S3 parquet read performance by @sgrebnov in #10064
  • fix: Prevent write-through stalls and preserve PartitionTableProvider during catalog refresh by @Jeadie in #10066
  • feat: spice completions auto-detects shell directory and writes file by @lukekim in #10068
  • fix: Bug in DynamoDB, GraphQL, and ISO8601 refresh data handling by @sgrebnov in #10063
  • fix partial aggregation deduplication on string checking by @lukekim in #10078
  • fix: add MetastoreTransaction support to prevent concurrent transaction conflicts by @phillipleblanc in #10080
  • fix: Use GreedyMemoryPool, add spidapter query memory limit arg by @phillipleblanc in #10082
  • feat: Add metrics for EXPLAIN ANALYZE in FlightSQLExec by @lukekim in #10084
  • Use strict cast in try_cast_to to error on overflow instead of silent NULL by @sgrebnov in #10104
  • feat: Implement MERGE INTO for Cayenne catalog tables by @peasee in #10105
  • feat: Add distributed MERGE INTO support for Cayenne catalog tables by @peasee in #10106
  • Improve JSON format auto-detection for single multi-line objects by @lukekim in #10107
  • Add mode: file_update acceleration mode by @krinart in #10108
  • Coerce unsupported Arrow types to Iceberg v2 equivalents in REST catalog API by @peasee in #10109
  • fix: Update default query memory limit to 90% from 70% by @phillipleblanc in #10112
  • feat: Add mTLS client auth support to spice sql REPL by @lukekim in #10113
  • fix(datafusion-federation): report error on overflow instead of silent NULL by @sgrebnov in #10124
  • fix: Prevent data loss in MERGE when source has duplicate keys by @peasee in #10126
  • feat: Add ClickHouse Date32 type support by @sgrebnov in #10132
  • Add Delta Lake column mapping support (Name/Id modes) by @sgrebnov in #10134
  • fix: Restore Turso numeric BETWEEN rewrite lost in DML revert by @lukekim in #10139
  • fix: Enable arm64 Linux builds with fp16 and lld workarounds by @lukekim in #10142
  • fix: remove double trailing slash in Unity Catalog storage locations by @sgrebnov in #10147
  • fix: Improve GitHub GraphQL client resilience and performance by @lukekim in #10151
  • Enable reqwest compression and optimize HTTP client settings by @lukekim in #10154
  • fix: executor startup failures by @Jeadie in #10155
  • feat: Distributed runtime.task_history support by @Jeadie in #10156
  • fix: Preserve timestamp timezone in DDL forwarding to executors by @peasee in #10159
  • feat: Per-model rate-limited concurrent AI UDF execution by @Jeadie in #10160
  • fix(Turso): Reject subquery/outer-ref filter pushdown in Turso provider by @lukekim in #10174
  • Fix linux/macos spice upgrade by @phillipleblanc in #10194
  • Improve CREATE TABLE LIKE error messages, success output, EXPLAIN, and validation by @peasee in #10203
  • fix: chunk MERGE delete filters and update Vortex for stack-safe IN-lists by @peasee in #10207
  • Propagate runtime.params.parquet_page_index to Delta Lake connector by @sgrebnov in #10209
  • Properly mark dataset as Ready on Scheduler by @Jeadie in #10215
  • fix: handle Utf8View/LargeUtf8 in GitHub connector ref filters by @lukekim in #10217
  • fix(databricks): Fix schema introspection and timestamp overflow by @lukekim in #10226
  • fix(databricks): Fix schema introspection failures for non-Unity-Catalog environments by @lukekim in #10227
  • feat: Add pagination support to HTTP data connector by @lukekim in #10228
  • feat(databricks): DESCRIBE TABLE fallback and source-native type parsing for Lakehouse Federation by @lukekim in #10229
  • fix(databricks): harden HTTP retries, compression, and token refresh by @lukekim in #10232
  • feat[helm chart]: Add support for ServiceAccount annotations and AWS IRSA example by @peasee in #9833
  • fix: Log warning and fall back gracefully on Cayenne config change by @krinart in #9092
  • fix: Handle engine mismatch gracefully in snapshot fallback loop by @krinart in #9187
  • fix: Full Text Search schema mismatch with ADBC connector by @lukekim in #10235
  • docs: Update v2.0.0-rc.2 release notes with latest changes by @lukekim in #10238
  • Fix append refresh dedup failure when refresh_sql selects column subset by @sgrebnov in #10225
  • Revert "Properly mark dataset as Ready on Scheduler (#10215)" by @sgrebnov in #10242
  • Fix failing merge conflicts for benchmarks by @krinart in #10247
  • fix(github): fetch commits for dynamic and slash refs by @lukekim in #10233
  • Upgrade DataFusion to v52.5.0-rc1 by @lukekim in #10249
  • Merge develop to trunk (2026-04-09) by @claudespice in #10248
  • fix: Validate embedding row_id columns during dataset init (fixes #8226) by @claudespice in #10208
  • fix: Update tpch benchmark snapshots for federated/glue[csv].yaml by @app/github-actions in #10244
  • feat(databricks): add resilience controls, UC awareness, and task history instrumentation by @lukekim in #10246
  • fix: Make PartitionManager resilient to bare vs fully qualified table references by @sgrebnov in #10257
  • fix: Update tpch benchmark snapshots for accelerated/s3[parquet]-cayenne[file].yaml by @app/github-actions in #10256
  • Merge develop to trunk (2026-04-10) by @claudespice in #10251
  • Improve Snowflake/ADBC dataset registration performance and observability by @lukekim in #10266
  • Fixes for kafka connector by @krinart in #10263
  • fix(runtime): gate otel code tags, suppress aws sdk noise, and unblock connector init by @lukekim in #10260
  • fix(runtime): avoid regionless AWS SDK loads by @lukekim in #10271
  • Add versioned release install workflow coverage by @lukekim in #10276
  • fix(runtime): handle HTTP JSON unions and spicepod reloads by @lukekim in #10277
  • Databricks UC permission prechecks: explicit denial as permanent error, ambiguous cases advisory by @lukekim in #10274
  • Revert component status changes re-introduced by develop merge (#10248) by @sgrebnov in #10293
  • Fix broken CI workflows by @ewgenius in #10294
  • Group dependabot updates by ecosystem by @lukekim in #10296
  • fix(tests): Replace flaky S3 Vectors snapshot tests with structural validation by @lukekim in #10301
  • Update test_github_workflows snapshot by @lukekim in #10304
  • fix(ci): fix Bedrock runner mismatch and snapshot auto-merge failure by @ewgenius in #10306
  • feat(http): Add map-to-array conversion and query-parameter pagination by @lukekim in #10295
  • New crate: datafusion-ddl by @Jeadie in #10205
  • Make Databricks UC permission checks advisory with structured error reporting by @lukekim in #10283
  • build(deps): bump the github-actions-dependencies group with 4 updates by @app/dependabot in #10298
  • fix: Clear cached plans on view updates by @peasee in #10312
  • build(deps): bump the aws-sdk group with 7 updates by @app/dependabot in #10299
  • Code out of runtime. by @Jeadie in #10178
  • fix: Respect function registry denies for accelerated table filter pushdown by @peasee in #10311
  • fix: Don't block heartbeat when all slots acquired by @peasee in #10322
  • fix: strip only outer parens in get_table_partition_expr_from_ctx by @Jeadie in #10323
  • Upgrade datafusion-table-providers with MongoDB SRV support by @lukekim in #10317
  • fix: Avoid pushing down bucketing partition expressions into executors by @peasee in #10324
  • Upgrade datafusion-table-providers to d1b911a5 and bump adbc to 0.23 by @lukekim in #10329
  • fix: Update Search integration test snapshots by @app/github-actions in #10308
  • Handle foreign table + Classic sql warehouse combination gracefully by @krinart in #10318
  • New crate datafusion-flightsql by @Jeadie in #10201
  • Set tantivy=warn unless very verbose logging by @Jeadie in #10338
  • Remove image registry and image name options from spidapter by @ewgenius in #10241
  • build(deps): bump sysinfo from 0.37.2 to 0.38.4 by @app/dependabot in #10291
  • build(deps): bump futures from 0.3.31 to 0.3.32 by @app/dependabot in #10289
  • New crate 'datafusion-dml' by @Jeadie in #10334
  • Jeadie/26 04 16/spice sql by @Jeadie in #10343
  • Add Teraswitch/Pittsburgh apt mirrors + retry config for CI runners by @lukekim in #10349
  • Implement sort pushdown and fix pushdown gaps across providers by @lukekim in #10337
  • Merge develop to trunk (2026-04-16) by @claudespice in #10345
  • Update candle and mistral.rs lock-step pins by @lukekim in #10278
  • docs: fix status badges in README by @lukekim in #10350
  • Migrate secrets to vars by @krinart in #10354
  • Add limit pushdown and improve sort pushdown for Oracle and MSSQL by @sgrebnov in #10351
  • Fix ubuntu mirror configuration by @ewgenius in #10359
  • fix: Increase throughput test default ready_wait from 30s to 300s (fixes #8207) by @claudespice in #10344
  • Add auth headers support to OTEL metrics exporter by @lukekim in #10347
  • fix(github): shrink GraphQL page size on gateway errors; lower comment defaults by @lukekim in #10355
  • Relax apt mirror substitution failure to warning in CI action by @ewgenius in #10361
  • feat(http): Add OAuth2 refresh-token auth to HTTP connector by @lukekim in #10348
  • Upgrade Rust toolchain to 1.94.1 by @lukekim in #10353
  • Handle order by and sort in PartitionedTableScanRewrite by @Jeadie in #9656
  • Fix OTEL Exporter by @krinart in #10363
  • Pin spiceai candle / TEI forks to merged revs; drop local [patch] overrides by @lukekim in #10362
  • Integrate spiceio and makefile_targets into pr.yml by @lukekim in #10357
  • ci: skip artifact compression for test binaries/archives by @lukekim in #10381
  • chore(deps): bump spiceai/candle, spiceai/mistral.rs, aws-lc-rs, tantivy, rand by @lukekim in #10379
  • Bump datafusion-table-providers (#10375) by @lukekim in #10384
  • fix: Update Search integration test snapshots by @app/github-actions in #10376
  • v2.0.0-rc.3 preparation by @ewgenius in #10382
  • fix(spicepod): JSON schema accepts string or {name: expr} for partition_by by @lukekim in #10352
  • fix: Use ROUND for Turso decimal BETWEEN comparisons (fixes #9872) by @claudespice in #10360
  • Revert "v2.0.0-rc.3 preparation" from trunk by @ewgenius in #10386
  • Add on_schema_resolved dataset ready state by @lukekim in #10368
  • feat: Add Elasticsearch data connector with hybrid search support by @lukekim in #10258
  • ci: bump test archive upload compression-level to 1 by @lukekim in #10388
  • feat(git-connector): promote Git connector to RC status by @lukekim in #10385
  • feat(postgres): stream WAL directly to Spice accelerators by @lukekim in #10364
  • Add schema decomposition to the HTTP connector by @lukekim in #10393
  • fix(cayenne): Skip catalog refresh state reload for existing providers by @sgrebnov in #10396
  • Make cayenne-flightsql tool by @Jeadie in #10356
  • build(deps): bump the github-actions-dependencies group with 2 updates by @app/dependabot in #10398
  • Update openapi.json by @app/github-actions in #10272
  • Merge develop to trunk โ€” 2026-04-19 by @claudespice in #10407
  • feat(otel): default OTLP push exporter to delta temporality by @phillipleblanc in #10412
  • fix: Restore analyzer rule ordering to run federation before type coercion by @sgrebnov in #10415
  • fix: Map Utf8/LargeUtf8 to STRING in Databricks/Spark SQL dialects by @sgrebnov in #10420
  • feat(otel): add metric name prefix at runtime.telemetry.metric_prefix by @phillipleblanc in #10418
  • fix: Map LargeUtf8 to VARCHAR in Athena ODBC dialect by @sgrebnov in #10419
  • feat(cluster): connector-driven object store registration on executors by @phillipleblanc in #10414
  • build(deps): bump ubuntu from 22.04 to 24.04 in the docker-dependencies group by @app/dependabot in #10397
  • fix: Update benchmark snapshots Apr 20 by @app/github-actions in #10417
  • feat(otel): apply runtime.telemetry.properties as resource attributes on exported metrics by @phillipleblanc in #10416
  • Publish RC releases to DockerHub; upgrade runners to ubuntu-24.04 by @lukekim in #10428
  • feat: Add Azure Cosmos DB (NoSQL) data connector (RC) by @lukekim in #10392
  • feat(datafusion): flatten_json_properties + json_tree UDTFs by @lukekim in #10406
  • Harden /v1/tools and /v1/nsql against unauthenticated / LLM-driven SQL by @lukekim in #10365
  • feat(embeddings): multi-vector embeddings with MaxSim + late-interaction by @lukekim in #10408
  • Update GH runners for CUDA builds by @ewgenius in #10432
  • fix(delta_lake): register object stores on cluster executors by @phillipleblanc in #10436
  • DF-native DML by @krinart in #10327
  • ci: run Build and Test on spiceai-macos; split install jobs by profile by @lukekim in #10434
  • Improve search UDTFs: text_search, vector_search, rrf by @lukekim in #10387
  • fix(model2vec): Improve robustness of model loading for sentence-transformers layouts by @sgrebnov in #10444
  • Merge develop to trunk โ€” 2026-04-21 by @claudespice in #10448
  • Enable filter pushdown for vector_search UDTF by @sgrebnov in #10447
  • Support Snowflake OBJECT, MAP, GEOGRAPHY, GEOMETRY, VECTOR, TIMESTAMP_LTZ types by @lukekim in #10451
  • Fix Databricks tests by @krinart in #10449
  • fix(cluster): forward register_object_stores through connector wrappers by @phillipleblanc in #10460
  • Fixes for vector-search by @krinart in #10455
  • Add expand_maps option and flatten_json UDTF by @lukekim in #10452
  • fix: Update Search integration test snapshots by @app/github-actions in #10458
  • Fix physical codec decode ambiguity for empty protobuf messages by @sgrebnov in #10466
  • chore(logging): demote s3_single_file_cached skip refresh log to debug by @phillipleblanc in #10467
  • Enable filter pushdown for rrf UDTF by @sgrebnov in #10465
  • feat(cluster): consolidate distributed state into cluster.json by @phillipleblanc in #10463
  • feat(cayenne): Add column statistics and data inlining by @lukekim in #10314
  • docs(copilot): flag missing wrapper delegation when adding default trait methods by @phillipleblanc in #10461
  • Wire Elasticsearch vector engine write path through acceleration by @lukekim in #10453
  • Add helm lint CI by @ewgenius in #10468
  • Fix Azure and GCS acceleration snapshot object store credential handling by @phillipleblanc in #10486
  • Update spicepod.schema.json by @app/github-actions in #10485
  • fix(secrets): harden AWS Secrets Manager secret store by @lukekim in #10478
  • Update datafusion-ballista crate by @sgrebnov in #10488
  • feat(secrets): add ParameterSpec and more params for AWS secrets manager by @phillipleblanc in #10487
  • Add rerank UDTF for hybrid search with query auto-propagation by @lukekim in #10469
  • Fix flatten_json_properties by @krinart in #10475
  • fix: preserve field and schema metadata in expand_views_schema by @claudespice in #10494
  • Upgrade rmcp to upstream 1.5.0; switch MCP server to Streamable HTTP by @lukekim in #10491
  • fix: handle Snowflake TIMESTAMP_LTZ wire format and prevent nanosecond overflow by @claudespice in #10493
  • Lint parity in Makefile by @krinart in #10492
  • Add connect_timeout/client_timeout params to Databricks sql_warehouse mode by @lukekim in #10495
  • fix(tracing): suppress opentelemetry INFO logs at all verbosity levels by @lukekim in #10497
  • DynamoDB DML by @krinart in #10470
  • feat(cayenne): native vector search via SIMD similarity UDFs by @lukekim in #10456
  • fix(cli): suppress banner for all JSON-producing cloud subcommands (fixes #10498) by @claudespice in #10510
  • fix(deps): bump openssl to 0.10.78 by @phillipleblanc in #10509
  • fix(s3): quiet AWS SDK credential probe when no region is configured by @phillipleblanc in #10506
  • fix(cdc): emit ready signal on caught-up Kafka/Debezium streams (#5201) by @phillipleblanc in #10504
  • runtime-cluster crate + Run partition discovery before forwarding refresh to executors by @krinart in #10490
  • Update lint-rust target to use --keep-going by @Jeadie in #10508
  • Add TPC-H SF100 s3[parquet]-duckdb[file] benchmark spicepod by @lukekim in #10524
  • Remove dev-profile install steps from pr.yml by @Jeadie in #10507
  • fix: add missing NULL check on Timestamp path in append refresh by @claudespice in #10518
  • fix: return error on Decimal128/256 overflow instead of silently dropping scale by @claudespice in #10519
  • fix: delegate update and delete_from in IndexedTableProvider and EmbeddingTable by @claudespice in #10520
  • feat(devx): make config errors, CLI, and REPL lead users to success by @lukekim in #10489
  • fix(rerank): defer execution to RerankExec, enable filters and projection pushdown by @sgrebnov in #10514
  • fix(llms): support Gemma models with missing attention_bias config field by @lukekim in #10523
  • Fix vector_search silently ignoring named limit/column/include_score args by @sgrebnov in #10527
  • fix: split unsupported filters locally in scan() for UseSource mode by @ewgenius in #10528
  • feat(secrets): add Azure Key Vault secret store by @lukekim in #10496
  • Bump mistralrs by @krinart in #10532
  • Fix benchmark configurations and CI build issues by @sgrebnov in #10535
  • Fix catalog query overrides for MySQL and MSSQL benchmarks by @sgrebnov in #10543
  • For Cayenne, preserve matched columns for MERGE ... ON <cols> by @Jeadie in #10340
  • build(deps): bump the aws-sdk group across 1 directory with 5 updates by @app/dependabot in #10538
  • docs: update AI agent instructions (git workflow + Rust 1.94) by @lukekim in #10544
  • fix: Update tpch benchmark snapshots by @app/github-actions in #10529
  • fix: Update tpch benchmark snapshots for accelerated/s3[parquet]-duckdb[file].yaml by @app/github-actions in #10525
  • Extract runtime-datafusion from runtime by @krinart in #10545
  • Use generic DML extension planner for Cayenne by @Jeadie in #10437
  • fix: Update Search integration test snapshots by @app/github-actions in #10552
  • Fix security and correctness audit issues by @lukekim in #10526
  • fix(MySQL): revert MySQL result column reorder to fix federated query failures by @sgrebnov in #10557
  • Fix protoc installation by @krinart in #10566
  • fix: Disable Ballista dynamic filters on HashJoinExec by @peasee in #10548
  • Support views on DDL catalogs by @Jeadie in #10554
  • Update datafusion by @Jeadie in #10422
  • Improve full-text search indexing performance by @sgrebnov in #10464
  • feat(mysql): add mysql_zero_date_behavior parameter (null|error) by @phillipleblanc in #10573
  • fix(snowflake): declare private_key in connector PARAMETERS (fixes #10517) by @claudespice in #10559
  • Honour CARGO_TARGET_DIR in Makefiles by @Jeadie in #10569
  • Enable cosine_distance pushdown to DuckDB accelerator via array_cosine_distance by @sgrebnov in #10564
  • fix: Update test snapshots by @app/github-actions in #10570
  • fix: Update tpch benchmark snapshots by @app/github-actions in #10560
  • feat(snapshots): make snapshots an optional feature by @phillipleblanc in #10574
  • Enforce read-only API key restrictions on Flight DoGet and async query paths by @Jeadie in #10551
  • Improved security posture on Github workflows by @Jeadie in #10556
  • fix: Update datafusion-table-providers to improve SqlTable filter pushdown by @sgrebnov in #10595
  • feat(secrets): add HashiCorp Vault secret store by @phillipleblanc in #10561
  • fix: delegate update() in UpsertDedupTableProvider to inner provider by @claudespice in #10593
  • Add DuckDB vector engine support by @lukekim in #10562
  • Sharepoint - add object-store listing connector with expanded auth and write support by @lukekim in #10473
  • fix: Install protoc from source by @peasee in #10597
  • Enable DML support for PostgreSQL data connector by @phillipleblanc in #10446
  • feat(postgres): support inline PEM sslrootcert by @claudespice in #10578
  • Add foreign key metadata discovery to PostgreSQL Catalog by @sgrebnov in #10849
  • Add Snowflake DML support by @lukekim in #10747
  • Add MongoDB Change Streams support by @lukekim in #10813
  • Add user-defined functions by @lukekim in #10571
  • Add table user functions and gate HTTP servers by @lukekim in #10675
  • feat: add on-demand dataset loading by @phillipleblanc in #10629
  • feat(runtime): declared-schema deferred datasets by @phillipleblanc in #10669
  • feat(spicepod, runtime): add columns[].type / nullable + lenient type parser by @phillipleblanc in #10661
  • Replace external smb crate with internal SMB 3.1.1 client by @phillipleblanc in #10516
  • Add unified query cancellation across all paths by @lukekim in #10390
  • Add dynamic HTTP request headers by @lukekim in #10604
  • feat(http): Support dynamic HTTP connector request params from subqueries by @lukekim in #10636
  • feat(http): pass through HTTP metadata columns with JSON schema decomposition by @lukekim in #10679
  • Add nolimit HTTP pagination max pages by @lukekim in #10673
  • Add shared HTTP rate control for connectors by @lukekim in #10648
  • Use origin label instead of name for HTTP rate control metrics by @lukekim in #10689
  • fix(http): reject OR across different HTTP filter columns by @lukekim in #10625
  • Add provider-aware LLM prompt caching by @lukekim in #10645
  • Add searchable registry mode for LLM tools by @lukekim in #10647
  • feat: refresh_mode: snapshot + SQLite/Turso WAL flush + Cayenne metastore slice by @phillipleblanc in #10651
  • feat: per-principal cache namespacing for SQL/search/caching-accelerator by @lukekim in #10702
  • Add self-hosted Spice connector support by @phillipleblanc in #10546
  • Add Delta Lake Azure tenant parameter by @phillipleblanc in #10671
  • Support OAuth2 client credentials in 'spice cloud login' by @ewgenius in #10586
  • Add configurable allowed_hosts for MCP by @lukekim in #10638
  • fix: make Helm chart probes configurable by @peasee in #10696
  • Strip high-cardinality datasets dim from anonymous telemetry by @lukekim in #10711
  • feat(elasticsearch): direct FTS engine config + index lifecycle and ingestion controls by @lukekim in #10672
  • Add DuckDB HNSW vector index support for accelerated views by @sgrebnov in #10695
  • Rewrite DuckDB vector search SQL to activate HNSW_INDEX_SCAN by @sgrebnov in #10674
  • Fix DuckDB HNSW vector indexes lost after data refresh by @sgrebnov in #10668
  • Fix DuckDB DELETE/UPDATE on full and caching refresh mode datasets by @phillipleblanc in #10632
  • Fix DuckLake connector: downcast, module registration, schema discovery, and S3 credentials by @sgrebnov in #10650
  • Fix federation pushing denied functions inside subqueries to remote engines by @phillipleblanc in #10692
  • fix(caching): honour refresh_on_startup: always in caching mode by @phillipleblanc in #10594
  • fix(iceberg): rebuild storage factory when Hadoop catalog scheme is inferred by @sgrebnov in #10601
  • Pipeline CDC ingestion: overlap source reads with batch apply by @lukekim in #10676
  • fix: add NULL check to CDC primary key extraction by @lukekim in #10684
  • Properly handle nullability during CDC processing by @krinart in #10803
  • Flatten scheduler config and rename partition management โ†’ partition assignment by @lukekim in #10450
  • Improve NSQL UX and harden internal LLM tools by @lukekim in #10715
  • Support Responses API across model providers by @lukekim in #10724
  • Update xAI default model and handle Grok model retirements by @Jeadie in #10723
  • Improve cli table layout by @krinart in #10725
  • TLS cert hot-reload (mTLS plan M1) by @phillipleblanc in #10727
  • Fix DuckLake catalog include filter being ignored by @phillipleblanc in #10738
  • Promote DuckLake Catalog and Data Connector to Beta quality by @sgrebnov in #10743
  • feat(ducklake): Support INSERT on catalog tables with read_write access by @sgrebnov in #10744
  • perf(cdc): coalesce envelopes and overlap commits in apply pipeline by @lukekim in #10745
  • feat: Allow full version tags in spicepod version by @peasee in #10748
  • Add Arrow primary key upserts by @lukekim in #10749
  • fix(snapshot): keep refresh_mode snapshot read-only by @phillipleblanc in #10752
  • feat(tls): public mTLS for HTTP and Flight (channel + identity modes) by @phillipleblanc in #10753
  • perf(cayenne): lock-free deletion caches with bloom-prefiltered probe by @lukekim in #10756
  • fix(security): close API key timing-position leak and remote-UDF SSRF by @lukekim in #10757
  • Fix 'wait_until_dependent_tables_are_ready' for catalogs by @phillipleblanc in #10758
  • Fixes for views and resolved tables on 'spice refresh' CLI by @phillipleblanc in #10759
  • Implement FlightSQL CommandStatementSubstraitPlan support by @lukekim in #10761
  • feat(connectors): mTLS client cert support for flightsql and spiceai connectors by @phillipleblanc in #10764
  • Allow arbitrary filenames when specifying spicepod path + kind validation by @krinart in #10777
  • fix: ignore field metadata in schema compatibility check in index_table_scan by @Jeadie in #10778
  • Display pushed-down limits in EXPLAIN TREE output by @lukekim in #10779
  • fix: enable streaming append for Kafka with Cayenne accelerator by @lukekim in #10780
  • fix: bound chunked-index intermediate batch size to prevent OOM by @phillipleblanc in #10783
  • fix: label all columns in spice cloud metrics table output by @claudespice in #10784
  • fix: use checked arithmetic for Turso integer-millis timestamp read path by @claudespice in #10786
  • fix: use checked arithmetic in timestamp-to-nanosecond conversions by @claudespice in #10666
  • Upgrade to DuckDB v1.5.2 by @sgrebnov in #10788
  • Improve CDC ingestion performance by @lukekim in #10789
  • Fix tool_search/tool_invoke spans by @lukekim in #10791
  • Add Cayenne inline mutations and benchmark coverage by @lukekim in #10792
  • Ensure we always resolve table names in distributed mode/metadata by @Jeadie in #10793
  • Remove permanent errors from DynamoDB Streams by @krinart in #10794
  • Add expanded view mode for wide table display in SQL REPL by @lukekim in #10797
  • Fix Cayenne CDC schema mismatch error by @sgrebnov in #10800
  • Executors should create catalog tables on join by @Jeadie in #10807
  • Add compressed file support for listing connectors by @lukekim in #10809
  • Improve Cayenne mutation, scan, and inline memtable scaling by @lukekim in #10811
  • Add range fallback for large join filters by @lukekim in #10816
  • Improve Cayenne join filter pushdown by @lukekim in #10818
  • Synchronize Cayenne partition commits across partitions by @phillipleblanc in #10819
  • fix: Deny nondistributed cayenne catalog by @peasee in #10821
  • Enable parallel Cayenne Vortex writes by @lukekim in #10822
  • Expand Arrow type handling in formatting and Elasticsearch by @lukekim in #10825
  • Add response.output_text.delta to responses API by @krinart in #10828
  • feat(cayenne): add join filter propagation and no-spill Q21 planning by @lukekim in #10840
  • Upgrade Turso to v0.6.0 by @sgrebnov in #10843
  • feat(cli): add spice feedback command to open community Slack by @lukekim in #10856
  • Upgrade iceberg to v0.9.1 by @sgrebnov in #10859
  • feat(cluster): per-request executor readiness gate on /v1/ready by @phillipleblanc in #10860
  • fix: Require dim-side statistics for CayennePropagateFilterAcrossEquiJoinKeys by @sgrebnov in #10863
  • fix: Debezium schema evolution breaks dataset init on reload by @claudespice in #10144
  • fix(mssql): Push topK limit to SQL Server for non-nullable sort columns by @Jeadie in #10621
  • fix(ScyllaDB): disable physical filter pushdown by @sgrebnov in #10772
  • fix: handle typed NULLs and prevent overflow in DynamoDB DML type conversions by @krinart in #10511
  • fix: use InsertOp::Overwrite in DynamoDB bootstrap scan_and_overwrite_accelerator by @krinart in #10639
  • Improve DynamoDB Bootstrap performance by @krinart in #10616
  • fix: preserve field and schema metadata in Vortex type transformation by @lukekim in #10628
  • fix: GH connector - explicitly use AWS LC RS crypto provider for jwt by @phillipleblanc in #10619
  • fix: add snapshot mode guards to delete_from/update and delegate DML in SwappableTableProvider by @phillipleblanc in #10685
  • Persist HTTP rate-control state in object storage by @lukekim in #10697
  • Rate limit metrics HTTP endpoint by @lukekim in #10162
  • feat(geo): add optional spatial SQL UDF support by @lukekim in #10833
  • feat(cayenne): CDC throughput, compaction, scan caching, and benchmarks by @lukekim in #10852
  • fix(cayenne): fix Vortex panic on highly compressible data by @sgrebnov in #10855
  • fix(cayenne): Read live protected snapshots after cleanup grace period by @sgrebnov in #10901
  • fix: Disable Cayenne HashJoin rewriter optimizer by @sgrebnov in #10882
  • Fix GetFlightInfo vs DoGet Flight Schema by @krinart in #10864
  • fix(search): preserve column casing in /v1/search primary key plumbing by @claudespice in #10909
  • fix(object-store): dedupe s3 url style auto-detection log by @phillipleblanc in #10898
  • Improve Spice CLI manifest editing and direct command modes by @lukekim in #10815
  • Persist Kafka CDC offsets in sidecar tables by @lukekim in #10823
  • feat(task-history): record Ballista stages for distributed queries by @phillipleblanc in #10831
  • Add '#[deny(clippy::missing_trait_methods)]' to wrapper/delegation trait impls by @Jeadie in #10795
  • Optimize Cayenne catalog maintenance paths by @lukekim in #10904
  • Centralize DuckDB settings for accelerator by @ewgenius in #10895
  • deps(ballista): bump to 47e2b494 to fix S3 shuffle reads under cluster mode by @phillipleblanc in #10910
  • Authorization header + Bump async-openai + responses_adapter fix by @krinart in #10911
  • Tune accelerators by storage profile by @lukekim in #10913
  • feat: add dataset-level on_schema_change config by @lukekim in #10908
  • Handle NULL sentinel for nullable partition expressions by @Jeadie in #10880
  • fix: Remove Cayenne Catalog from catalog registration by @peasee in #10914
  • Add catalog name to foreign key metadata in postgres catalog by @Jeadie in #10917
  • Cayenne perf: eliminate redundant clones, PK point-lookup fanout fix, IN-list rewrite + microbench coverage by @lukekim in #10916
  • fix(turso-shared): retry on Turso BEGIN CONCURRENT "Write-write conflict" by @lukekim in #10946
  • Vendor Vortex DataFusion for Cayenne by @lukekim in #10933
  • perf(cayenne): background retention + enable CDC pipelining for retention-configured tables by @lukekim in #10936
  • feat(cayenne): scale metastore pool to 32 + vs_duckdb_scaling benches (1โ†’128 concurrency, sqlite + turso lanes) by @lukekim in #10943
  • feat(mcp): support auth for streamable HTTP tools by @phillipleblanc in #10927
  • Explicit error if v1/search requests a table without search index by @Jeadie in #10968
  • Fix spicepod loading failure when directory name contains dots by @sgrebnov in #10958
  • Extend append tests with arrow engine configurations by @sgrebnov in #10959
  • Remove dataset on_schema_change Policy from rc.5 release notes by @sgrebnov in #10964
  • Skip tpcds_q78 for Cayenne engine at SF100 by @sgrebnov in #10966
  • fix: Update benchmark snapshots May-20 by @app/github-actions in #10952
  • Fix #10951: UdtfExec invariant Vec lengths must match children count by @phillipleblanc in #10953
  • docs(release): update v2.0.0-rc.5 notes with latest trunk PRs by @lukekim in #10949
  • Remove eval related things for v2.0.0 by @Jeadie in #10945
  • build(deps): bump ubuntu from 24.04 to 26.04 in the docker-dependencies group by @app/dependabot in #10883
  • fix: Add publish = false to chbench-driver by @sgrebnov in #10939
  • [Bug] Timing between reconnect and AllocateInitialPartitions leaves connection without flight_sql_client by @Jeadie in #10805
  • Fix: refresh_mode: snapshot reports Ready with empty data when no snapshot exists by @sgrebnov in #10979
  • fix(cluster): gate scheduler readiness on executor partition loads by @phillipleblanc in #10992
  • fix: handle EXISTS/NOT EXISTS subqueries in federation analyzer by @sgrebnov in #10996
  • Refactor spice dataset configuration command by @Jeadie in #10999
  • fix: preserve field and schema metadata in Vortex physical schema calculation by @claudespice in #11013
  • fix: validate Snowflake account identifiers and auth config by @Jeadie in #11024
  • Fix Unity Catalog connector deserialization failure with OSS Unity Catalog by @ewgenius in #11026
  • feat(cayenne): allow inline writes with pending deletions (deletes/upserts) by @sgrebnov in #11031
  • Expose metadata descriptions via PostgreSQL UDFs by @lukekim in #11032
  • Remove default runtime features - enable explicitly in spiced by @phillipleblanc in #11037
  • feat(cayenne): fast-path CDC deletes by extracting PK values from filters by @sgrebnov in #11049
  • Cayenne optimizer rules: auto relevance test for q21-shape (all-Cayenne CH-Bench) and runtime rule selection by @lukekim in #11050
  • refactor(cdc): reduce CDC sub-batch splits for interleaved upsert/delete workloads by @sgrebnov in #11051
  • fix(snowflake): enforce function deny-list in federation pushdown by @claudespice in #11057
  • fix(mcp): trace external server tool calls in task history by @ewgenius in #11058
  • perf(cdc): Last-write-wins dedup in group_into_sub_batches to reduce sub-batch splits by @sgrebnov in #11059
  • PM edits to v2.0.0-rc5 by @lukekim in #11067
  • fix(snowflake): wire deny-list in extracted connector crate (#10703) by @claudespice in #11071
  • perf(cayenne): keep CDC upsert PK keysets resident to avoid per-batch full-table rebuilds by @lukekim in #11074
  • Fix metadata on search indexing by @Jeadie in #11080
  • feat(cayenne): merge-on-read position deletes for PK upsert tables + memory-pool accounting by @lukekim in #11085
  • perf(cayenne): scale CDC inline flush caps with memory + storage class by @lukekim in #11087
  • feat(cluster): report per-executor table statistics so distributed JoinSelection can size joins by @phillipleblanc in #11089
  • Improve Cayenne CDC write and compaction path tracing by @sgrebnov in #11091
  • Support tuple-IN composite PK extraction in Cayenne delete fast-path by @sgrebnov in #11093
  • feat(cluster): NDV-aware executor stats so CDC q18 join swap fires by @phillipleblanc in #11098
  • feat(cayenne): maintain join-sizing stats on the write path by @phillipleblanc in #11104
  • fix(cache): run periodic moka maintenance for idle caches by @phillipleblanc in #11106
  • Upgrade to DuckDB 1.5.3 + statically link the VSS (HNSW) extension by @sgrebnov in #11107
  • Fix fetched_at for HTTP connector by @Jeadie in #11116
  • fix(cayenne): tombstone inline-checkpointed rows on upsert to prevent duplicate PKs by @sgrebnov in #11129
  • feat: dedicated compaction runtime for Cayenne + CDC pipelining, protected snapshots, and test coverage by @lukekim in #11130
  • Add datasets dimension to the query_executions metric by @phillipleblanc in #11138
  • Fix #11137: localpod child not tracking parent refreshes with in-memory (arrow) parent accelerator by @phillipleblanc in #11139
  • Fix Windows build: vendor the VSS extension (drop nested submodule) by @phillipleblanc in #11140
  • fix(spiceai): keep correlated subqueries out of JOIN ON for Spice Cloud federation by @phillipleblanc in #11143
  • Refactor spice dataset configuration command by @Jeadie in #10999
  • feat(cayenne): sharded parallel Vortex encode with key/time clustering by @lukekim in #11144
  • fix(cluster): prevent DoPut write pipeline self-deadlock under ingest backpressure by @phillipleblanc in #11160
  • fix(cayenne): only warn on genuine protected-snapshot amplification by @lukekim in #11158

Full Changelog: https://github.com/spiceai/spiceai/compare/v1.11.6...v2.0.0