🧑🍳 Spice.ai OSS Cookbook
122 guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more.
Contribute to the Cookbook on GitHub!
Sample Applications and Guides
Example apps and guides for real-world Spice.ai usage and best practices.
Core Features
Start with the core capabilities: federated SQL query, data acceleration, async queries, hybrid search, and AI inference in SQL. Join data from S3 and PostgreSQL in a single SQL query, then accelerate both locally. Combine full-text and vector search using Reciprocal Rank Fusion (RRF) for improved search results.Federated SQL Query
Hybrid-Search with RRF
Models, AI, and Agents
Connect to hosted and local AI models, and build intelligent agents using Spice.ai. Use Azure OpenAI models for vector search and chat over structured and unstructured data. Ask natural language (NLP) questions of your datasets using the built-in text-to-SQL tool. Generate SQL queries and interactive charts from natural language questions. Deploy Nvidia NIM infrastructure on Kubernetes with GPUs, connected to Spice. Deploy Nvidia NIM on a GPU-optimized AWS EC2 instance, connected to Spice. Run Spice as an MCP server and connect AI assistants such as Claude Desktop, Cursor, or VS Code.Azure OpenAI Models
Text to SQL (NSQL)
Generative Visualizations
Nvidia NIM on Kubernetes
Nvidia NIM on AWS EC2
Spice as an MCP Server
Data Acceleration, Materialization, and Federation
Optimize query performance with local acceleration, data materialization, and federation techniques. Materialize data into an attached PostgreSQL instance. Available in Spice.ai Enterprise. Prune data on categorical columns, such as IDs, using hashed partitioning. Partition accelerated datasets so queries skip partitions they do not need. Bootstrap accelerations from snapshots in object storage to skip cold starts. Available in Spice.ai Enterprise. Serve queries immediately while a large table accelerates in the background. Commit gated, serializable transactions on Cayenne tables and write the results back to PostgreSQL.PostgreSQL Data Accelerator
Hashed Partitioning with Cayenne
Dataset Partitioning
Acceleration Snapshots
Dual-Dataset Registration
Serializable Transactions
Change Data Capture (CDC)
Stream inserts, updates, and deletes from source databases to keep accelerated datasets current. Stream changes from PostgreSQL using logical replication, without Debezium or Kafka. Discover every table in a PostgreSQL database and keep a local copy of each current with CDC. Stream inserts, updates, and deletes from a DynamoDB table using DynamoDB Streams. Stream MySQL changes using Debezium with SASL/SCRAM authentication.PostgreSQL CDC
PostgreSQL Catalog CDC
DynamoDB Streams
Debezium CDC with SASL/SCRAM
Search & Embeddings
Search data with full-text, vector, and hybrid search, using embeddings and vector engines. Combine full-text and vector search using Reciprocal Rank Fusion (RRF) for improved search results. Use Elasticsearch for both full-text and vector search. Available in Spice.ai Enterprise.Hybrid-Search with RRF
Elasticsearch Full-Text and Vector Search
Data Connectors
Connect to databases, data warehouses, data lakes, and APIs, and query them with SQL. Connect to and query Databricks instances using Delta Lake or Spark Connect. Query Elasticsearch indices using federated SQL. Available in Spice.ai Enterprise. Query ScyllaDB clusters using federated SQL. Available in Spice.ai Enterprise. Combine real-time data streaming from Kafka with other datasets using Spice.Databricks Connector
Elasticsearch Connector
ScyllaDB Connector
Live Orders Analytics with Apache Kafka Data Connector
Catalog Connectors
Connect to data catalogs to discover and query their tables. Discover every table in a PostgreSQL database and keep a local copy of each current with CDC. Connect to Iceberg catalog with support for reading and writing Iceberg tables. Connect to Iceberg Hadoop catalogs, locally or on S3-compatible object storage. Discover and query all schemas and tables in a SQL Server database.PostgreSQL Catalog CDC
Iceberg Catalog Connector
Iceberg Hadoop Catalog
Microsoft SQL Server Catalog
Visualization
Visualize data with BI and analytics tools.
API Clients
Use API clients for data access and integration.
Deployment
Deploy Spice.ai in different environments. Run Spice alongside the application on the same host for low-latency access. Run Spice as an independent service, optionally with replicas behind a load balancer. Connect a local Spice instance to Spice Cloud, deploy changes without restarting, and deliver secrets.Sidecar Deployment Architecture
Microservice Deployment Architecture
Cloud Connect on a Development Machine
Performance and Benchmarking
Measure and optimize performance with benchmarks and best practices for your Spice.ai deployment.
Configuration
Configure data refresh, retention, scheduling, and data quality for accelerated datasets.
SDKs
Use SDKs for different programming languages. Query Spice.ai using the JavaScript (Node.js) SDK with examples.Spice.js JavaScript (Node.js) SDK
Security
Secure your Spice.ai deployment and data access with encryption, authentication, and authorization. Authenticate clients with certificates using mutual TLS. Available in Spice.ai Enterprise. Add multi-tenancy, row-level security, PII masking, and RBAC with Cedar policies. Available in Spice.ai Enterprise.Mutual TLS (mTLS)
Authorization
Advanced Topics
Replicate datasets locally, distribute queries across nodes, and work with JSON data.
