Snowflake Separates Batch Inference from the Warehouse — SPCS + Ray Jobs
Snowflake’s job-based batch inference on Snowpark Container Services with Ray lets large-scale ML/LLM inference run as dedicated distributed jobs — not inside interactive warehouses. The trade is performance and cost control for heavy batches while keeping data and compute in the data cloud; claims are vendor architecture, not third-party latency benches.

As models grow, inference stops fitting neatly inside interactive warehouse workloads. Snowflake’s move is to run heavy batch inference as a separate job — keep data in-place, pull compute out of the query path.
Snowflake (May 20) announced support for job-based batch inference, enabling distributed, dedicated inference workloads on Snowpark Container Services (SPCS) using Ray. Users can run large-scale inference as a separate workload for better performance and cost efficiency on complex models and unstructured data.
What shipped
- Dedicated jobs: Inference runs independently on SPCS with Ray for distribution.
- Scale: Batch workloads for ML models, including LLMs.
- Integration: Works with Snowflake ML for training and serving.
- Benefit pitch: Better performance for heavy inference; separation from interactive queries.
Claims vs checks
Architecture and feature availability are Snowflake primary. Performance and cost-efficiency benefits are vendor claims without published independent latency/$ benchmarks in the announcement. Treat “better performance and cost efficiency” as design intent until customer case numbers appear.
Limits
- No public head-to-head vs Databricks / SageMaker Batch / Vertex batch in the launch post.
- Pricing for SPCS Ray jobs not restated as a full TCO table here.
- Unstructured-data examples are illustrative, not audited throughput.
Sources
- Snowflake: “How Snowflake Executes Distributed Batch Inference Workloads at Scale” (May 20, 2026).
- Related Snowflake ML documentation and coverage confirming the May 20 launch.