3 Sources
[1]
Snowflake adds AI model routing to cut costs | VentureBeat
Enterprise teams running AI agents at scale are finding that a single model handles every task poorly -- either the model is too expensive for simple questions or not capable enough for hard ones. Model routing, which picks the right model for each task automatically, is becoming the
[2]
Snowflake Unlocks Better AI Economics with Dynamic Model Routing, Delivering More Value to Customers
New capabilities in Cortex AI Gateway help enterprises lower AI costs with targeted model choice for each workload while keeping data governed in Snowflake Snowflake today announced dynamic model routing2 within Cortex AI Gateway and Snowflake's flagship AI products, alongside expanded access to
[3]
Snowflake Announces Dynamic Model Routing and Expands Access to Deepseek-V4-Flash 0731 and Glm-5.3 in Cortex Ai Gateway
Snowflake announced dynamic model routing within Cortex AI Gateway and Snowflake?s flagship AI products, alongside expanded access to leading open models. New capabilities in Cortex AI Gateway help enterprises lower AI costs with targeted model choice for each workload while keeping data governed
Share
Copy Link
Snowflake unveiled dynamic model routing within Cortex AI Gateway, enabling enterprises to automatically select the most cost-effective AI model for each task. Internal testing shows up to 3x greater token efficiency while maintaining quality. The capability addresses rising AI costs by routing simple queries to efficient models and complex tasks to frontier models.
Snowflake announced dynamic model routing within Cortex AI Gateway, a capability designed to automatically select the most cost-effective AI model for each task without compromising quality
1
2
. The move addresses a critical pain point for enterprise AI deployments: teams running AI agents at scale find that using a single model for every task either proves too expensive for simple questions or insufficiently capable for complex ones1
. Snowflake's internal testing indicates the capability can reduce token costs by as much as 3x on some workloads, after discovering that simple questions were often routed to the most capable model, making responses unnecessarily expensive and slower1
. In one evaluation, agents using dynamic model routing with Cortex AI Gateway built a dbt pipeline with up to 3x greater token efficiency than a frontier-model-only approach while maintaining the same quality2
3
. A separate test showed engineering teams completed the same number of pull requests with 25% greater token efficiency2
3
.The capability builds on Cortex AI Gateway, which Snowflake launched in July 2026 as a governance layer for agent and model traffic
1
2
. Enterprises can now select "auto" instead of a fixed model, and the system routes each task to whichever model offers the best combination of quality and cost1
. Dynamic model routing operates through two mechanisms, according to Baris Gultekin, vice president of AI at Snowflake1
. Under an advisor pattern, a smaller model attempts a task first, and if it cannot finish the job, it calls a larger model as a tool and continues from there1
. A separate classifier, trained on past queries, automatically routes straightforward questions to simpler models1
. This new capability directs lower-complexity or repetitive tasks to more efficient models, while work that requires deeper reasoning is routed to frontier models2
3
. Customers retain control: auto routing is optional, and they can restrict routing to one model or a defined set of models1
. Snowflake prices AI purely on token usage, with no separate fee for the routing decision itself1
.
Source: CXOToday
Dynamic model routing is integrated across Snowflake's flagship AI products, including Snowflake CoCo and Snowflake CoWork, and is available to third-party AI agents using Cortex AI Gateway
2
3
. Snowflake ties routing to the same access controls it already uses for data governance1
. Governance starts at the data level with role-based access controls, extends to models where customer roles map to buckets of approved models, and extends again to agents, where an agent can be restricted to narrower privileges than the user invoking it1
. "For high quality, enterprise grade agents to be built, it's crucial to get the context and the governance right," Gultekin told VentureBeat. "Context, trust and model choice all go hand in hand"1
. Open models can run from a customer's own region to satisfy data residency requirements, with all inference staying inside Snowflake's security boundary rather than routing out to an external provider1
. Cortex AI Gateway gives administrators visibility into token usage and costs, while allowing organizations to establish spending limits across AI apps and agents3
. Snowflake CoCo extends those controls through role-based access and tagging framework, allowing administrators to set default models, attribute usage to teams or cost centers, establish per-user quotas, and receive notifications as consumption approaches defined limits3
.Snowflake is expanding customers' access to leading open models, including DeepSeek-V4-Flash 0731 and GLM-5.3, through Snowflake Cortex AI
2
3
. This adds to Snowflake's extensive model library, giving customers more options to balance model quality and cost while keeping governed data secure within Snowflake2
. Both DeepSeek-V4-Flash and GLM-5.3 are developed in China, and the regional setup matters specifically for open models with non-U.S. origins1
. Snowflake's AI Research Team evaluated DeepSeek-V4-Flash on enterprise-focused tasks, with recent testing showing that DeepSeek-V4-Flash outperformed the proprietary models used for the evaluation, scoring 74.4% on data engineering tasks3
. GLM-5.3 also performed strongly at 62.8%, while using fewer tokens than any other model tested3
. These results demonstrate that open models are increasingly capable of supporting mission-critical tasks at a lower cost3
. Snowflake continuously evaluates and optimizes how these models are served and used across its AI products, helping customers take advantage of advances in the open model ecosystem without having to repeatedly rework their apps or infrastructure3
.Snowflake recently announced its Horizon Context and Cortex Sense tools that provide context capabilities
1
. Without good context, a model has to do the exploratory work itself, writing and testing SQL, searching through data and retrying when something does not work1
. Gultekin explained that the process is expensive, and getting it right typically requires a more capable model1
. Packaging the context in advance removes that exploratory step, which means a simpler, cheaper model can often handle the same task1
. Snowflake also builds agent memory into that context, and as an agent is used repeatedly, its memory updates and gets folded back into future queries1
. The system does not re-solve the same problem from scratch each time, and memory becomes part of the context passed to the model1
. Snowflake's recent acquisition of Natoma adds another layer, bringing more than 100 MCP connectors with scoped, governed access1
. An agent could get read-only access to a connected tool like email, for example, rather than broader permissions1
.Related Stories
The move lands amid a broader industry shift toward automated model routing
1
. Databricks, AWS, Google Cloud and Nvidia have all announced some form of model routing technology1
. Databricks has an offering with Smart Routing for its Unity AI Gateway, while Nvidia on August 11 announced Switchyard as a technology layer to help route AI model choice1
. OpenRouter is one of the most widely known options, providing a platform that enables organizations to route based on cost and performance1
. Snowflake argues that model routing is more complex than just price and performance—it's also about governance and context1
. "The interesting part is what it says about where differentiation has moved," Sanjeev Mohan, Principal and Founder, SanjMo, told VentureBeat. "Snowflake isn't really selling routing, it's selling routing that never leaves the governed data"1
. Mohan also noted that "enterprises are drowning in model choices, but the real problem isn't which model to pick. It's the operational overhead of picking the right one for every task, at scale"2
. He added that "the ability to match workload complexity to model cost, without rebuilding your infrastructure every time a new model drops, is exactly the kind of efficiency enterprises need to move from AI experimentation to AI at scale"2
."Enterprises are becoming much more rigorous about the economics of AI. The question is no longer how much AI they are using, but whether that AI is translating into meaningful business value," said Sridhar Ramaswamy, CEO, Snowflake
2
. He added that "achieving intelligence efficiency requires the flexibility to use the best model for each task as the landscape evolves. Snowflake's role is to absorb that complexity so customers can focus on outcomes while we optimize model choice underneath"2
. As model performance and pricing change, Cortex AI Gateway can update routing decisions across Snowflake's AI products like Snowflake CoCo and Snowflake CoWork so customers don't need to rebuild their apps or agents2
3
. This allows enterprises to take advantage of new model options as they emerge while Snowflake manages the complexity of model choice and optimization underneath2
. Organizations can automatically match workloads to the right models, take advantage of new model options as they emerge, and manage how AI resources are consumed across the business3
. Watch for enterprises to scrutinize token efficiency metrics more closely as they scale AI deployments, and expect further competition among cloud providers to offer similar routing capabilities with tighter governance controls.Summarized by
Navi
[1]
[2]
28 Jul 2026•Technology

02 Oct 2025•Technology

03 Jun 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
