Blog/Product and Technology/Building the Modern AI Infrastructure Stack with Cortex AI Gateway
Oct 8, 2026/9 min readProduct and Technology

Building the Modern AI Infrastructure Stack with Cortex AI Gateway

Over the last few years, the enterprise AI stack has evolved into a sophisticated but fragmented and unwieldy ecosystem. It now encompasses everything including vector databases, monitoring systems and pipeline tuning, connected through a sprawling mix of coding agents, embedded agents and Model Context Protocol (MCP) servers interfacing with enterprise systems — all supported by a model portfolio that shifts every few weeks. Different teams adopt pieces of the stack on different timelines, potentially creating a management and governance gap that can impact security. 

In this fractured environment, few enterprises have the full context they need to accurately answer basic questions across these surfaces: Which models are we calling, who is calling them, what actions are they taking, and what is the associated cost? 

The instinct in response is to restrict, but this isn’t a viable long-term solution. Hindering employees’ ability to use AI will only lead them to resort to unsanctioned AI tools, which amplifies the risk even further.

That gap isn’t a tooling problem; it’s a missing layer. The modern AI stack needs a control point that sits between every AI client and the enterprise systems they reach. Cortex AI Gateway (public preview3) addresses this by acting as that centralized control plane, providing the unified visibility, governance and oversight that this modern, multifaceted infrastructure requires. In our view, six elements are nonnegotiable in the modern enterprise’s AI infrastructure stack:

  • Unified, governed inference 

  • Dynamic model routing 

  • MCP and tool governance

  • Central observability 

  • Security policies and guardrails 

  • Cost governance and controls

Now in public preview, Cortex AI Gateway gives platform teams one place to grant access to models and MCP servers, attribute and control spend, and see what AI has done on behalf of their users. Application teams get a single endpoint to build against. Let’s take a look at how it maps onto each layer of the stack.

Cortex AI Gateway is Snowflake's centralized control plane for enterprise AI infrastructure. It provides a single endpoint for governed model inference, dynamic model routing, MCP tool governance, observability, security guardrails and cost controls — giving platform teams unified visibility across every AI client and agent in the organization.

Unified, governed inference: Use one endpoint for your agents

The first thing that breaks in a growing AI stack is the number of paths to a model: An engineer configures a coding agent against one provider. A platform team wires an application to a second. Someone else scripts against a third. Each path has its own credential, its own log and its own bill.

Cortex AI Gateway collapses those into one endpoint. It implements the Chat Completions API for OpenAI, Grok, GLM, Gemini, Llama, Mistral and DeepSeek models and the Messages API for Claude, which are the two protocols model serving has converged on. Streaming, tool calling, structured output, prompt caching, image input and reasoning all work the same way they do natively, so pointing an existing client at Cortex AI Gateway requires just another base URL and a token. Support for the OpenAI Responses API is also coming very soon. The following example shows all the setup you need to configure OpenCode with Cortex AI Gateway:

export PAT="<pat>"
export BASE_URL="<base_url>"
export OPENCODE_CONFIG_CONTENT='{
  "$schema": "https://opencode.ai/config.json",
  "model": "snowflake-cortex/openai-gpt-5.4",
  "small_model": "snowflake-cortex/openai-gpt-5.4",
  "provider": {
    "snowflake-cortex": {
      "options": {
        "baseURL": "${BASE_URL}",
        "account": "<ACCOUNT>.snowflakecomputing.com",
        "apiKey": "<PAT>",
        "headers": {
          "snow-agent-name": "opencode"
        }
      }
    }
  }
}'

Access control uses Snowflake’s existing role-based access control (RBAC) framework. An admin can configure what models are available through the gateway, making sure only those approved by an organization are accessible.

To make it even easier to use coding agents with Cortex AI Gateway, the Snowflake CLI will be able to take care of setting up your traffic to go through the gateway. One command is all it takes to read credentials from your active Snowflake connection, inject them into the coding agent’s config and launch the agent with its traffic routed through Cortex AI Gateway. For example, launching OpenCode using the Snowflake CLI is now as simple as running snow ai opencode.

Dynamic model routing in Cortex AI Gateway: Pick the right model for each task

Most teams have settled on a single frontier model for everything, which is a reasonable way to start and an expensive way to scale. A large share of agent steps are file edits, formatting and lookups that a cheaper model handles with the same quality and accuracy.

Cortex AI Gateway offers dynamic model routing (private preview soon).
Figure 1: Cortex AI Gateway offers dynamic model routing (private preview soon).

 

Dynamic model routing, coming to private preview soon through Cortex AI Gateway, selects the most affordable model (from the available options) that can confidently complete the task at each step of agent execution. It’s been fine-tuned by the Snowflake research team, so lower-complexity and repetitive work goes to efficient models, and work that needs deeper reasoning goes to frontier models. In an internal evaluation, it completed a dbt pipeline workload with up to 3x greater token efficiency than a frontier-model-only approach at comparable quality, and in a separate coding test engineering teams held pull-request throughput steady while using roughly 25% fewer tokens.1 Those results can translate into real-world cost savings for enterprises.

Routing is governed just like manual model selection. Only approved models are used, existing data residency settings are respected, and every routing decision is logged, so a compliance team can see which model handled which request. Because routing lives in Cortex AI Gateway rather than in each application, Snowflake can update decisions as model pricing and performance change without agent configurations ever having to change.

The open source model side matters here too. Evaluated on ADE-bench using Snowflake CoCo as the agent harness, DeepSeek-V4-Flash scores 74.4%,2 ahead of the leading proprietary model we tested. GLM-5.3 is coming soon to private preview, but its predecessor GLM-5.2 scored 66% on ADE-bench2 with the lowest token footprint of any model in the benchmark,2 which is exactly the profile a high-volume workload wants. Snowflake serves these models within the secure Snowflake perimeter, so inference happens next to governed data, inside the same RBAC and audit trail that already covers it.

Learn more about dynamic model routing.

MCP and tool governance: Control what your agents can do

The moment an agent can call tools, it becomes an actor in the enterprise — and MCP has made that trivially easy. Every unmanaged MCP server an employee installs is a new path into production systems, operated by something that decides at runtime what to do with the access it has. This is shadow IT, except the unauthorized software can read production databases and modify records in enterprise SaaS.

An API gateway asks whether a client is authorized to hit an endpoint. A tool gateway has to answer a different question: Is this agent authorized to perform this tool call, on behalf of this user, right now? Following Snowflake's acquisition of Natoma, Tools in Cortex AI Gateway answers it. Now in private preview, it ships with a curated catalog of more than 100 MCP servers, each with OAuth handled for you, which helps eliminate the sprawl of MCP servers. 

Administrators choose between enabling specific tools, all currently published tools or all current and future tools. Tools that are not enabled are not visible to the client and cannot be invoked, so a denied capability is usually never advertised rather than failing at call time. Tool calls can be logged so that a full agent sequence from one prompt reads as a sequence. Tool calls also emit spans into the same trace table as inference, so a single trace covers the reasoning and the actions.

That audit trail is the defense against attacks specific to this layer. Tool poisoning, tool shadowing and rug pulls, where an approved server quietly updates a tool's definition to something malicious and clients pick up the change automatically, are invisible to end users by design. They are not invisible to a system that records every tool definition and every call.

Observability: See your agents’ traces in one place

Ask organizations which agents are running against their systems, and the honest answer most of the time comes from a spreadsheet of agents that’s out of date. This is reinforced by results from a survey we conducted, where roughly one in three customers surveyed used custom-built internal dashboards for agent observability. A survey by the Cloud Security Alliance found that in the past year, 82% of organizations have discovered at least one AI agent or autonomous workflow that security or IT did not previously know about. Furthermore, organizations with high shadow AI use saw an average breach cost of $5.39 million, with shadow AI incidents affecting 43% of breached organizations according to the Cost of a Data Breach Report 2026. 

The usual response is to instrument every client, which fails for the ordinary reason that you cannot instrument what you don't know about. Cortex AI Gateway takes the other route and makes observability a property of the serving layer. A model or MCP tool call that traverses the gateway, whether it’s from an agent running locally on a developer’s laptop or in a remote Kubernetes cluster, is recorded as an OpenTelemetry span in Snowflake. This gives you full execution visibility across third-party runtimes without modifying agent code.

Those spans land in the AGENT_TRACE_TABLE using OpenTelemetry generative AI semantic conventions. Each record carries the model requested, input and output tokens, cache reads and writes, duration, status code, MCP servers and tools, and the client that made the call down to its version and session. Prompts, responses and tool call inputs and outputs are captured only when an administrator turns payload capture on, and when they do, that text sits inside Snowflake under the access control the organization already runs. Because telemetry lands directly inside customer-owned event tables, such as AGENT_TRACE_TABLE, it remains strictly within your Snowflake security boundary.

Cortex AI Gateway also accepts client-side OTLP traces at its own endpoint. Visibility into both server-side and client-side traces provides the level of introspection that any leader in agentic observability needs.

Using CoCo makes investigating the telemetry in AGENT_TRACE_TABLE a one-prompt step. The example below showcases how you can find drops in cache hit rate and identify opportunities to lower your model costs.

Security and guardrails: Protect your most crucial surfaces

Agentic AI expanded the attack surface in a way that is genuinely new, and the clearest evidence comes from the incidents that have made headlines in the last few months.

The conclusion that’s easy to draw from these incidents: Guardrails are essential to a secure AI stack. They belong at the boundary the agent has to cross, not inside the agent, because the agent's own judgment is exactly the thing that fails.

Snowflake Cortex AI Guardrails already provide runtime protection against prompt injection and jailbreak attempts across CoCo, Snowflake CoWork and Cortex Agents. They scan tool outputs for indirect prompt injections that try to override system instructions, detect attempts to bypass model safety boundaries and identify previously unseen attack patterns in real time.

Cortex AI Gateway is bringing this infrastructure to gateway traffic very soon. It extends protection from Snowflake's own AI surfaces to any third-party agent and harness routed through the gateway, minimizing the exposed attack surface.

Review the Cortex AI Guardrails docs for more details.

Cost governance: Make every AI dollar count

AI spend broke a lot of budgets in 2026, but the real issue is poor forecasting, not extravagance. Alongside that came tokenmaxxing, the practice of treating token volume as a proxy for productivity. Some companies published internal leaderboards that they eventually shut down because users would attempt to reach the top of these leaderboards by overspending on AI. It became evident quite quickly that higher token usage does not generally equate to a higher ROI.

Snowflake's existing cost controls now extend to Cortex AI Gateway in an effort to combat this. 

Budgets provide established cost governance controls, and they are now available to any AI traffic routed through Cortex AI Gateway. The level of control really comes down to the administrator. Budgets can be set at a gateway level to keep things granular, or through resource and user tags to slice and dice cost controls in whatever dimension required. Whether it’s by department, role, cost center, user ID, agent or application, Cortex AI Gateway will support cost visibility and controls across all of the above. 

Per-user quotas allow setting monthly and optional daily limits per individual and enforce blocks within minutes. 

Gateway traffic reports into AI_GATEWAY_USAGE_HISTORY with credits broken out per model and a request ID that joins directly to the AGENT_TRACE_TABLE. This opens up fine-grained visibility into exactly what agent conversations, workflows, tasks, model requests and tool calls led to runaway spend. CoCo is a great way for you to easily dig into your trace data and understand your AI cost.

Learn more about AI cost governance.

How to get started with Cortex AI Gateway

Snowflake is dedicated to making AI safe and economical for enterprises, and Cortex AI Gateway is the starting point. Experience it for yourself: Point one client at Cortex AI Gateway and look at the traces, start implementing granular cost controls, or dig into model usage trends.

Cortex AI Gateway is in public preview in AWS commercial regions.3 Review our Cortex AI Gateway docs to learn more about its features and functionality and how to get started. 

 

This content contains forward-looking statements, including about our future product offerings, and are not commitments to deliver any product offerings. Actual results and offerings may differ and are subject to known and unknown risk and uncertainties. See our latest 10-Q for more information.

1 Token-efficiency and ADE-bench results are based on Snowflake internal testing; methodology and conditions available upon request, and individual results may vary by workload and configuration. https://www.snowflake.com/en/blog/dynamic-model-routing-open-models-cortex-ai/

2 Efficiency score based on internal testing using ADE-bench, a framework created by dbt for evaluating AI agents on real-world analytics and data engineering tasks.

3 Available to accounts in Amazon Web Services (AWS) commercial regions only, excluding Asia Pacific (New Zealand), Asia Pacific (Malaysia) and Europe (Spain).

Share this post

Subscribe to our blog newsletter

Get the best, coolest and latest delivered to your inbox each week

Where Data Does More