Expedition. Free, virtual, Nov 3–6.

Technical tracks for practitioners, outcomes for leaders.

How to Build an AI Agent: 8 Steps to Production

This eight-step guide moves past the usual AI agent overview to show what teams actually have to engineer for production, including context, tools, memory, evaluation, cost and access controls.

A prototype AI agent only has to show that it can complete the intended workflow. Production asks a much broader question: can it keep doing that reliably across changing data, different users, varying permissions and a wider range of real-world requests?

That gap helps explain why so many agent projects stall after the demo. In a 2026 Snowflake-sponsored study conducted by Omdia, organizations reported an average 49% ROI from generative and agentic AI, even as executives expected roughly 41% of the agentic initiatives they sponsor to fail over the following three years. Teams are getting value from agents, but they’re also discovering which projects can support real workloads and which ones can’t.

Getting an agent through the pilot-to-production transition requires a series of engineering decisions that rarely show up in quick-start tutorials. The task needs a measurable boundary. The agent needs enough governed context to interpret the data correctly, permissions that constrain what it can access and change, and evaluation that shows where a run went wrong rather than simply whether the final answer looked right. Cost enters the picture too, because one request might trigger several rounds of reasoning, retrieval and tool calls before the user sees a result.

This guide walks through those decisions in the order they usually surface when an AI agent moves from a working prototype toward production.

What you’re actually building when you build an AI agent

When you build an AI agent, you’re building a system that coordinates model reasoning with context, actions, state and the sequence of decisions required to complete a task.

Most agent architectures use an LLM as the reasoning engine. But instead of taking one input and producing one response, the model works through a loop: interpret the goal, choose or develop a plan, select a tool, inspect the result and decide what to do next.

Tools give that reasoning loop something to act on. Depending on the application, the agent might query a database, search documents, invoke an API, run code or update another system.

The agent’s ability to take action has practical consequences when it’s given permission to modify records, send messages or trigger transactions, so defining what it’s allowed to do — and when human approval is required — is part of the build.

See how data, application development and agent architecture come together in an end-to-end agentic artificial intelligence workflow:

Components of an AI agent

Production agents typically use the same six broad components: a model, memory, tools, orchestration, evaluation and observability. The model handles reasoning, while memory keeps information the agent needs during or across tasks. Tools connect the agent to data and external systems. Orchestration controls the sequence of decisions and handoffs. Evaluation checks whether the run met the expected standard, and observability records the traces, latency, token consumption and tool activity engineers need when something goes wrong.

In practice, model selection is often one of the simpler decisions. Teams spend much more engineering time preparing business context, deciding which tools the agent can use, determining what should persist in memory and building tests that reflect the actual workflow.

Step 1: Scope the agent to one task with a measurable baseline

Start with a named workflow and record how that workflow performs today. For a support agent, for example, useful measures might include average resolution time, escalation rate and first-response accuracy. For an analytics agent, the team might measure analyst turnaround time and accuracy against a fixed set of known business questions.

A concrete baseline gives the project a target to beat. Suppose analysts currently answer 85% of a defined question set correctly in 15 minutes. The agent now has an accuracy target and a time target.

Define the boundaries just as clearly. What inputs will the agent receive? Which actions is it allowed to take? What will count as success?

It also helps to define a stopping condition before development starts. The team might decide that accuracy below an agreed threshold after several iterations, cost above a certain limit or an inability to constrain a high-risk action will halt the project. If the pilot ends, there’s a documented technical reason rather than an open-ended sense that the team should keep trying.

QUICK TIP

Write down the baseline and the condition that would stop the project before you start building. If the existing workflow is 85% accurate, for example, define what accuracy the agent needs to reach, how much it can cost and which failures would prevent a broader rollout. That gives the pilot something concrete to prove.

Step 2: Ground the agent in governed context

Giving an agent access to a database simply makes the data available to query. It doesn’t tell the agent what that data means to the business. The agent needs context to understand what fields mean and how they should be used.

Take a revenue question. The agent might find several columns named revenue, but it needs to be told whether the business means recognized revenue, booked revenue or annual recurring revenue. It also needs the right date field, the correct entity relationships and any exclusions that apply to the calculation.

Semantic context gives the agent those definitions and relationships, including business entities, metrics, joins and authoritative sources. Catalog metadata, lineage and freshness information fill in other parts of the picture: lineage shows where a value came from, while freshness helps determine whether the source is still appropriate for the question being asked.

The quality of that context affects both accuracy and cost. In a Snowflake internal benchmark of 58 product analytics questions, Snowflake CoWork using Cortex Sense (currently in private preview) reached 86.3% accuracy at $0.59 per query. An agent querying Snowflake through MCP without that supplied context reached 24.1% accuracy at $1.76 per query.

The reason shows up in the agent’s behavior. When business and data context is available up front, the agent spends fewer steps inspecting schemas, testing relationships and searching for the right source.

Data quality has an outsized impact on agent results. Automated reasoning will keep working with whatever information the system supplies unless the agent has some way to recognize that the source is no longer appropriate.

Step 3: Give it tools and constrain every tool call

Tools are where an agent’s reasoning turns into work: querying data, calling APIs, running code or changing the state of another system. How those tools are defined, connected and constrained affects both what the agent can accomplish and what happens when a step goes wrong.

For each tool, decide which users are allowed to invoke it, which resources it exposes and whether the action needs approval. Least privilege is important here. Once an agent has permission to modify records, send messages or trigger transactions, a bad decision can affect systems outside the conversation. For example, if an agent only needs order status, it shouldn’t get general database access. If it only needs to draft an email; it doesn’t need permission to send one.

These limits also reduce the consequences of prompt injection. An agent that reads an untrusted webpage, document or message might encounter instructions designed to redirect its reasoning. What happens next depends heavily on the tools and credentials available to it. Read-only access limits what the injected instruction can accomplish.

Model Context Protocol (MCP) provides a standards-based way to expose tools and data to compatible agent applications. Teams can define the tool interface separately from the agent instead of building a custom integration for every client. MCP doesn’t decide what should be exposed, though. The team still has to define which tools the agent can discover and which user or service identity authorizes each invocation.

For exploratory tasks such as ad hoc calculations, chart generation or temporary data manipulation, a contained code environment provides another option. Rather than creating a separate tool for every possible calculation, teams can give the agent a sandbox with tightly controlled access to the files, data and network resources required for the task.

Step 4: Add memory deliberately

Agent memory refers to three different things:

  • Working memory holds information needed during the current task, typically through the model’s context window.
  • Episodic memory keeps useful information about previous sessions or events.
  • Semantic memory retains longer-lived knowledge about the environment, such as conventions, relationships or lessons learned from earlier work.

Each type of memory requires different retention strategies, with clear separation between active tasks, past experiences, and long-term facts.

Arun Agarwal, Snowflake’s Principal Product Marketing Manager, AI/ML, outlines how to structure retention without diluting context quality: “Keep the current task in working memory, past attempts in episodic memory, and confirmed, reusable knowledge in semantic memory. It’s not a good idea to load the full history into every request. An AI system has to retrieve what’s relevant and be selective about what becomes a lasting lesson. For example, a correction about one customer in your CRM software shouldn’t become a rule for every customer.”

Quote Icon

Keep the current task in working memory, past attempts in episodic memory, and confirmed, reusable knowledge in semantic memory. It’s not a good idea to load the full history into every request — an AI system has to retrieve what’s relevant and be selective about what becomes a lasting lesson.

Arun Agarwal
Principal Product Marketing Manager, AI/ML, Snowflake

Memory architecture affects efficiency and cost. Loading an agent’s complete interaction history into every request degrades reasoning speed, inflates token overhead and introduces noise. In contrast, if the agent remembers where a resource lives or which approach has failed in the past, it doesn’t have to rediscover that information every time it runs. In Snowflake AI Research’s ArcticMem experiments, agents that retrieved relevant knowledge saved from previous sessions before and during a metrics bug fix used 18% fewer tool calls on the task and achieved a higher pass rate than the baseline agent that wasn’t memory-enabled.

Anything that persists also needs access and retention rules. If memory contains user-specific data, proprietary context or earlier outputs, the system needs to know who’s allowed to retrieve it and how long it should remain available. Otherwise, information from one user or task could surface in another session.

Step 5: Choose an orchestration pattern

Start with one agent when one reasoning loop can handle the workflow. A single agent with focused instructions, good context and a clearly defined tool set is easier to trace when something goes wrong. Engineers can inspect the reasoning steps and tool calls without reconstructing several handoffs between different agents.

Add another agent when that separation produces a measurable improvement in quality, latency, cost or control. Multiple agents make sense when the workflow has distinct responsibilities. One agent might research a question while another independently reviews the evidence. An orchestrator-worker design can split a large task into smaller pieces, while an evaluator-optimizer pattern separates generation from review. A router can send different request types to different specialists.

Each additional agent, however, creates another handoff, another context boundary and another intermediate state. That means the team needs a clear reason for introducing it.

Whatever agent orchestration option you use, define the handoffs. Which agent owns each decision? What information moves between them? Which permissions travel with the task? Where will engineers find the full trace?

Learn more about agent orchestration patterns and how to choose >

Step 6: Evaluate the trajectory, not just the answer

An agent might make dozens of decisions before it produces a final answer, and the quality of the intermediate steps affects the quality of the result. It could omit a relevant source, misinterpret part of the request or draw a conclusion that isn’t fully supported by the evidence it retrieved — none of which is necessarily obvious from the response alone.

That’s why production evaluation needs to look at the trajectory as well as the outcome. By examining how the agent interpreted the request, which tools it chose, what evidence it retrieved and how it used that evidence, teams can distinguish a data problem from a planning problem, a tool-use problem or a response-generation problem.

Agarwal highlights the necessity of inspecting the full execution trace: “Look at what the agent actually did: which sources it retrieved, which tools it called, how it reasoned over intermediary results and how it responded to errors. The model’s explanation of its reasoning can help with debugging, but it isn’t proof that the steps were correct.”

Quote Icon

The model’s explanation of its reasoning can help with debugging, but it isn’t proof that the steps were correct.

Arun Agarwal
Principal Product Marketing Manager, AI/ML, Snowflake

In some cases, evaluation criteria need to reflect the domain. For example, a support agent needs to cite current information and policies. Subject matter experts should help define tests so the evaluation measures the behavior people will rely on in production.

Production usage will also expose cases the original evaluation set missed. New data becomes available, tools change and users ask questions the team didn’t anticipate. Representative cases from production logs can feed back into the evaluation suite and become regression tests for future changes.

Observability supplies the evidence underneath those evaluations: traces, tool calls, latency, token consumption, errors and retrieved context. Evaluation tells the team whether the run met the expected standard, while observability gives engineers something to inspect when it didn’t.

COMMON PITFALL

A common mistake is checking only the final response. An answer can look accurate while omitting a required source, using the wrong metric definition or relying on an unsupported inference. Evaluating the trajectory helps verify that the agent used the data, tools and evidence required for an accurate and complete result.

Step 7: Bound cost before you scale

The cost of an AI agent depends on how much work each request triggers. Reasoning steps, retrieval, tool calls, code execution and evaluation all consume model or compute resources, while implementation cost includes the engineering required to prepare data, context, integrations and controls.

The primary cost drivers include:

  • Reasoning and model inference: Longer plans, retries and reflection loops consume more tokens.
  • Retrieval and context: Search operations, large context windows and repeated schema exploration require additional processing.
  • Tool calls and compute: Database queries, APIs and code execution consume resources outside the model call.
  • Evaluation and observability: Production systems run tests, traces and monitoring alongside user requests.
  • Engineering and operations: Teams prepare data, define semantic context, integrate tools, maintain policies and investigate failures.

Request volume alone won’t tell you much about spend. One question might trigger several model calls, searches and database queries before the user sees an answer. An agent with weak grounding might do even more work as it explores unfamiliar schemas or retries an unsuccessful route.

Keeping spend under control means setting hard guardrails before a workflow ever executes. Agarwal explains the practical limits AI administrators need to establish: “It’s critical to set limits before users are able to start using a tool’s capabilities. Settings like how long an agent can run for; how much it can spend per request, user, or billing department; and how often it can retry in a timeframe are important cost guardrails. For enterprise data, it’s really critical to have the agent start with read-only access, require approval for high-impact changes and test the recovery paths on your data if things go awry.”

Quote Icon

It’s critical to set limits before users are able to start using a tool’s capabilities.

Arun Agarwal
Principal Product Marketing Manager, AI/ML, Snowflake

One effective approach is to set several budget thresholds rather than just one hard ceiling. The first might generate an alert. A higher threshold could restrict further usage, while an exceptional increase might require approval. Team-level and user-level limits also keep one workload from consuming shared resources unnoticed.

Snowflake’s 2026 research found that 60% of organizations reported data storage and compute costs pushing AI projects over budget, reinforcing the need to track consumption as agents move from pilots into broader use.

Step 8: Preserve authorization, identity and auditability

Tool permissions determine what an agent is allowed to do, but once the same agent starts serving many users, there’s another question to address: whose authority is the agent acting under?

An analytics agent, for example, might be available to both a finance executive and a regional sales manager. Those users shouldn’t automatically see the same rows or invoke the same actions simply because they’re using the same agent.

Carry user authorization through the workflow

The requesting user’s authorization should follow the task through database queries and tool calls so the underlying systems continue to enforce row-, object- and action-level controls. This keeps access decisions in the systems responsible for protecting the data rather than asking the agent to infer from its instructions what a particular user should be allowed to see.

Keep the user and agent identities visible

When an agent takes an action for someone, both identities (the agent’s and the user’s) are relevant. The system needs to know which user initiated the request and which agent performed the subsequent work, particularly when a workflow crosses multiple tools or services.

That distinction becomes increasingly important as agents perform work without a person approving every intermediate step. An administrator investigating an account change, for example, should be able to tell whether a user performed it directly or an agent performed it under that user’s authority.

Record what the agent did

For consequential actions, audit records should capture enough detail to reconstruct what happened: the initiating user, the agent, the tool invoked, the authorization context applied and the resulting operation.

This gives security, governance and engineering teams a concrete record to work from when investigating an unexpected result or reviewing how an agent is being used. Snowflake’s 2026 research found that roughly 57% of respondents acknowledged using unapproved AI tools, underscoring the value of keeping agent activity visible rather than allowing AI-assisted work to happen outside established controls.

What a production rollout actually looks like

Moving an agent into production doesn’t require solving every possible use case before launch. A staged rollout gives teams a way to answer different questions in sequence: Does the agent work well enough for a narrow group? Do users find enough value to return? Which additional data and capabilities does real usage show they need? And once adoption grows, what engineering practices are required to keep the system reliable?

Snowflake’s internal go-to-market AI assistant provides a useful example of a real-world production rollout. The project expanded from a narrow pilot to roughly 6,000 intended users across Snowflake’s sales and marketing organization, answering more than 330,000 questions by the end of 2025.

Use different gates for pilot, beta and GA

During the initial pilot, the Snowflake team focused on correctness, reliability and basic usefulness. Once those requirements were met, beta testing shifted toward feature completeness and retention: Did the assistant solve enough real problems that people continued using it week after week?

Beta users reported more than 92% Net Promoter Score and weekly active user retention above 70%, which gave the team evidence to proceed with broader deployment. At GA, the measurement question changed again, this time toward activation and adoption across the larger user population.

The broader lesson is that the evidence required to expand an agent changes with the stage of deployment. Quality can gate the move from pilot to beta, retention can show whether the product has enough utility to scale and adoption measures whether the broader population is actually incorporating it into work.

Expand scope using observed demand

Snowflake’s sales and marketing organization included more than 15 potential user personas, but the initial GA release concentrated on account executives, solution engineers and SDRs. Those groups represented roughly half of the intended user population and gave the team a sizable audience without requiring it to support every workflow at once.

The data scope followed a similar pattern. At launch, the assistant used six semantic views covering 48 tables and approximately 1,400 columns. Post-GA usage helped identify where additional coverage was warranted; the team later expanded to 10 semantic views, 64 tables and more than 1,750 columns, along with additional data sources and capabilities.

For teams planning their own rollout, that argues for defining enough context to support the initial workflow well, then using production questions and failed or unsupported requests to prioritize what gets modeled next.

Plan for the engineering work that starts after launch

Once thousands of people were using the assistant, the team’s work changed. Feature requests increased, reliability expectations rose and changes in the underlying AI platform sometimes required earlier design decisions to be revisited.

Snowflake responded by moving to sprint-based development with a defined process for intake, triage and prioritization. The team added analytics engineering and backend engineering capacity, invested in automated testing and CI/CD and reserved development time for refactoring, architectural changes and platform upgrades.

That post-launch work is easy to underestimate when planning an agent project. A production agent continues to encounter new questions, changing data, updated tools and new model or platform capabilities. Teams therefore need ownership and regression testing after GA, along with enough capacity to address defects without putting every new request onto the roadmap.

Measure whether broader use is producing value

By the end of 2025, Snowflake’s assistant was answering more than 35,000 questions per week for more than 2,500 weekly active users. Based on a conservative estimate of five minutes saved per question, Snowflake calculated annual productivity equivalent to more than 65 full-time employees and ROI above 5x.

Those figures are specific to Snowflake’s implementation, but the measurement approach is transferable: adoption alone doesn’t establish value. Pair usage with a measure tied to the original workflow — time saved, analyst workload reduced, resolution time improved or another baseline established during Step 1 — and revisit that calculation as usage expands.

How to build AI agents on Snowflake

For enterprise agents, one of the biggest architectural questions is how much data, context, governance and operational infrastructure you rebuild around every agent. Snowflake’s role is to keep those data-facing pieces connected to the same governed environment rather than recreating them application by application.

Cortex Agents handles multi-step orchestration across structured and unstructured data. Teams define the agent’s instructions and tools around a bounded workflow, then use traces and evaluations to see where the design needs adjustment.

For structured analytics, Cortex Analyst works with governed Semantic Views that define business entities, metrics and relationships. Cortex Search retrieves relevant information from unstructured enterprise content. An agent can use both during the same interaction when a question crosses structured and unstructured sources.

Snowflake Managed MCP Servers expose approved Snowflake resources through MCP, with Snowflake authentication and role-based access control governing tool discovery and invocation. Cortex Agents can also use custom tools and external MCP connections when the workflow needs to reach systems outside Snowflake.

For exploratory calculations and data processing, the Cortex Agent code execution tool (public preview) runs generated Python in an isolated sandbox. Data retrieval still goes through the agent’s approved data tools.

Cortex Agent Evaluations measures agent runs using metrics such as answer correctness and logical consistency, along with custom criteria defined for a particular workflow. Traces capture the intermediate steps engineers need when they’re trying to understand why a run received a poor score.

Snowflake governance policies continue to enforce access at the data layer, while resource budgets and per-user quotas provide controls over AI consumption as usage grows.

Organizations that want a ready-made interface can use Snowflake CoWork for conversational access across structured and unstructured enterprise data. Teams building their own application can work directly with Cortex Agents.

Whichever interface the organization chooses, the engineering sequence stays largely the same: start with a narrow task, establish the business context, expose only the tools the task requires, decide what information should persist and evaluate the full trajectory. Once the team understands how the agent behaves, it has evidence for deciding how widely to deploy it.

Build for the conditions the agent will actually face

A prototype answers an important first question: Can the agent complete the intended task? Production raises the standard. The agent has to keep producing accurate, complete results as the data changes, different users bring different permissions and requests move beyond the examples the development team anticipated.

This changes what teams need to engineer around the agent. Governed context helps it interpret the data correctly, carefully designed tools control how it interacts with other systems, and trajectory-level evaluation shows whether its answers are supported by the work that produced them. As usage grows, budgets, authorization and audit records provide the controls needed to understand what the agent is consuming, who it’s acting for and what it actually did.

The strongest production designs avoid recreating those foundations for every new agent. Business definitions, access policies, lineage and other shared context are more useful when agents can draw on the same governed sources rather than maintaining their own versions. From there, teams can expand scope based on evidence from real usage — adding data, tools and capabilities as the agent proves that it can handle them reliably.

KEY TAKEAWAY

A production-ready AI agent is a system, not simply a model connected to tools. Teams need to engineer the context it reasons over, the actions it can take, the state it retains and the way its work is evaluated, then add cost, authorization and audit controls before expanding its reach.

 

Forward-looking statements: This article contains forward-looking statements, including about our future product offerings, and are not commitments to deliver any product offerings. Actual results and offerings may differ and are subject to known and unknown risk and uncertainties. See our latest 10-Q for more information.

Frequently Asked Questions

Your common questions about building AI agents, answered by Snowflake experts.

No-code and low-code tools work for relatively narrow agents over prepared data and predefined actions. An agent that integrates with enterprise systems typically still requires engineering work around data preparation, permissions, tool definitions, evaluation and operations, even if a visual interface handles the initial configuration.

Total cost depends on model usage, reasoning steps, retrieval, tool calls and supporting compute, along with the engineering required to prepare context and operate the system. Runtime cost varies considerably by architecture: repeated retrieval, unnecessary exploration and long reasoning loops increase consumption, while stronger grounding and tighter tool routing reduce the amount of work each request triggers.

A chatbot primarily responds to an input with information. An AI agent works toward a goal across multiple steps, choosing tools, inspecting results and adjusting what it does next. If those tools include APIs, databases or other systems with write access, the agent can also change system state.

A prototype over prepared data might take days. Production timing depends on how much work is required to establish governed context, integrate tools, define permissions and build an evaluation suite, followed by enough testing to establish acceptable quality and cost thresholds.

Explore AI Resources

Explore AI Topics

Deep dives into every aspect of artificial intelligence