Expedition. Free, virtual, Nov 3–6.

Technical tracks for practitioners, outcomes for leaders.

Snowflake for Developers/Developer Blog/Decisions Are All You Need: Run Jev-Class Decision Models on Snowflake

Decisions Are All You Need: Run Jev-Class Decision Models on Snowflake

Cost and Performance Optimization
Josh Reini

A lot of what we ask large language models to do in data pipelines isn't open-ended. Is this support ticket about billing or a bug? Does this refund follow the policy text? Did the agent's answer stay on topic? Each one is a judgment with a handful of possible answers, and we hand it to a general model that can write essays and code, then parse one word out of its reply.

There is a huge advantage to using LLMs compared to classical ML models: they work with zero-shot: you write the labels and the instructions and it starts answering; no training or retraining needed. But this advantage came at a high cost that decision models aim to reduce.

TypeSafe named its decision model, Jev, after William Stanley Jevons, the economist who observed that more efficient steam engines led to more coal use, not less. The idea is simple: if decisions costs a small fraction of what it costs today but can still be performed without training, we can apply these decisions to a lot of work that is currently left not done.

Decision models are the new frontier, designed to accomplish this class of tasks. And there is an abundance of this new class of models, both proprietary and with open weights. This post gives you a practical look on how to run an open-weight decision model in Snowflake that’s designed to help you produce fast, cheap decisions where your data lives without going to a new, potentially untrusted, API.

decider-2b performance on JevBench with Snowflake ML

To test this I chose a small high-performing jev-class model, decider-2b, and a popular benchmark, JevBench. decider-2b is a 1.9B-parameter fine-tune of Qwen3.5-2B-Base that performed reasonably well on the benchmark at a very compelling price and small size.

In this set of tests, decider-2b gave a valid answer on all 231 of JevBench's public decisions and match the model's published results to within one hard item at a similarly low cost; a fraction of the cost of calling a general-purpose LLM API for the same single-pass classification. I also tested two paths for serving: batch inference and a real-time inference service. Batch inference and real-time inference generally cost about the same per decision in GPU time; the difference is when you pay. A batch job starts its own GPU node for each run and releases it afterward, which suits applying decisions to an entire dataset at once. The real-time service keeps a node running between requests and answered each one in about a tenth of a second, which suits interactive queries and small, frequent requests.

On JevBench's 231 public decisionsBatch inference jobReal-time inference (REST)
Correct answers175 of 231 (75.8%)175 of 231 (75.8%)
Easy / standard / hard48 of 48 / 64 of 72 / 63 of 11148 of 48 / 64 of 72 / 63 of 111
Total time (after startup)18.1 s17.7 s with 8 concurrent clients; 37.4 s with one
Latency per decision (median)n/a103 ms
Cost per 1,000 decisions$0.025$0.024 with 8 clients; $0.051 with one

Results as of October 2, 2026. This material contains hypothetical scenarios created for informational purposes only. Scenarios may include general recommendations, but they are not intended to predict or guarantee any specific results or cost savings. The potential accuracy of these scenarios is contingent upon numerous factors, including information provided by you, and Snowflake makes no guarantee that you will achieve any stated results.

Scaling decision models

To score more decisions per second, add GPUs. With run_batch(replicas=4) on a compute pool of four GPU_NV_S nodes started before the job, a 50,000-decision batch job scored 50.2 decisions per second, 3.6 times the rate of one A10G. That's near-linear scaling: 91% of the 55.2 decisions per second that four times one GPU would give. It finished in 22.5 minutes for $1.71, or $0.034 per 1,000 decisions including start-up.

decider-2b batch inference throughput: 13.8 decisions per second on one A10G, 50.2 on four, against 55.2 for linear scaling

Results as of October 2, 2026.

Real-time inference scales the same way. With create_service(min_instances=4, max_instances=4), four GPU_NV_S instances behind one REST endpoint answered in about 120 ms per decision (median) with up to four concurrent clients. At 16 clients the service reached 40 decisions per second with a median latency of 356 ms. At 64 clients it leveled off at 47.5 decisions per second, about 4 million decisions a day, with no errors. At full load it cost $0.027 per 1,000 decisions.

decider-2b real-time inference on four A10G instances: decisions per second rise from 5 with one client to 48 with 64, while median latency rises from 115 ms to 1.1 s

Results as of October 2, 2026.

Running decider-2b from the Model Registry

The Model Registry stores the weights and Python wrapper as one model, DECIDER_2B. Snowpark Container Services runs it on a GPU compute pool. Query history, metering views and client-side timers provide the timings and credits, and JevBench's harness scores the stored answers against ground truth.

The Model Registry's standard Hugging Face pipeline tasks don't return decider-2b's distribution over option letters, so I registered the weights with a CustomModel wrapper. Its system_one method takes JevBench's STATE_JSON and QUESTIONS_JSON and returns ANSWER_JSON, letting the harness score stored Snowflake answers. The wrapper verifies flash-linear-attention's Triton kernels, runs with torch.compile off as JevBench did, and warms CUDA graphs before scoring.

The companion repo has the full calls, the dependency list and the compute pool SQL.

MethodWhat it didKey settings
Registry.log_modelRegistered the weights and wrapper as DECIDER_2Btarget_platforms=["SNOWPARK_CONTAINER_SERVICES"], Python 3.12, cuda_version="12.8", relax_version=False
ModelVersion.run_batchScored the 231 items as a batch inference jobA GPU pool of one GPU_NV_S node (1x NVIDIA A10G, 24 GB) that suspends after 300 idle seconds; ResourcesSpec(gpu_requests="1"); InferenceSpec(num_workers=1, max_batch_rows=32); OutputSpec writing Parquet to a stage
ModelVersion.create_serviceRan the same model as a real-time inference service with a public REST endpointThe same GPU pool; gpu_requests="1", num_workers=1, max_batch_rows=32 (the minimum allowed); one instance; ingress_enabled=True; image built on a separate CPU_X64_S pool; requests authenticated with a programmatic access token

Deploy in App Runtime to Get Work Done Fast

Because decider-2b answers in about a tenth of a second, it's fast enough to sit inside an application and make decisions as work arrives.

To show this, I built a support ticket triage board on Snowflake App Runtime. For each ticket, the app asks which team should handle it and whether to escalate it. It sends the same question to decider-2b and to claude-sonnet-5 through AI_COMPLETE, side by side.

The same support tickets triaged by decider-2b and claude-sonnet-5 side by side

In this run, decider-2b made 6.5 decisions per second for $0.049 per 1,000 tickets, against 0.5 per second and $4.98 per 1,000 for the LLM. Decisions that fast and cheap can go anywhere an application needs to choose what happens next: routing, flagging, escalating.

What I learned running decider-2b on Snowflake

I wanted to see whether decider-2b's reported performance would hold up on Snowflake infrastructure, and it did: 175 of 231 correct, within one hard item of the model card's per-tier results. That one item came out as an exact tie between two options (0.4724 each), where small numerical differences between GPUs or library versions decide which label wins.

More importantly, I was able to efficiently run a jev-class model to produce these decisions at an incredibly low cost compared to performing the same work with LLMs. And I did it on Snowflake infrastructure, next to my data, without a new API or model provider. Scaling it was simple, too: four A10G nodes scored 50,000 decisions in 22.5 minutes for $1.71 in batch, and for real-time inference 64 clients made 47.5 decisions per second, about 4 million decisions a day, with no errors. At full load, real-time inference of decider-2b cost $0.027 per 1,000 decisions. Putting it to work was just as quick: a triage app on Snowflake App Runtime deployed with one command and made 6.5 decisions per second for $0.049 per 1,000 tickets.

To try it on your own decisions, the companion repo has the wrapper, the registration, batch, scaling and REST scripts, the triage board app, and the scoring code.

Updated Oct 5, 2026

This content is provided as is, and is not maintained on an ongoing basis. It may be out of date with current Snowflake instances