Expedition. Free, virtual, Nov 3–6.

Technical tracks for practitioners, outcomes for leaders.

Snowflake for Developers/Guides/Scientific Workbench for Life Sciences R&D
Quickstart

Scientific Workbench for Life Sciences R&D

Cortex LLM
Deven Atnoor

Overview

The Scientific Workbench is an enterprise AI platform for life sciences research and development, built entirely on Snowflake. It brings together Cortex Agents, NVIDIA BioNeMo NIMs, Snowpark Python tools, and a Snowflake App Runtime (SAR) web application into a single governed environment where computational biologists, medicinal chemists, structural biologists, and clinical data scientists can run AI-driven discovery workflows.

By the end of this guide you will have a fully deployed Scientific Workbench with 28 agent-callable tools, a multi-agent Discovery Agent, 4 Cortex Analyst semantic views, reference datasets across 5 scientific domains, and a React web application — all running in your Snowflake account.

Architecture Diagram

Prerequisites

  • Snowflake account (Enterprise edition or higher recommended)
  • ACCOUNTADMIN access (for initial setup; can be reduced after deployment)
  • Snowflake CLI (snow) installed
  • NVIDIA API key from build.nvidia.com (free tier available)
  • Node.js 20+ and npm (for local development only)

What You'll Learn

  • How to deploy a multi-agent AI platform on Snowflake
  • How to configure NVIDIA BioNeMo NIM tools for drug discovery
  • How to set up role-based access control for scientific teams
  • How to build and deploy a Snowflake App Runtime (SAR) web application
  • How to run end-to-end drug discovery workflows through a conversational agent
  • How to extend the platform with new tools and data sources

What You'll Need

ResourceDetails
Storage (reference data)~40 GB
Storage (platform tables)< 1 GB
Warehouses3 (XS, S, ML) with auto-suspend
Cortex Search2 services (catalog + tools)
Marketplace subscriptionsPubMed CKE, ClinicalTrials.gov CKE
GPU computeNone required (NVIDIA hosted API)
Monthly credits (estimate)500 - 1,000

What You'll Build

  • 28 agent-callable tools: 7 native Snowpark Python, 11 NVIDIA BioNeMo NIM wrappers, 3 search tools, 2 MONAI imaging tools, 3 Parabricks genomics tools, 2 multi-NIM pipelines
  • 7 Cortex Agents: Genomics, Chemistry, Structural, Clinical, Workflow, Orchestrator, Discovery
  • 4 Cortex Analyst semantic views: Clinical, Genomics, Compounds, Chemistry
  • Reference data across 5 domains: Genomics (HGNC), ChEMBL (bioactivity), Clinical (ClinicalTrials.gov), Pathways (Reactome, MSigDB, GO), Protein (UniProt)
  • A React web application deployed on Snowflake App Runtime

Environment Setup

Clone the Repository

git clone https://github.com/Snowflake-Labs/sf-hcls-solutions.git
cd sf-hcls-solutions/solutions/scientific-workbench

Install Snowflake CLI

If you haven't already, install the Snowflake CLI:

pip install snowflake-cli

Configure a Connection

Add a named connection for deployment. You need ACCOUNTADMIN role for the initial setup:

snow connection add \
  --connection-name workbench-deploy \
  --account <your-account> \
  --user <your-user> \
  --role ACCOUNTADMIN \
  --warehouse COMPUTE_WH

Test the connection:

snow connection test --connection workbench-deploy

Obtain NVIDIA API Key

  1. Go to build.nvidia.com and create an account
  2. Navigate to any BioNeMo NIM (e.g., Boltz-2) and click "Get API Key"
  3. Copy the key — you'll provide it during deployment

Quick Start (One Command)

The fastest path is the automated deploy script:

bash deploy.sh --connection workbench-deploy --nvidia-key <your-nvidia-key>

This runs all 7 phases automatically. If you prefer step-by-step deployment, follow Steps 3 through 7 below. If you used the one-command deploy, skip to Step 8.

For non-interactive environments (CI, automation):

SWB_ASSUME_YES=1 bash deploy.sh --connection workbench-deploy --nvidia-key <your-nvidia-key>

Configuration File (Optional)

For repeatable deployments or custom naming, create a config file:

cp workbench.config.yaml.example workbench.config.yaml

The config file controls all deployment parameters. Every field is optional — deploy.sh auto-detects defaults if omitted. CLI flags always take precedence over config values.

snowflake:
  # Snowflake CLI connection name from ~/.snowflake/config.toml
  connection: default

secrets:
  # NVIDIA API key from build.nvidia.com (for hosted NIM API calls)
  nvidia_api_key: ""
  # NGC API key from ngc.nvidia.com (for SPCS NIM container weight downloads)
  ngc_api_key: ""

databases:
  workbench: SCIENTIFIC_WORKBENCH      # Platform database
  reference: WORKBENCH_REFERENCE       # Read-only reference data
  projects: WORKBENCH_PROJECTS         # Per-project schemas

warehouses:
  xs: WORKBENCH_XS                     # Lightweight queries, admin tasks
  small: WORKBENCH_S                   # Tool execution, data loading
  ml: WORKBENCH_ML                     # Heavy ML workloads, large data scans

# Enable specific NIM containers on SPCS GPU compute pools
nim_services:
  boltz2: false                        # L40S GPU required
  genmol: false                        # A10G GPU sufficient
  diffdock: false                      # L40S GPU required
  rfdiffusion: false                   # L40S GPU required
  proteinmpnn: false                   # A10G GPU sufficient
  molmim: false                        # A10G GPU sufficient
  openfold2: false                     # A10G GPU sufficient
  openfold3: false                     # L40S GPU required
  msa_search: false                    # A10G + 500 GiB block volume

agents:
  deploy_legacy_sql: true              # Deploy agents via SQL scripts
  deploy_agent_studio: true            # Deploy agents via Agent Studio YAML

marketplace:
  pubmed_cke: false                    # Set true if PubMed CKE is subscribed
  clinicaltrials_cke: false            # Set true if ClinicalTrials.gov CKE is subscribed

Deploy with the config file:

bash deploy.sh --config workbench.config.yaml

Validate the config before deploying:

bash scripts/validate-config.sh

deploy.sh Reference

The deploy script runs 7 phases in sequence. Use flags to skip or isolate phases:

FlagEffect
--connection NAMESnowflake CLI connection name (auto-detected if omitted)
--config PATHPath to workbench.config.yaml
--nvidia-key KEYNVIDIA API key (or NVIDIA_API_KEY env var)
--ngc-key KEYNGC API key for SPCS NIMs (or NGC_API_KEY env var)
--skip-prereqsSkip Phase 1 EAI/secret creation (if already done)
--skip-dataSkip Phase 2 reference data loading (~45 min saved)
--skip-appSkip Phase 7 app deployment
--skip-testsSkip post-deploy verification
--backend-onlyDeploy only NVIDIA procedures, agents, and app
--agents-app-onlyResume at agents + app (skips infrastructure and data)
--app-onlyDeploy only the SAR app (fastest redeploy)
--mirror-nimsMirror NIM container images from nvcr.io (requires Docker + --ngc-key)
--dry-runPrint what would be executed without running

Deployment phases:

PhaseNameScriptsDuration
1Infrastructure00-prerequisites.sql through 13-execute-tool-by-name.sql (11 scripts)2-5 min
1bNIM Mirroring (opt-in)Docker pull/tag/push of 9 NIM images60-120 min
2Reference Datadata/loaders/*.sql, data/synthetic/load_synthetic_data.sql30-45 min
3Toolstools/native/*.sql, tools/nvidia/*.sql, tools/search/*.sql, 14-custom-tool-onboarding.sql5-10 min
3.5NotebooksUpload notebooks/*.ipynb to workspace1-2 min
4Agentsengine/sql/agents/*.sql, Agent Studio YAML specs3-5 min
5Semantic Viewsdata/semantic_views/*.sql, workflow runner2-3 min
6App Deploymentsnow app deploy (Docker build + SPCS service)5-8 min
7Verificationtests/*.sql, CALL RUN_VALIDATION_TESTS()1-2 min

Total fresh deployment: ~50-80 minutes (without NIM mirroring). Subsequent deploys with --app-only take 5-8 minutes.

Teardown

To remove the entire deployment:

bash scripts/teardown.sh --connection workbench-deploy -y
FlagEffect
-ySkip confirmation prompt
--dry-runPrint what would be dropped without executing
--keep-dataDrop platform objects but preserve WORKBENCH_REFERENCE data

Deploy Infrastructure

This step creates the Snowflake objects that the workbench runs on: databases, schemas, warehouses, compute pools, RBAC roles, external access integrations, and secrets.

Phase 1: Run Infrastructure Scripts

If deploying manually (not using deploy.sh), execute these scripts in order using Snowsight or the Snowflake CLI:

OrderScriptWhat It Creates
1scripts/00-prerequisites.sqlExternal access integrations (NVIDIA API, PDB, NIM runtime), secrets, network rules
2scripts/01-databases-and-schemas.sqlSCIENTIFIC_WORKBENCH (6 schemas), WORKBENCH_REFERENCE (5 schemas), WORKBENCH_PROJECTS
3scripts/02-warehouses.sqlWORKBENCH_XS, WORKBENCH_S, WORKBENCH_ML warehouses; NIM_GPU_A10G_POOL, NIM_GPU_L40S_POOL compute pools
4scripts/03-rbac.sqlWORKBENCH_ADMIN, WORKBENCH_SCIENTIST, WORKBENCH_VIEWER roles with appropriate grants
5scripts/05-tool-registry.sqlCATALOG.TOOLS table, CATALOG.ASSETS table, CATALOG.CHAT_SESSIONS table, REGISTER_TOOL and SEED_ASSETS procedures

Database Layout

After Phase 1, you have three databases:

SCIENTIFIC_WORKBENCH          -- Platform database
  CATALOG                     -- Tool registry, assets, agents, chat sessions
  WORKFLOWS                   -- Workflow templates and run history
  GOVERNANCE                  -- Promotion log, canary assertions
  PROVENANCE                  -- Provenance audit trail
  PROJECTS                    -- Shared project space

WORKBENCH_REFERENCE           -- Read-only reference data
  GENOMICS                    -- HGNC gene nomenclature
  CHEMBL                      -- Bioactivity measurements
  CLINICAL                    -- ClinicalTrials.gov
  PATHWAYS                    -- Reactome, MSigDB, GO
  PROTEIN                     -- UniProt entries

WORKBENCH_PROJECTS            -- Per-project schemas
  DEMO_NSCLC                  -- NSCLC patient cohort demo
  DEMO_COMPOUNDS              -- Compound screening demo
  DEMO_CLINICAL               -- Clinical trials demo
  DEMO_STRUCTURAL             -- Protein targets demo

RBAC Roles

RolePurposeTypical User
WORKBENCH_ADMINFull access: manage tools, approve submissions, configure agents, share dataPlatform administrators
WORKBENCH_SCIENTISTRun tools, execute workflows, view all data, share resultsComputational biologists, chemists
WORKBENCH_VIEWERRead-only access to results and catalogManagers, reviewers

Verify Infrastructure

-- Check databases exist
SHOW DATABASES LIKE 'SCIENTIFIC_WORKBENCH';
SHOW DATABASES LIKE 'WORKBENCH_REFERENCE';
SHOW DATABASES LIKE 'WORKBENCH_PROJECTS';

-- Check roles
SHOW ROLES LIKE 'WORKBENCH_%';

-- Check external access integrations
SHOW EXTERNAL ACCESS INTEGRATIONS LIKE 'NVIDIA_%';

Deploy NIMs on SPCS

This optional step deploys NVIDIA BioNeMo NIM containers locally on Snowpark Container Services (SPCS) GPU compute pools. By default, all NIM tools call the NVIDIA hosted API — SPCS deployment provides an alternative for customers who need air-gapped operation or want to avoid per-call API costs.

Prerequisites for SPCS NIMs

  • Docker installed locally (for image mirroring)
  • NGC API key from ngc.nvidia.com (separate from the build.nvidia.com key)
  • GPU compute pool quotas enabled in your account (A10G and/or L40S)

GPU Requirements

Each NIM has specific GPU and storage requirements. SPCS nodes have a 93.13 GiB storage cap across all instance families.

NIMContainer ImageCompressed SizeGPUMemoryStatus
Boltz-2nvcr.io/nim/mit/boltz2:1.8.09.93 GiBL40S 48 GiB32-90 GiBVerified
GenMolnvcr.io/nim/nvidia/genmol:latest9.44 GiBA10G 24 GiB8-24 GiBTemplate
DiffDocknvcr.io/nim/mit/diffdock:latest15.33 GiBL40S 48 GiB32-64 GiBTemplate
RFdiffusionnvcr.io/nim/nvidia/rfdiffusion:2.3.016.37 GiBL40S 48 GiB32-90 GiBTemplate
ProteinMPNNnvcr.io/nim/nvidia/proteinmpnn:latest8.60 GiBA10G 24 GiB8-24 GiBTemplate
MolMIMnvcr.io/nim/nvidia/molmim:latest13.04 GiBA10G 24 GiB16-48 GiBTemplate
OpenFold2nvcr.io/nim/nvidia/openfold2:latest8-10 GiBA10G 24 GiB16-48 GiBTemplate
OpenFold3nvcr.io/nim/nvidia/openfold3:latest10.19 GiBL40S 48 GiB32-90 GiBTemplate
MSA-Searchnvcr.io/nim/nvidia/msa-search:latest8.64 GiBA10G 24 GiB16-48 GiBSpecial (needs block volume)
Evo2nvcr.io/nim/nvidia/evo2:latest15.25 GiBH100/H200TBDBlocked (needs measurement)

Compute Pools

The infrastructure scripts (Step 3) create two GPU compute pools:

-- Already created by 02-warehouses.sql:
-- NIM_GPU_A10G_POOL: GPU_NV_S (A10G 24 GiB) — GenMol, ProteinMPNN, MolMIM, OpenFold2
-- NIM_GPU_L40S_POOL: GPU_L40S_G1_16 (L40S 48 GiB) — Boltz-2, DiffDock, RFdiffusion, OpenFold3

-- Verify pools exist
SHOW COMPUTE POOLS LIKE 'NIM_GPU_%';

Step 1: Mirror NIM Images

NIM containers must be mirrored from NVIDIA's NGC registry (nvcr.io) to your Snowflake image repository. This is required because SPCS pulls images from the Snowflake registry, not directly from external registries.

Automated (via deploy.sh):

bash deploy.sh --connection workbench-deploy \
  --nvidia-key <your-nvidia-key> \
  --ngc-key <your-ngc-key> \
  --mirror-nims

Manual (per image):

# Login to NGC registry
docker login nvcr.io -u '$oauthtoken' -p <your-ngc-key>

# Login to Snowflake image registry
snow spcs image-registry login --connection workbench-deploy

# Get your Snowflake registry URL
REPO_URL=$(snow sql --connection workbench-deploy \
  --query "SHOW IMAGE REPOSITORIES IN SCHEMA SCIENTIFIC_WORKBENCH.CATALOG" \
  --format json | python3 -c "
import json, sys
data = json.load(sys.stdin)
print([r['repository_url'] for r in data if 'nim_gpu_images' in r.get('repository_url','').lower()][0])
")

# Mirror Boltz-2 (example — repeat for each NIM)
docker pull nvcr.io/nim/mit/boltz2:1.8.0
docker tag nvcr.io/nim/mit/boltz2:1.8.0 $REPO_URL/boltz2:1.8.0
docker push $REPO_URL/boltz2:1.8.0

Important — Apple Silicon users: Docker on ARM Macs pulls the arm64 variant by default. SPCS requires amd64. Pull by digest or use --platform linux/amd64:

docker pull --platform linux/amd64 nvcr.io/nim/mit/boltz2:1.8.0

Mirroring all 9 images takes 60-120 minutes depending on bandwidth.

Step 2: Create SPCS Services

The service definitions are in scripts/10-nim-spcs-services.sql. Only Boltz-2 is uncommented (verified). To deploy it:

USE DATABASE SCIENTIFIC_WORKBENCH;
USE SCHEMA CATALOG;

CREATE SERVICE SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC
  IN COMPUTE POOL NIM_GPU_L40S_POOL
  FROM SPECIFICATION $$
spec:
  containers:
    - name: boltz2
      image: /SCIENTIFIC_WORKBENCH/CATALOG/NIM_GPU_IMAGES/boltz2:1.8.0
      env:
        NIM_HTTP_API_PORT: "8000"
        NIM_LOG_LEVEL: "INFO"
      secrets:
        - snowflakeSecret: SCIENTIFIC_WORKBENCH.CATALOG.NGC_API_KEY
          secretKeyRef: SECRET_STRING
          envVarName: NGC_API_KEY
      volumeMounts:
        - name: dshm
          mountPath: /dev/shm
      resources:
        requests:
          nvidia.com/gpu: 1
          memory: 32Gi
        limits:
          nvidia.com/gpu: 1
          memory: 90Gi
      readinessProbe:
        port: 8000
        path: /v1/health/ready
  endpoints:
    - name: boltz2
      port: 8000
      public: true
  volumes:
    - name: dshm
      source: memory
      size: 16Gi
  $$
  EXTERNAL_ACCESS_INTEGRATIONS = (NIM_RUNTIME_EAI)
  MIN_INSTANCES = 1
  MAX_INSTANCES = 1;

Key points about the service spec:

  • /dev/shm volume: All PyTorch-based NIM containers require a shared memory volume. SPCS uses source: memory for this.
  • NGC_API_KEY secret: Mounted as an environment variable for model weight downloads on first startup.
  • Readiness probe: /v1/health/ready — the service reports READY only after weights are loaded.
  • Startup time: ~7.5 minutes (3.5 min image pull + 4 min weight download). Add 11-15 minutes if the compute pool must provision a new node.

Step 3: Verify Service Health

-- Check service status
SELECT SYSTEM$GET_SERVICE_STATUS('SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC');

-- View container logs
SELECT SYSTEM$GET_SERVICE_LOGS('SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC', 0, 'boltz2', 50);

-- Check endpoint URL
SHOW ENDPOINTS IN SERVICE SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC;

Endpoint Path Reference

The hosted API and self-hosted containers use different URL structures. For the hosted API, the base is https://health.api.nvidia.com. For self-hosted containers (including SPCS), the base is http://localhost:8000 (or the SPCS service DNS).

NIMHosted API PathSelf-Hosted Container PathDocs
Boltz-2/v1/biology/mit/boltz2/predict/biology/mit/boltz2/predictbuild.nvidia.com/boltz2
GenMol/v1/biology/nvidia/genmol/generate/biology/nvidia/genmol/generatebuild.nvidia.com/genmol
DiffDock/v1/biology/mit/diffdock/biology/mit/diffdockbuild.nvidia.com/diffdock
RFdiffusion/v1/biology/ipd/rfdiffusion/generate/biology/ipd/rfdiffusion/generatebuild.nvidia.com/rfdiffusion
ProteinMPNN/v1/biology/ipd/proteinmpnn/predict/biology/ipd/proteinmpnn/predictbuild.nvidia.com/proteinmpnn
MolMIM/v1/biology/nvidia/molmim/generate/biology/nvidia/molmim/generatebuild.nvidia.com/molmim
OpenFold2/v1/biology/openfold/openfold2/predict-structure-from-msa-and-template/biology/openfold/openfold2/predict-structure-from-msa-and-templatebuild.nvidia.com/openfold2
OpenFold3/v1/biology/openfold/openfold3/predict/biology/openfold/openfold3/predictbuild.nvidia.com/openfold3
MSA-Search/v1/biology/colabfold/msa-search/predict/biology/colabfold/msa-search/predictbuild.nvidia.com/msa-search
Evo2/v1/biology/arc/evo2-40b/generate/biology/arc/evo2-40b/generatebuild.nvidia.com/evo2

The pattern: the self-hosted container path is the hosted path minus the /v1 prefix. The hosted API requires a Bearer token (Authorization: Bearer <NVIDIA_API_KEY>); self-hosted containers do not require authentication.

To verify a container's actual endpoint after deployment, query its OpenAPI spec:

curl http://<service-endpoint>:8000/openapi.json | python3 -m json.tool | grep '"paths"' -A 20

Cost Management

SPCS GPU services consume credits continuously while running. To manage costs:

-- Suspend a service (stops GPU billing)
ALTER SERVICE SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC SUSPEND;

-- Resume when needed
ALTER SERVICE SCIENTIFIC_WORKBENCH.CATALOG.NIM_BOLTZ2_SVC RESUME;

The compute pools have auto-suspend configured (300 seconds by default). However, auto-suspend only triggers after all services in the pool are suspended. Suspend services explicitly when not in use.

Special Cases

  • MSA-Search: Requires a 500 GiB+ block volume for reference sequence databases (~1.2 TB for full UniRef30/ColabFold). The service spec includes a block volume sourced from a stage.
  • Evo2: The 40B parameter model's runtime footprint (~89 GiB) is close to the 93.13 GiB node storage cap. Do not deploy until measured on your target instance family. May require H100 80 GiB or H200 141 GiB VRAM.

Load Reference Data

Phase 2 loads ~40 GB of reference data across 5 scientific domains. This is the longest phase (~30-45 minutes).

Phase 2: Reference Data

# If deploying manually, run these scripts:
# scripts/04-reference-data.sql (schema setup)
# data/loaders/load_chembl.sql
# data/loaders/load_remaining.sql
# data/synthetic/load_synthetic_data.sql

Data Domains

DomainSchemaKey TablesSourceRows (approx)
GenomicsGENOMICSHGNC_GENESHGNC43,000
BioactivityCHEMBLACTIVITIES, ASSAYS, MOLECULE_DICTIONARY, TARGET_DICTIONARYChEMBL2M+
ClinicalCLINICALSTUDIES, CONDITIONS, INTERVENTIONSClinicalTrials.gov500K+
PathwaysPATHWAYSPATHWAYS, PATHWAY_GENES, GENE_SETS, GENE_SET_MEMBERS, GO_ANNOTATIONS, GO_TERMSReactome, MSigDB, GO1M+
ProteinPROTEINUNIPROT_ENTRIESUniProt570K+

Verify Data Loading

-- Check row counts across reference schemas
SELECT TABLE_SCHEMA, TABLE_NAME, ROW_COUNT
FROM WORKBENCH_REFERENCE.INFORMATION_SCHEMA.TABLES
WHERE TABLE_SCHEMA IN ('GENOMICS', 'CHEMBL', 'CLINICAL', 'PATHWAYS', 'PROTEIN')
  AND TABLE_TYPE = 'BASE TABLE'
ORDER BY TABLE_SCHEMA, TABLE_NAME;

Register Tools

Phase 3 creates the 28 agent-callable tool procedures and registers them in the CATALOG.TOOLS table.

Tool Taxonomy

The workbench supports four types of tools:

TypeRuntimeExampleCount
Native (Snowpark Python)Warehousevalidate_molecule, run_differential_expression7
NIM Wrapper (NVIDIA API)Warehouse + NVIDIA hosted APIrun_boltz2, run_genmol, run_diffdock11
Search (Cortex Search / CKE)Cortex Search servicesearch_pubmed, search_clinical_trials3
Pipeline (multi-tool chain)Warehouserun_drug_discovery_pipeline, run_msa_structure_pipeline2
Imaging (MONAI NIM)Warehouse + NVIDIA hosted APIrun_monai_vista3d, run_monai_lung_nodule2
Genomics (Parabricks)SPCS GPU (pending)run_parabricks_fq2bam, run_parabricks_deepvariant3

How Tools Are Structured

Every tool is a Snowflake stored procedure or function with a JSON annotation in its COMMENT field:

CREATE OR REPLACE PROCEDURE CATALOG.MY_TOOL(...)
RETURNS VARCHAR
LANGUAGE PYTHON
...
COMMENT = 'TOOL:{"display_name":"My Tool","description":"...","domains":"genomics","params":{...},"return_type":"VARCHAR","example":"CALL CATALOG.MY_TOOL(...)"}'
AS $$
...
$$;

The REGISTER_TOOL procedure adds the tool to CATALOG.TOOLS:

CALL CATALOG.REGISTER_TOOL(
  'my_tool',                                    -- name (unique)
  'My Tool',                                    -- display_name
  'Description of what the tool does.',         -- description
  'genomics,drug-discovery',                    -- domains (comma-separated)
  'procedure',                                  -- tool_type
  'SCIENTIFIC_WORKBENCH.CATALOG.MY_TOOL',       -- function_reference
  '{"param1":"STRING","param2":"INT"}',         -- parameters (JSON)
  'VARCHAR',                                    -- return_type
  'CALL CATALOG.MY_TOOL(''value'', 42)'         -- example_usage
);

NVIDIA BioNeMo NIM Tools

These 11 tools call the NVIDIA hosted API via the NVIDIA_API_EAI external access integration:

ToolNIM ModelWhat It Does
run_boltz2Boltz-2Protein structure prediction + binding affinity (pIC50)
run_genmolGenMolDe novo drug-like molecule generation
run_diffdockDiffDockSmall-molecule docking (binding pose prediction)
run_proteinmpnnProteinMPNNInverse folding (backbone to sequence design)
run_rfdiffusionRFdiffusionDe novo protein backbone generation
run_molmimMolMIMLatent-space molecule optimization
run_openfold2OpenFold2Single-chain protein structure prediction
run_openfold3OpenFold3Multi-chain complex structure prediction
run_msa_searchMSA-SearchMultiple sequence alignment via ColabFold
run_evo2Evo2DNA/RNA sequence generation (40B parameter model)
run_drug_discovery_pipelineGenMol + DiffDockEnd-to-end: generate, validate, dock, rank

Verify Tool Registration

SELECT tool_type, COUNT(*) AS count
FROM SCIENTIFIC_WORKBENCH.CATALOG.TOOLS
WHERE status = 'active'
GROUP BY tool_type
ORDER BY count DESC;

-- Expected: 28 total active tools
SELECT COUNT(*) AS total_active_tools
FROM SCIENTIFIC_WORKBENCH.CATALOG.TOOLS
WHERE status = 'active';

Configure AI Services

Phase 4-6 sets up Cortex Search services, the governance framework, workflow engine, semantic views, and the agent layer.

Cortex Search and Governance

ScriptWhat It Creates
scripts/06-cortex-services.sqlCortex Search services for tool discovery and asset catalog
scripts/07-governance.sqlGOVERNANCE.PROMOTION_LOG, GOVERNANCE.CANARY_ASSERTIONS
scripts/08-workflow-engine.sqlWORKFLOWS.TEMPLATES, WORKFLOWS.RUNS, workflow execution procedures

Semantic Views (Cortex Analyst)

Four semantic views enable natural-language queries over structured data:

Semantic ViewUnderlying DataExample Query
SV_CLINICALDEMO_NSCLC patient cohort"How many patients are in stage IIIA?"
SV_GENOMICSGene expression + pathways"Which genes are upregulated in the NSCLC cohort?"
SV_COMPOUNDSCompound IC50 assay results"Show me compounds with IC50 below 100 nM against EGFR"
SV_CHEMISTRYChEMBL molecule reference"List all approved kinase inhibitors"

Agent Architecture

The workbench uses a multi-agent architecture with specialized domain agents coordinated by an orchestrator:

                    DISCOVERY_AGENT
                    (user-facing)
                         |
               uses semantic views +
               search + code execution
                         |
                  ORCHESTRATOR_AGENT
                  (routes to domains)
                   /     |     \
        GENOMICS   CHEMISTRY  STRUCTURAL
         AGENT      AGENT      AGENT
           |          |          |
        5 tools    4 tools    9 tools
                         |
                   CLINICAL_AGENT
                   (4 semantic views)

Each domain agent has access only to its relevant tools. The ORCHESTRATOR_AGENT routes queries to the correct domain and synthesizes cross-domain results. The DISCOVERY_AGENT is the user-facing agent that also has direct access to semantic views and code execution.

Verify Agent Configuration

-- Test the Discovery Agent with a simple query
SELECT SNOWFLAKE.CORTEX.DATA_AGENT_RUN(
  'SCIENTIFIC_WORKBENCH.CATALOG.DISCOVERY_AGENT',
  '{"query": "What tools are available for protein structure prediction?"}',
  TRUE
) AS response;

Deploy the Application

Phase 7 deploys the React web application to Snowflake App Runtime (SAR).

SAR Deployment

From the app/ directory:

cd solutions/scientific-workbench/app
snow app deploy --connection workbench-deploy

This performs three steps:

  1. Uploads source files to a Snowflake workspace
  2. Builds a Docker container in the cloud (via SPCS build service)
  3. Creates or upgrades the APPLICATION SERVICE

Build time is approximately 3-5 minutes. The output includes the app URL:

App ready at https://<hash>-<account>.snowflakecomputing.app

Grant Access to Users

After deployment, grant access to the appropriate roles:

-- Scientists can use the app
GRANT USAGE ON APPLICATION SERVICE SCIENTIFIC_WORKBENCH_APP
  TO ROLE WORKBENCH_SCIENTIST;

-- Viewers get read-only access
GRANT USAGE ON APPLICATION SERVICE SCIENTIFIC_WORKBENCH_APP
  TO ROLE WORKBENCH_VIEWER;

App Modules

The application includes 10 modules:

ModulePathDescription
Home/Dashboard with tool count, agent status, recent activity
Chat/chatConversational interface to the Discovery Agent
Explore/exploreSchema browser with column profiling and data preview
Experiments/experimentsWorkflow run history and results viewer
Asset Catalog/catalogSearchable catalog of all platform assets with data preview
Tools/toolsRegistry of all 28 agent-callable tools
Tool Onboard/tools/onboardCustom tool submission wizard
Governance/governancePromotion log, canary assertions, provenance trail
Notebooks/notebooksSnowflake notebook launcher
Share/shareCross-account data sharing (admin/scientist)

Verify Deployment

Open the app URL in your browser. You should see:

  • Home page with "28 Active Tools" (or similar count)
  • Chat page where you can talk to the Discovery Agent
  • Tools page listing all registered tools

Local Development

For development and testing, you can run the app locally against your Snowflake account.

Setup

cd solutions/scientific-workbench/app
npm install

Environment Variables

Create a .env.local file (or rely on ~/.snowflake/config.toml default connection):

# Option 1: Key-pair auth (recommended)
SNOWFLAKE_ACCOUNT=<account>
SNOWFLAKE_USER=<user>
SNOWFLAKE_PRIVATE_KEY_PATH=~/.snowflake/rsa_key.p8
SNOWFLAKE_WAREHOUSE=WORKBENCH_XS
SNOWFLAKE_ROLE=WORKBENCH_ADMIN

# Option 2: config.toml (zero config — uses default connection)
# No env vars needed if ~/.snowflake/config.toml is configured

# Enable admin features in local dev
SWB_LOCAL_DEV_ADMIN=true

Run the Dev Server

npm run dev

The app starts at http://localhost:3000. Hot reload is enabled — changes to pages and API routes take effect immediately.

Key Differences from SPCS

AspectLocal DevSPCS (Deployed)
AuthPassword or config.tomlSPCS service token
Caller's rightsNot available (use SWB_LOCAL_DEV_ADMIN)Reads sf-context-current-user-token header
URLlocalhost:3000<hash>-<account>.snowflakecomputing.app
SecretsEnvironment variablesgetSecret() API from app.yml

Testing Changes

Before deploying, verify your changes build successfully:

npx tsc --noEmit       # Type checking
npx next build         # Full production build

Then deploy with snow app deploy --connection workbench-deploy.

Run a Discovery Workflow

Now that the workbench is deployed, walk through an end-to-end drug discovery workflow.

Open the Chat

Navigate to the Chat page in the web application. You'll see the Discovery Agent interface.

Example: Multi-Step Drug Discovery

Try this prompt:

Find compounds in ChEMBL that target EGFR with IC50 below 100 nM, validate their drug-likeness, and predict binding poses for the top 3 candidates against PDB structure 1M17.

The agent will:

  1. Query the SV_COMPOUNDS semantic view to find EGFR compounds with IC50 < 100 nM
  2. Call validate_molecule on each candidate (RDKit sanitization + PAINS/BRENK filters)
  3. Call run_diffdock to dock the top 3 validated molecules against PDB 1M17
  4. Return structured results with confidence scores and binding poses

Example: Genomics Analysis

Run differential expression analysis comparing Stage I vs Stage III patients in the NSCLC cohort, then perform pathway enrichment on the upregulated genes.

The agent will:

  1. Call run_differential_expression on the NSCLC gene expression data
  2. Filter for significantly upregulated genes (log2FC > 1, FDR < 0.05)
  3. Call run_pathway_enrichment against MSigDB Hallmark gene sets
  4. Return enriched pathways with Fisher exact test p-values

Example: Protein Structure Prediction

Predict the structure of the first 200 residues of human TP53 protein and assess confidence.

The agent will:

  1. Look up the TP53 sequence (or ask you to provide it)
  2. Call run_openfold2 for single-chain structure prediction
  3. Return the predicted structure with pLDDT confidence scores

View Results

After each workflow, results are saved to output tables. View them in:

  • Experiments page — workflow run history with parameters and status
  • Explore page — browse the output table schema and preview rows
  • Asset Catalog — newly created result tables appear as assets

Extend the Platform

Adding a New Native Tool

Create a Snowpark Python procedure in SCIENTIFIC_WORKBENCH.CATALOG:

CREATE OR REPLACE PROCEDURE CATALOG.MY_NEW_TOOL(
    "P_INPUT" VARCHAR,
    "P_OUTPUT_TABLE" VARCHAR
)
RETURNS VARCHAR
LANGUAGE PYTHON
RUNTIME_VERSION = '3.11'
PACKAGES = ('snowflake-snowpark-python')
HANDLER = 'run'
COMMENT = 'TOOL:{"display_name":"My New Tool","description":"What it does","domains":"genomics","params":{"input":"STRING","output_table":"STRING"},"return_type":"VARCHAR","example":"CALL CATALOG.MY_NEW_TOOL(input, output)"}'
AS $$
import json

def run(session, p_input: str, p_output_table: str) -> str:
    # Your tool logic here
    result = {"status": "success", "output_table": p_output_table}
    return json.dumps(result)
$$;

Register it:

CALL CATALOG.REGISTER_TOOL(
    'my_new_tool', 'My New Tool',
    'What it does',
    'genomics', 'procedure',
    'SCIENTIFIC_WORKBENCH.CATALOG.MY_NEW_TOOL',
    '{"input":"STRING","output_table":"STRING"}',
    'VARCHAR',
    'CALL CATALOG.MY_NEW_TOOL(''test'', ''RESULTS.OUTPUT'')'
);

Refresh the asset catalog:

CALL CATALOG.SEED_ASSETS();

Adding a New NIM Wrapper

Follow the same pattern as existing NIM tools. Key requirements:

  1. Add EXTERNAL_ACCESS_INTEGRATIONS = (NVIDIA_API_EAI) to the procedure
  2. Add SECRETS = ('nvidia_key' = SCIENTIFIC_WORKBENCH.CATALOG.NVIDIA_API_SECRET)
  3. Use the _snowflake.get_generic_secret_string('nvidia_key') API to read the key
  4. Call the NIM endpoint via requests.post() with Bearer token auth

Custom Tool Onboarding

Scientists can submit custom tools through the web application:

  1. Navigate to Tools > Onboard in the app
  2. Fill in the tool specification (name, description, parameters, code)
  3. Submit for review
  4. An admin reviews and approves/rejects via Tools > Submissions
  5. Approved tools are automatically registered in the catalog

Adding Reference Data

To add a new reference dataset:

  1. Create a schema in WORKBENCH_REFERENCE (or use an existing one)
  2. Load the data
  3. Register it as an asset:
INSERT INTO SCIENTIFIC_WORKBENCH.CATALOG.ASSETS
    (asset_id, asset_name, asset_type, description, domain, schema_name, owner)
VALUES
    ('asset-data-WORKBENCH_REFERENCE.MY_SCHEMA.MY_TABLE',
     'MY_TABLE', 'dataset', 'Description of the dataset',
     'MY_SCHEMA', 'MY_SCHEMA', CURRENT_USER());

Or run CALL CATALOG.SEED_ASSETS() to auto-discover tables in registered schemas.

Security Model

Platform Roles and Hierarchy

The workbench uses three Snowflake database roles in a hierarchical inheritance chain:

SYSADMIN
  └── WORKBENCH_ADMIN       (full platform administration)
        └── WORKBENCH_SCIENTIST   (run tools, execute workflows, share results)
              └── WORKBENCH_VIEWER      (read-only access to results and catalog)

Each higher role inherits all privileges of the roles below it. WORKBENCH_ADMIN inherits SCIENTIST, which inherits VIEWER.

RoleCan DoCannot Do
WORKBENCH_ADMINEverything: manage tools, approve custom tool submissions, configure agents, share data cross-account, manage SPCS services, run governance proceduresN/A (full access)
WORKBENCH_SCIENTISTRun all tools, execute workflows, view all data, create result tables, share results, write to provenance log, submit custom toolsApprove/reject tool submissions, manage compute pools, run governance procedures
WORKBENCH_VIEWERRead catalog, view results, view tool registry, use XS warehouseRun tools, execute workflows, create tables, share data

Access Control by Schema

Database.SchemaADMINSCIENTISTVIEWER
SCIENTIFIC_WORKBENCH.CATALOGFull DDLSELECT + CALL proceduresSELECT only
SCIENTIFIC_WORKBENCH.RESULTSFull DDLSELECT + INSERT + CREATE TABLESELECT only
SCIENTIFIC_WORKBENCH.WORKFLOWSFull DDLSELECT + CALL proceduresNo access
SCIENTIFIC_WORKBENCH.GOVERNANCEFull DDL + proceduresSELECT onlyNo access
SCIENTIFIC_WORKBENCH.PROVENANCEFull DDLSELECT + INSERTNo access
WORKBENCH_REFERENCE.*Full DDLSELECT (read-only)No access
WORKBENCH_PROJECTS.*Full DDLSELECT (read-only)No access

Scientific Personas

In a typical deployment, the platform roles map to scientific personas. Assign users to the appropriate role based on their function:

PersonaSnowflake RoleWhat They Do
Computational BiologistWORKBENCH_SCIENTISTDifferential expression analysis, pathway enrichment, survival analysis, gene symbol validation
Medicinal ChemistWORKBENCH_SCIENTISTMolecule generation (GenMol), optimization (MolMIM), validation (RDKit + PAINS), molecular descriptors
Structural BiologistWORKBENCH_SCIENTISTProtein structure prediction (Boltz-2, OpenFold2/3), backbone design (RFdiffusion), sequence design (ProteinMPNN), molecular docking (DiffDock)
Clinical Data ScientistWORKBENCH_SCIENTISTCohort queries via semantic views, clinical trial search, patient outcome analysis
AI Drug Discovery ScientistWORKBENCH_SCIENTISTEnd-to-end pipelines, multi-tool workflows, agent-driven discovery
Platform AdministratorWORKBENCH_ADMINTool onboarding, agent configuration, data sharing, governance, compute management
Manager / ReviewerWORKBENCH_VIEWERReview experiment results, browse catalog, audit provenance

Assigning Roles to Users

-- Grant a scientist role to a user
GRANT ROLE WORKBENCH_SCIENTIST TO USER jane_doe;

-- Grant admin role
GRANT ROLE WORKBENCH_ADMIN TO USER platform_admin;

-- Grant viewer role
GRANT ROLE WORKBENCH_VIEWER TO USER manager_smith;

Users switch to their workbench role in a Snowflake session:

USE ROLE WORKBENCH_SCIENTIST;

In the web application, caller's rights automatically detects the user's active role via the SPCS user token — no manual USE ROLE needed.

Named Query Allowlist

The web application does not execute arbitrary SQL from the client. All queries go through a named query registry (lib/named-queries.ts) that maps query keys to pre-defined SQL templates:

  • Static queries (no parameters): SQL is a compile-time constant — zero attack surface
  • Parameterized queries: Parameters are validated with Zod schemas and used as bind variables (?) — no string interpolation
Client: { query: "column_info", params: { schema: "GENOMICS", table: "HGNC_GENES" } }
                                    |
                              resolveNamedQuery()
                                    |
Server: SELECT COLUMN_NAME, DATA_TYPE FROM INFORMATION_SCHEMA.COLUMNS
        WHERE TABLE_SCHEMA = ? AND TABLE_NAME = ?
        binds: ['GENOMICS', 'HGNC_GENES']

Authorization Patterns

API routes use caller's rights to check the user's Snowflake role:

// Check caller's role via SPCS user token
const [row] = await querySnowflake("SELECT CURRENT_ROLE() AS role", { callersRights: true })
const callerRole = String(row?.ROLE ?? "").toUpperCase()

if (!ALLOWED_ROLES.has(callerRole)) {
  return Response.json({ error: "Forbidden" }, { status: 403 })
}
EndpointRequired Role
GET /api/tools/submissionsAny authenticated user
POST /api/tools/submissions (approve/reject)WORKBENCH_ADMIN
POST /api/share (grant/create)WORKBENCH_ADMIN or WORKBENCH_SCIENTIST
GET /api/chat/sessionsScoped to CURRENT_USER()

External Access Integrations

EAIPurposeAllowed Hosts
NVIDIA_API_EAINVIDIA BioNeMo NIM API callsintegrate.api.nvidia.com, api.nvcf.nvidia.com
PDB_API_EAIRCSB Protein Data Bank lookupsdata.rcsb.org
NIM_RUNTIME_EAISPCS NIM container weight downloads*.nvidia.com, *.nvcr.io

Conclusion and Resources

Congratulations! You've deployed a complete Scientific Workbench for life sciences R&D on Snowflake. The platform provides:

  • A multi-agent AI system with 28 tools spanning genomics, chemistry, structural biology, and clinical data
  • NVIDIA BioNeMo NIM integration for state-of-the-art molecular and protein AI
  • A governed, role-based environment with audit trails and provenance tracking
  • A modern web application for interactive discovery workflows

What You Learned

  • Deploying a multi-database, multi-schema Snowflake platform with automated scripts
  • Configuring NVIDIA BioNeMo NIMs as agent-callable tools
  • Setting up Cortex Agents with domain-specific tool routing
  • Building and deploying a Snowflake App Runtime (Next.js) application
  • Extending the platform with new tools, data, and semantic views
  • Implementing security patterns: named query allowlists, RBAC, caller's rights

Related Resources

Updated Oct 9, 2026

This content is provided as is, and is not maintained on an ongoing basis. It may be out of date with current Snowflake instances