Blog/Core Platform/Introducing S3-compatible storage support for Iceberg tables in Snowflake Horizon Catalog
Oct 5, 2026/2 min readCore Platform

Introducing S3-compatible storage support for Iceberg tables in Snowflake Horizon Catalog

AI is only as good as the data it can reach. This is why teams are adopting Apache Iceberg™: to access data in place from their preferred engines and models.

But landing data in Iceberg isn’t enough. You also have to securely manage access to those tables from every engine. This is where Snowflake Horizon Catalog comes in. Built on Apache Polaris™, it leverages open security mechanisms based on the Iceberg REST Catalog API to grant query-specific secure access to engines like Spark, Trino, Flink, PyIceberg, DuckDB, and more for both read and write operations. This means engines can get the same level of access to these tables as Snowflake's own query engine does.

This setup has worked for Amazon S3, Azure Storage, and Google Cloud Storage out of the box. However, it did not work for any other storage that speaks the S3 wire protocol without being AWS. Until now, MinIO, Ceph, Wasabi, Backblaze B2, Cloudflare R2, IDrive e2, or a self-hosted cluster in your own data center could not participate in this architecture pattern. The problem isn’t about how Snowflake writes these tables. Rather, other engines couldn't read the file paths in the Iceberg metadata.

This ends today. With this private preview, Snowflake Horizon managed Iceberg tables on S3-compatible storage are now readable and writable from any compatible engine.

Why this matters

Plenty of data cannot move into public cloud storage, often for reasons that have nothing to do with technology: a data residency requirement, a sovereign cloud mandate, an existing on-premises investment, or an egress bill that makes a copy uneconomic. Until now, those teams had to choose between keeping the data where it belongs or getting multi-engine access to it.

Horizon Catalog can already govern those tables and query with Snowflake’s compute. Engines outside Snowflake couldn’t access them.

Now they can access these tables too.

 

Figure 1
Figure 1

Horizon as a translation layer

Horizon Catalog is designed to completely abstract storage location and type for engines. The result is a consistent experience for your teams regardless of where the data is stored. It works by first identifying the storage location by where it lives, not just how to talk to it. This works because commercial AWS, government regions, sovereign clouds, and S3-compatible systems each carry a distinct identity inside Snowflake. All your preferred engines see is a pointer to the table’s metadata file.

Powering this experience is Horizon Catalog's built-in storage intelligence layer. For every table, it resolves the full access contract an engine needs. Part of that contract is the table's location, in the form that the reader expects. Azure resolves to abfss://, Google Cloud Storage to gs://, and S3-compatible storage to standard s3://. Snowflake's internal storage model isn’t exposed in the open metadata, so standard Iceberg engines can work with these tables without Snowflake-specific support.

[The two storage models Horizon translates between]

Figure 2
Figure 2

 

Your storage stays the same. Your configuration stays the same. Your S3-compatible table now behaves like any other Iceberg table.

Horizon: The catalog for universal governance, including diverse storage types

Interoperability is only useful if it does not cost you your controls. Your existing access controls come along with these tables. The same roles and grants apply. The catalog remains the single place access is decided. Operations external engines perform through the Iceberg REST Catalog APIs land in Access History next to your Snowflake SQL activity, giving these tables, including from external engines, one audit trail.

Updating your existing tables

Tables predating this feature have their older metadata files written with the internal scheme. They need a one-time metadata update before external engines can access them. During this preview, our team runs this update for you on request. Tables created after S3-compatible storage support is enabled for your account need no additional work.

Two things to know about this update. First, Horizon Catalog makes it easy to see which tables need it: until a table is updated, the catalog returns a clear error instead of metadata the engine can't resolve, so you get a clear error up front rather than a scan that fails partway through. Second, the update rebuilds only the Iceberg metadata layer, which means rewriting a few metadata files without moving any data. Snowflake Time Travel is unaffected.

How it works

Point Horizon Table at an S3-compatible volume:

 

CREATE OR REPLACE EXTERNAL VOLUME minio_vol
  STORAGE_LOCATIONS = (
    (
      NAME = 'minio-loc'
      STORAGE_PROVIDER = 'S3COMPAT'
      STORAGE_BASE_URL = 's3compat://analytics-bucket/iceberg/'
      STORAGE_ENDPOINT = 'minio.internal.example.com'
      CREDENTIALS = (AWS_KEY_ID = '...' AWS_SECRET_KEY = '...')
    )
  );

CREATE ICEBERG TABLE orders (o_orderkey NUMBER, o_totalprice NUMBER)
  CATALOG = 'SNOWFLAKE'
  EXTERNAL_VOLUME = 'minio_vol'
  BASE_LOCATION = 'orders/';

 

The engine connects with your endpoint and keys. Here with PyIceberg:

 

from pyiceberg.catalog import load_catalog

catalog = load_catalog("horizon", **{
    "type": "rest",
    "uri": "https://<account>.snowflakecomputing.com/polaris/api/catalog",
    "warehouse": "<database>",
    "token": "<token>",
    "s3.endpoint": "https://minio.internal.example.com",
    "s3.access-key-id": "<key>",
    "s3.secret-access-key": "<secret>",
})

print(catalog.load_table("analytics.orders").scan().to_arrow().num_rows)

Spark, Trino and others take the same three properties under their own configuration prefixes.

What's coming next

Vended credentials: Today, the engine manages its own bucket access because it interacts with S3-compatible volumes using static keys and secrets that we cannot scope down. Even though some providers issue short-lived credentials, they force us to use different APIs. To solve this, we are actively building a way to vend scoped credentials for each storage platform.

Get Started

S3-compatible storage support is in private preview. If you have Iceberg data on MinIO, Ceph, Cloudflare R2, or a cluster in your own data center, talk to your account team about switching it on.


Apache Iceberg™ is a trademark of the Apache Software Foundation in the United States and/or other countries.

Forward-looking statements

This article contains forward-looking statements, including about our future product offerings, and are not commitments to deliver any product offerings. Actual results and offerings may differ and are subject to known and unknown risk and uncertainties. See our latest 10-Q for more information.

Learn more about the author

Sushant Raikar

Senior Software Engineer
Share this post

Subscribe to our blog newsletter

Get the best, coolest and latest delivered to your inbox each week

Where Data Does More