Blog/Catch Snowflake Engineering at Community Over Code EU 2026 in Glasgow!
Oct 8, 2026/7 min read

Catch Snowflake Engineering at Community Over Code EU 2026 in Glasgow!

Community Over Code is the Apache Software Foundation’s annual conference, and this year it's in Glasgow from Oct. 11 through 14. If you work in or around the Apache data ecosystem, there's a lot happening this year that's worth showing up for.

On our end, Snowflake engineers are showing up throughout the week to connect with the community and share the work we're doing in and around Apache Iceberg™, Apache Polaris™, Apache Spark™ and Apache Kafka® with talks, keynotes and hands-on opportunities.

Here’s the full rundown!

Lakehouse Day

Lakehouse Day is a one-day event taking place on Oct. 10, the day before the main conference. It’s focused on lakehouse architectures and the Apache projects that underpin them. You can register for Lakehouse Day when you register for the main conference.

Herding Cat(alogue)s: Understanding the Apache Iceberg™ Catalog Landscape Elizabeth Christensen, Clyde, 11:45 a.m.

The Iceberg catalog landscape has proliferated fast — Hive Metastore, JDBC, Nessie, Polaris and more. In her session, Elizabeth surveys the implementations, compares their metadata storage approaches and table discovery mechanisms, covers how the REST catalog spec reshaped the ecosystem, and lays out the tradeoffs for choosing between them.

An Extremely Technical Overview of how the Apache Iceberg™ Planning Implementation Actually Works Russell Spitzer, Tay, 1:45 p.m.

Russell walks through the actual Iceberg codebase: how query predicates get transformed and applied to metadata, how manifest files are read, how file tasks get split and dispatched to execution engines, and which configuration properties control each stage. If you're an engine developer or want to understand why your scans perform the way they do, this is the one.

Primary key index for Apache Iceberg: Unlocking Fast Point Lookups and Efficient Upserts Huaxin Gao, Tay, 3:30 p.m.

Iceberg has no efficient way to answer: given a key, where is the row? That means full-table scans for point lookups and expensive equality deletes for change data capture (CDC) workloads. Huaxin presents a primary key index that maps keys to physical locations (file path + row position), enabling scan-time file pruning and converting key-based deletes into position deletes. She covers the design as an Iceberg-native table, async maintenance that doesn't impact write paths, and a working Spark-based prototype.

Community Over Code main conference

The main conference runs four days and covers the full breadth of the Apache ecosystem. Here's where you'll find us, from a keynote on Snowflake's open source journey to talks on Kafka, Spark, and LLMs in the maintainer workflow, plus a hands-on Spark hackathon.

Sessions

Topics as a Service: Implementing Kafka for the Snowflake Compute Stack Tyler Jones, Wee Dram, Oct. 11, 4:10 p.m.

Tyler walks through what happened when Snowflake built a Kafka-compatible system on separated compute and storage: which parts of the protocol turned out to be genuinely portable, which had undocumented assumptions baked in, and where things broke in unexpected ways.

Lessons from Building a High-Performance Kafka Connect Sink for Cloud-Native Storage Artem Minyaylov and Seb Kurella, Clyde, Oct. 12, 2:20 p.m.

Snowflake's streaming ingest went through a major architecture shift from Snowpipe Streaming V1 to V2. This talk covers how V1 maps to Kafka's topic/partition model, how exactly-once semantics work at scale, what production experience revealed about downstream query performance, and how those lessons drove V2. The V2 section digs into data validation at ingest time, schema evolution across heterogeneous topics, and the tradeoffs involved.

From Consumers to Contributors: Snowflake's Open Source Journey at the ASF Danica Fine and Russell Spitzer, Keynote Hall, Oct. 13, 8:40 a.m.

A candid look at Snowflake's evolution in open source — from building on top of open source projects and open standards to actively contributing upstream, donating projects to the ASF, and working alongside the communities that maintain them. We'll cover what it took to change the internal culture, what the community taught us along the way, and where that journey is headed next.

Making PySpark faster and removing JVM type safety bottlenecks with transpilation Holden Karau, Spey, Oct. 14, noon

Holden covers a newly approved Spark Project Improvement Proposal (SPIP) for transpilation — an experimental approach to removing JVM type safety bottlenecks from PySpark execution. She walks through how it works, what you can do with it today, and where the limitations are.

How I Gave In and Started Using LLMs in My OSS Flow Russell Spitzer, Clyde, Oct. 14, 5 p.m.

Russell went from AI skeptic to running a maintainer workflow built on agents orchestrating agents. He walks through the reluctant first steps, what he built for code review and development, and how he keeps human judgment in the loop when the automation gets deep.

Spark Hackathon

Huaxin Gao and Holden Karau are helping organize a Spark hackathon at the main conference. It’s a hands-on session aimed at getting new contributors into the Spark codebase. If you've been curious about contributing to Spark but haven't found the right entry point, this is it. Join us on Monday, Oct. 12, from 11:20 a.m. to 3 p.m.

Spark has always been a project where you can make a meaningful contribution, but the size of the codebase can make it hard to know where to start. The goal of this hackathon is to lower that barrier and give new contributors a practical way to get involved, so that people leave with more confidence, a better understanding of the community, and a clear next step for continuing to contribute to Spark after the conference.

—Huaxin Gao, Apache Spark Committer

Open source at Snowflake in 2026

A few milestones from this year that we're bringing to Glasgow with us:

  • Apache Polaris™ graduated to a Top-Level Project (TLP): Polaris started as an internal Snowflake project, was open-sourced in mid-2024, and entered the ASF Incubator. As a TLP, it now has fully independent governance where the project's direction is shaped by a diverse community of contributors, not a single company's roadmap.
  • Apache Ossie (incubating) entered the Apache Incubator: Ossie is an open specification for semantic layer interoperability, a vendor-neutral, YAML-based format for defining business metrics, dimensions, and their relationships so that BI platforms, query engines, and AI agents can consume the same semantic definitions without loss of meaning. The project already has contributions from over 50 organizations including Snowflake, Salesforce, Databricks, dbt Labs, and others. Incubating under the ASF ensures it stays vendor-neutral as it matures.
  • Excitingly, upstream contributions to projects like Iceberg, Polaris, and Spark are now a regular part of how Snowflake engineers work, not a side project.

Snowflake went from an OSS Skeptic to a central contributor to multiple Apache projects in just a few years. We now have contributors on many critical Apache Lakehouse technologies and we’ve donated a few to the foundation as well! The change internally has also been dramatic, and now you won’t find a product meeting without someone asking “How can we make it work with Apache Iceberg?” or “How can we make this interoperable?” It takes a long time to reorient company strategy, but Snowflake saw that users wanted a more open and interoperable future and went full speed ahead!

—Russell Spitzer, Apache Iceberg PMC Member

Come find us in Glasgow

Throughout the conference, we’ll have a booth in the expo hall. Stop by and talk to us about what you're building, what's not working, or what you'd like to see from these projects.

The Apache community is where a lot of our engineering work happens, and Glasgow is where we get to do it face-to-face.

See you there!

Share this post

Subscribe to our blog newsletter

Get the best, coolest and latest delivered to your inbox each week

Where Data Does More