
DataHub Simplified Observability and Cut Incident Duration by 80% with Observe by Snowflake
By consolidating four observability systems, DataHub cut over $100,000 in annual AWS costs, eliminated an outage blind spot and reduced median incident duration by 80%.
80%Reduction in median incident duration
$100K+Annual AWS savings from unifying its observability infrastructure


Industry
TechnologyLocation
Palo Alto, CaliforniaGiving organizations the context to trust their data
When a company is trying to answer a critical business question, the answer is only as reliable as the data behind it. For organizations building AI and analytics into everyday decisions, knowing where data came from, what it means and whether it can be trusted can determine whether those systems deliver useful results or simply produce more information.
This is the problem DataHub helps organizations solve. Its open source metadata platform gives companies the context they need to understand and trust their data, so the people and AI systems relying on that data can make better-informed decisions. Its foundation is the DataHub open source metadata project, maintained by more than 15,000 contributors and running across thousands of organizations worldwide.
That same idea was becoming increasingly important inside DataHub.
Story highlights
Faster incident response: Centralized telemetry gives engineers one place to search, correlate signals and investigate issues without switching between multiple observability tools.
A foundation for AI-powered operations: DataHub connected Observe to an internal AI agent so engineers can ask operational questions in Slack and get answers grounded in live telemetry.
Monitoring that survives outages: DataHub’s observability no longer runs on the infrastructure it monitors, so engineers keep visibility into DataHub Cloud during exactly the failures that once took their monitoring down with it.
Needing better context at home
As DataHub Cloud grew, so did the infrastructure behind it. The company operated a larger and more distributed production environment, with 22 EKS and GKE clusters and more than 70 isolated customer namespaces in its largest production cluster. Every day, roughly 500 GB of telemetry flows through that environment.
DataHub had plenty of data about what was happening in its infrastructure. The problem was finding the right context when something went wrong.
Its observability environment had grown into four separate systems. Logs lived in OpenSearch, while metrics lived in Prometheus and Thanos. Grafana handled dashboards, and Prometheus AlertManager handled alerting. During an incident, engineers had to move between these systems and connect the dots themselves.
The setup also made observability harder to scale across the company. Product teams depended on the Platform team to build dashboards and alerts, while the Platform team spent time maintaining the infrastructure behind the monitoring stack.
Then there was a more fundamental problem: The monitoring infrastructure lived on the same infrastructure it monitored. If the underlying environment failed, the systems DataHub needed to understand the failure could become unavailable too.
“That is a blind spot at the exact moment you can least afford one,” says Craig Rueda, member of the technical staff at DataHub. “It defeats the purpose of having a monitoring system at all.”
The cost was becoming difficult to justify as well. DataHub’s two OpenSearch clusters alone cost more than $100,000 a year in AWS infrastructure, before accounting for the engineering time required to operate them.
With roughly 90 of DataHub’s 110 employees using observability, this was no longer simply a tooling problem for the platform team. DataHub needed to rethink how its people accessed operational information.
Finding one platform for critical signals
DataHub spent seven weeks evaluating observability platforms and running hands-on trials. The team wanted more than a replacement for OpenSearch. It needed a platform that could handle its telemetry volume while giving engineers more control over how they used that data.
DataHub chose Observe by Snowflake and brought its telemetry into one platform. The team deployed the Observe agent across its Kubernetes fleet, created separate U.S. and EU tenants for data residency, and managed its configuration through Terraform.
For engineers, the biggest change was that the information needed to investigate an issue was now searchable in one place. Instead of jumping between systems, teams could investigate logs alongside metrics and traces. Product engineering teams could also create their own dashboards and alerts without waiting for the platform team.
That changed the role of the platform team too. Rather than spending its time maintaining separate observability systems or building every dashboard, the team could give other engineers the tools to answer more questions themselves.
Moving observability away from the infrastructure it monitored also eliminated the outage blind spot that had been one of DataHub’s biggest concerns. The result was more than just a simpler monitoring environment. It gave DataHub a different foundation for working with operational data.
We replaced four self-managed observability systems with one platform, retired over $100,000 a year of AWS infrastructure and the engineering toil that came with it, and cut median incident duration by about 80%. More importantly, our monitoring no longer runs on the infrastructure it monitors. We are no longer blind at the exact moment an outage starts.”
Chandrasekharan Iyer
Keeping visibility available when it matters most
The results showed up in both the cost of running observability and the time it took to respond when something went wrong. DataHub retired its two OpenSearch clusters along with Thanos, Grafana and Prometheus AlertManager. That eliminated over $100K in annual AWS infrastructure costs associated with its previous OpenSearch environment while removing the engineering work required to maintain those systems.
Incident response improved at the same time. Median incident duration fell from approximately 85 hours to approximately 16 hours in less than a year, an improvement of roughly 80%. DataHub attributes the change to several factors, including process improvements and its internal AI tooling. Centralized, searchable telemetry has nevertheless made it easier for engineers to find the information they need when an incident occurs.
“Engineers find things faster when the telemetry is in one place and searchable,” Rueda says. DataHub has also used that visibility to understand how its own telemetry is being consumed. The team reduced metrics volume by 16% in the U.S. and 19% in the EU by eliminating unused series.
The impact reaches beyond engineering. When a customer reports an issue, support and engineering teams can use Observe to investigate the underlying telemetry and provide evidence-based answers.
In one case, DataHub traced intermittent API failures to a customer’s database connection pool reaching its limit during a burst of concurrent traffic. The team found the root cause within an hour and provided a concrete fix.
The same centralized telemetry is now becoming useful in another way. Because Observe provides a CLI and a query language, OPAL, DataHub connected the platform to an internal AI agent. Engineers can ask operational questions in Slack, and the agent can generate and execute OPAL queries to find the relevant telemetry.
The agent can also combine that information with data from incident.io and Argo CD. DataHub built an internal reference corpus covering OPAL and the Observe Terraform provider so the agent had the context it needed to work with the platform.
Today, the agent uses Observe data to answer operational questions and generate a daily report covering capacity and error trends. “Observe is one tool in the agent’s quiver, but it is the one that holds the telemetry, so it is the one I reach for first,” Rueda says.
For DataHub, that makes Observe more than just a place to look at telemetry. It has become a source of operational context that engineers can use directly and that the company’s AI workflows can access programmatically.

19%
Reduction in metrics volume in the EU by eliminating unused series
Building from diagnosis to remediation
DataHub is now building on that foundation to move its AI workflows beyond assisted diagnosis.
The team plans to use AI for automated triage, well-understood failure modes and automated remediation. The goal is to take the same operational context that helps an engineer understand an incident and make it available to increasingly automated workflows.
“We replaced four self-managed observability systems with one platform, retired over $100,000 a year of AWS infrastructure and the engineering toil that came with it, and cut median incident duration by about 80%,” says Chandrasekharan Iyer, VP of Engineering at DataHub. “More importantly, our monitoring no longer runs on the infrastructure it monitors. We are no longer blind at the exact moment an outage starts.”
For DataHub, the broader value of Observe by Snowflake is that it goes beyond helping engineers diagnose problems. It also gives the company’s AI workflows the operational context to start resolving them.
Engineers find things faster when the telemetry is in one place and searchable.”

