Skip to main content

Elastic Delivers Best-in-Class Metrics With Native Prometheus Support and Agentic Investigation Experiences

Native PromQL, out-of-the-box Kubernetes agentic investigations, and automated migration from Datadog and Grafana — all in the platform SREs already run for logs

Elastic announced new capabilities that bring the same scale, performance, and operational simplicity that have made Elastic a trusted platform for logs to metrics. 

With native Prometheus and PromQL support, out-of-the-box Kubernetes investigation workflows, and automated migration from Datadog and Grafana, Elastic now delivers a unified platform for metrics and logs. Built on Elasticsearch's columnar metrics engine, the platform can query metrics up to 30x faster than Prometheus and store data up to 2.5x more efficiently, without cardinality limits or custom metric penalties.

The metrics landscape has changed dramatically. Kubernetes and microservices have already pushed observability systems from thousands to millions of time series. Now AI workloads are accelerating that growth, making metrics not only a scale challenge but also a strategic cost and reliability problem. Most platforms make that growth expensive: premium vendors increase costs as cardinality grows, while lower-cost alternatives fragment metrics and logs across separate backends and query languages. The result is that teams often reduce data collection to control costs, leaving engineers with less context when incidents occur.

Elastic Observability addresses both problems in a single platform that stores OpenTelemetry, Prometheus-native, and application-defined metrics at full resolution alongside logs and traces, with no separate backends and no retention trade-offs. The release spans the metrics engine and the capabilities built on it:

  • Native PromQL and Prometheus Remote Write: PromQL queries run natively in Kibana and Prometheus metrics arrive via Remote Write, so existing dashboards, alert rules, and scrape configs work without modification.
  • Out-of-the-box Kubernetes workflows and content: SREs now go straight from an alert to the root cause through out-of-the-box agentic workflows, alert templates, ML anomaly detection jobs, and pre-built dashboards that activate at ingest for Kubernetes. SRE teams do not need to configure infrastructure from scratch before they get value.
  • Agentic investigations: When an alert fires, Elastic correlates metrics, logs, and traces that already share a single backend, using workflows with ML anomaly detection to surface what changed and how severe the deviation is before anyone is paged. The Observability MCP App and agent skills bring the same investigation capabilities to Claude, Cursor, VS Code, and any MCP-compatible tool.
  • Automated migration from Datadog and Grafana: The Observability Migration Platform converts dashboards, alert rules, and PromQL queries into Kibana equivalents automatically, so teams move what they've already built rather than rebuilding it.

"Elastic was already the platform many SREs trusted for logs at scale. Now we're bringing that same impressive scale, performance, and operational simplicity to metrics, delivering up to 30x faster metric queries than Prometheus, native Prometheus compatibility, and a more predictable cost model for high-cardinality metrics," said Baha Azarmi, GM, Observability at Elastic. "With a single backend for every signal, a single query language, and investigations that start before anyone is paged, SREs get complete context at the moment they need it most — without the bills that have forced teams to compromise on the data they keep."

“As we’ve moved more applications into Kubernetes and expanded our cloud footprint, data is growing rapidly and our need for granular, high-cardinality metrics is increasing," said Jeff Beagley, manager of DevOps, SRE, and Cloud Engineering, Bass Pro Shops. “Elastic’s new metrics capabilities let us handle that volume and surface the insights we need. Coupled with Elastic’s OpenTelemetry support, we get visibility into an increasingly complex architecture — all while keeping performance up and costs down.”

“At Eurowings, the improved metrics performance, native Prometheus support, logsdb and incident-handling workflows in Elasticsearch have helped our teams achieve faster incident response times and a more unified view across signals without jumping between systems,” said Iosif Tournas, Cyber Security & Elastic Platform Lead, Eurowings Aviation. “These new metrics capabilities complement the millions of log events per minute and APM traces we’re already handling in Elastic Observability. This unified view reduces operational friction, breaks the silos between teams and the time it takes to detect and respond to issues.”

The columnar metrics engine (TSDS), ES|QL time series support, PromQL in Kibana, and Prometheus Remote Write ingest are generally available. Out-of-the-box Kubernetes infrastructure content including dashboards, alert templates, SLO and ML anomaly detection jobs are also generally available. The Observability MCP App, Agent Skills, and the Observability Migration Platform are available in tech preview. All capabilities run across Elastic Cloud, serverless, and self-managed deployments.

While Datadog does not offer an on-premises option and Grafana limits its highest-value features to hosted deployments, Elastic gives organizations the flexibility to run observability workloads where their data and operational requirements demand.

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

Elastic Delivers Best-in-Class Metrics With Native Prometheus Support and Agentic Investigation Experiences

Native PromQL, out-of-the-box Kubernetes agentic investigations, and automated migration from Datadog and Grafana — all in the platform SREs already run for logs

Elastic announced new capabilities that bring the same scale, performance, and operational simplicity that have made Elastic a trusted platform for logs to metrics. 

With native Prometheus and PromQL support, out-of-the-box Kubernetes investigation workflows, and automated migration from Datadog and Grafana, Elastic now delivers a unified platform for metrics and logs. Built on Elasticsearch's columnar metrics engine, the platform can query metrics up to 30x faster than Prometheus and store data up to 2.5x more efficiently, without cardinality limits or custom metric penalties.

The metrics landscape has changed dramatically. Kubernetes and microservices have already pushed observability systems from thousands to millions of time series. Now AI workloads are accelerating that growth, making metrics not only a scale challenge but also a strategic cost and reliability problem. Most platforms make that growth expensive: premium vendors increase costs as cardinality grows, while lower-cost alternatives fragment metrics and logs across separate backends and query languages. The result is that teams often reduce data collection to control costs, leaving engineers with less context when incidents occur.

Elastic Observability addresses both problems in a single platform that stores OpenTelemetry, Prometheus-native, and application-defined metrics at full resolution alongside logs and traces, with no separate backends and no retention trade-offs. The release spans the metrics engine and the capabilities built on it:

  • Native PromQL and Prometheus Remote Write: PromQL queries run natively in Kibana and Prometheus metrics arrive via Remote Write, so existing dashboards, alert rules, and scrape configs work without modification.
  • Out-of-the-box Kubernetes workflows and content: SREs now go straight from an alert to the root cause through out-of-the-box agentic workflows, alert templates, ML anomaly detection jobs, and pre-built dashboards that activate at ingest for Kubernetes. SRE teams do not need to configure infrastructure from scratch before they get value.
  • Agentic investigations: When an alert fires, Elastic correlates metrics, logs, and traces that already share a single backend, using workflows with ML anomaly detection to surface what changed and how severe the deviation is before anyone is paged. The Observability MCP App and agent skills bring the same investigation capabilities to Claude, Cursor, VS Code, and any MCP-compatible tool.
  • Automated migration from Datadog and Grafana: The Observability Migration Platform converts dashboards, alert rules, and PromQL queries into Kibana equivalents automatically, so teams move what they've already built rather than rebuilding it.

"Elastic was already the platform many SREs trusted for logs at scale. Now we're bringing that same impressive scale, performance, and operational simplicity to metrics, delivering up to 30x faster metric queries than Prometheus, native Prometheus compatibility, and a more predictable cost model for high-cardinality metrics," said Baha Azarmi, GM, Observability at Elastic. "With a single backend for every signal, a single query language, and investigations that start before anyone is paged, SREs get complete context at the moment they need it most — without the bills that have forced teams to compromise on the data they keep."

“As we’ve moved more applications into Kubernetes and expanded our cloud footprint, data is growing rapidly and our need for granular, high-cardinality metrics is increasing," said Jeff Beagley, manager of DevOps, SRE, and Cloud Engineering, Bass Pro Shops. “Elastic’s new metrics capabilities let us handle that volume and surface the insights we need. Coupled with Elastic’s OpenTelemetry support, we get visibility into an increasingly complex architecture — all while keeping performance up and costs down.”

“At Eurowings, the improved metrics performance, native Prometheus support, logsdb and incident-handling workflows in Elasticsearch have helped our teams achieve faster incident response times and a more unified view across signals without jumping between systems,” said Iosif Tournas, Cyber Security & Elastic Platform Lead, Eurowings Aviation. “These new metrics capabilities complement the millions of log events per minute and APM traces we’re already handling in Elastic Observability. This unified view reduces operational friction, breaks the silos between teams and the time it takes to detect and respond to issues.”

The columnar metrics engine (TSDS), ES|QL time series support, PromQL in Kibana, and Prometheus Remote Write ingest are generally available. Out-of-the-box Kubernetes infrastructure content including dashboards, alert templates, SLO and ML anomaly detection jobs are also generally available. The Observability MCP App, Agent Skills, and the Observability Migration Platform are available in tech preview. All capabilities run across Elastic Cloud, serverless, and self-managed deployments.

While Datadog does not offer an on-premises option and Grafana limits its highest-value features to hosted deployments, Elastic gives organizations the flexibility to run observability workloads where their data and operational requirements demand.

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...