Skip to main content

Dynatrace Intelligence Redefines Observability with Trusted Agentic Automation

First-of-its-kind system fusing deterministic and agentic AI to power safe, autonomous operations

At Perform, its flagship annual user conference, Dynatrace unveiled Dynatrace Intelligence, a new agentic operations system that fuses deterministic and agentic AI. 

This differentiated combination delivers reliable, agentic AI-powered observability to customers. Built to observe and optimize dynamic AI workloads, Dynatrace Intelligence empowers organizations to build more resilient applications, elevate customer experiences, and drive autonomous action across modern digital ecosystems.

Dynatrace Intelligence represents the next phase in the evolution of the Dynatrace platform, helping the world’s largest enterprises move from reactive to preventive and advances them toward autonomous operations while ensuring teams remain firmly in control.

Why Dynatrace Intelligence Matters

Organizations are confronting rising complexity as they continue to adopt new technologies. For example, global AI investment is expected to reach nearly $2 trillion in 2026 and organizations are under increasing pressure to show meaningful progress. Yet many struggle with the unpredictable and dynamic nature of AI and agentic systems. Teams must quickly identify unexpected behaviors, understand downstream impact, and deploy fixes before customer experience or business performance suffers.

Dynatrace Intelligence addresses these challenges with deep, real-time visibility into system behavior and performance across cloud and AI-native environments, creating a real-time digital twin. It removes guesswork by fusing the precise deterministic AI insights with reasoning from coordinated AI agents that drive self-healing systems. The result is reliable autonomous action, reduced operational burden, and more time for strategic decision-making.

The Dynatrace Difference: Fusing Deterministic with Agentic AI

Dynatrace Intelligence uniquely combines deterministic AI, grounded in real-time causal context, and agentic AI, capable of safe reasoning, decision-making, and action within defined guardrails.

This system is powered by the Dynatrace 3rd generation platform, including:

  • Grail, an industry leading, unified data lakehouse that stores metrics, logs, traces, events, user sessions, and business and security data with precise, contextual integrity.
  • Smartscape, the real-time dependency graph that continuously and automatically maps relationships and fuels trustworthy, causal insights.

By anchoring agents in environment-specific facts, Dynatrace Intelligence provides a safer, faster, more reliable foundation for autonomous operations.  When Dynatrace benchmarked an external SRE agent working together with its deterministic agents, problems were solved up to 12 times more often, three times faster, and at half the cost compared to tests that did not use deterministic agents.

An Ecosystem of AI Agents

Enterprises can orchestrate built‑in and partner agents, with bidirectional integrations across the broader ecosystem, including ServiceNow, AWS, Microsoft Azure, Google Cloud, Atlassian, GitHub, Red Hat, and more.

The agentic architecture includes:

  • Agents that deliver foundational capabilities for trusted operational context through causal reasoning, prediction, and real-time intelligence, and oversight.
  • Agents that expand teams by providing insight and guidance to targeted functional areas and personas.
  • Ecosystem agents that connect with partner platforms expanding the scope of autonomous action across complex environments.

Advancing Autonomous Operations

With Dynatrace Intelligence, organizations can achieve:

  • Self-healing systems in dynamic, AI-driven environments
  • Proactive prevention, remediation, and optimization
  • Reliable autonomous action, leveraging both built-in agents and collaboration with partner agents, with full visibility and control

Teams remain in command while the system continuously manages operational complexity in the background.

The Journey to Autonomous Operations

Dynatrace Intelligence supports customers on a phased journey toward autonomy. Organizations can start by using AI-driven insights and recommendations, progress to leverage automation for supervised operations with human oversight, and ultimately advance to fully autonomous operations with guardrails and controls. This approach allows customers to safely adopt auto-prevention, auto-remediation, and auto-optimization, while maintaining control and building trust at every step.

“Agentic AI offers enormous potential, but many businesses still struggle to ensure it operates reliably, securely, and with consistent performance in real‑world environments,” said Bernd Greifeneder, Chief Technology Officer and Founder at Dynatrace. “Dynatrace Intelligence fuses deterministic and agentic AI, removing the guesswork and delivering AI‑powered observability organizations can trust.”

“As our digital environment grows more complex, we’re looking to move beyond reactive operations and manual intervention,” said Alexander Bicalho, Senior Director of Engineering at Autodesk. “What Dynatrace is outlining with Dynatrace Intelligence aligns with where we want to go—using trusted data and insights to support more autonomous operations. An approach that connects insight to action, while keeping our teams in control, could significantly improve performance and reliability as we scale. It’s observability that doesn’t just detect problems—it understands them and acts on them reliably.”

“The evolution of observability platforms is moving from manual root cause analysis to preventive operations. Organizations are progressing beyond reactive monitoring toward autonomous operations models that combine deterministic AI with agentic AI systems, with AI agents operating at different autonomy levels to orchestrate workflows across integrated ecosystems that span cloud platforms, development tools, and IT service management systems,” said Stephen Elliot, Group Vice President at IDC.

The Latest

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers ...

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Dynatrace Intelligence Redefines Observability with Trusted Agentic Automation

First-of-its-kind system fusing deterministic and agentic AI to power safe, autonomous operations

At Perform, its flagship annual user conference, Dynatrace unveiled Dynatrace Intelligence, a new agentic operations system that fuses deterministic and agentic AI. 

This differentiated combination delivers reliable, agentic AI-powered observability to customers. Built to observe and optimize dynamic AI workloads, Dynatrace Intelligence empowers organizations to build more resilient applications, elevate customer experiences, and drive autonomous action across modern digital ecosystems.

Dynatrace Intelligence represents the next phase in the evolution of the Dynatrace platform, helping the world’s largest enterprises move from reactive to preventive and advances them toward autonomous operations while ensuring teams remain firmly in control.

Why Dynatrace Intelligence Matters

Organizations are confronting rising complexity as they continue to adopt new technologies. For example, global AI investment is expected to reach nearly $2 trillion in 2026 and organizations are under increasing pressure to show meaningful progress. Yet many struggle with the unpredictable and dynamic nature of AI and agentic systems. Teams must quickly identify unexpected behaviors, understand downstream impact, and deploy fixes before customer experience or business performance suffers.

Dynatrace Intelligence addresses these challenges with deep, real-time visibility into system behavior and performance across cloud and AI-native environments, creating a real-time digital twin. It removes guesswork by fusing the precise deterministic AI insights with reasoning from coordinated AI agents that drive self-healing systems. The result is reliable autonomous action, reduced operational burden, and more time for strategic decision-making.

The Dynatrace Difference: Fusing Deterministic with Agentic AI

Dynatrace Intelligence uniquely combines deterministic AI, grounded in real-time causal context, and agentic AI, capable of safe reasoning, decision-making, and action within defined guardrails.

This system is powered by the Dynatrace 3rd generation platform, including:

  • Grail, an industry leading, unified data lakehouse that stores metrics, logs, traces, events, user sessions, and business and security data with precise, contextual integrity.
  • Smartscape, the real-time dependency graph that continuously and automatically maps relationships and fuels trustworthy, causal insights.

By anchoring agents in environment-specific facts, Dynatrace Intelligence provides a safer, faster, more reliable foundation for autonomous operations.  When Dynatrace benchmarked an external SRE agent working together with its deterministic agents, problems were solved up to 12 times more often, three times faster, and at half the cost compared to tests that did not use deterministic agents.

An Ecosystem of AI Agents

Enterprises can orchestrate built‑in and partner agents, with bidirectional integrations across the broader ecosystem, including ServiceNow, AWS, Microsoft Azure, Google Cloud, Atlassian, GitHub, Red Hat, and more.

The agentic architecture includes:

  • Agents that deliver foundational capabilities for trusted operational context through causal reasoning, prediction, and real-time intelligence, and oversight.
  • Agents that expand teams by providing insight and guidance to targeted functional areas and personas.
  • Ecosystem agents that connect with partner platforms expanding the scope of autonomous action across complex environments.

Advancing Autonomous Operations

With Dynatrace Intelligence, organizations can achieve:

  • Self-healing systems in dynamic, AI-driven environments
  • Proactive prevention, remediation, and optimization
  • Reliable autonomous action, leveraging both built-in agents and collaboration with partner agents, with full visibility and control

Teams remain in command while the system continuously manages operational complexity in the background.

The Journey to Autonomous Operations

Dynatrace Intelligence supports customers on a phased journey toward autonomy. Organizations can start by using AI-driven insights and recommendations, progress to leverage automation for supervised operations with human oversight, and ultimately advance to fully autonomous operations with guardrails and controls. This approach allows customers to safely adopt auto-prevention, auto-remediation, and auto-optimization, while maintaining control and building trust at every step.

“Agentic AI offers enormous potential, but many businesses still struggle to ensure it operates reliably, securely, and with consistent performance in real‑world environments,” said Bernd Greifeneder, Chief Technology Officer and Founder at Dynatrace. “Dynatrace Intelligence fuses deterministic and agentic AI, removing the guesswork and delivering AI‑powered observability organizations can trust.”

“As our digital environment grows more complex, we’re looking to move beyond reactive operations and manual intervention,” said Alexander Bicalho, Senior Director of Engineering at Autodesk. “What Dynatrace is outlining with Dynatrace Intelligence aligns with where we want to go—using trusted data and insights to support more autonomous operations. An approach that connects insight to action, while keeping our teams in control, could significantly improve performance and reliability as we scale. It’s observability that doesn’t just detect problems—it understands them and acts on them reliably.”

“The evolution of observability platforms is moving from manual root cause analysis to preventive operations. Organizations are progressing beyond reactive monitoring toward autonomous operations models that combine deterministic AI with agentic AI systems, with AI agents operating at different autonomy levels to orchestrate workflows across integrated ecosystems that span cloud platforms, development tools, and IT service management systems,” said Stephen Elliot, Group Vice President at IDC.

The Latest

IT organizations have historically measured success by how quickly they can respond when something goes wrong. The entire discipline of Incident Management has been optimized around mean time to resolution, first-response SLAs and ticket closure rates. But new research suggests that even though this is a well-executed playbook, it's no longer enough to retain customers ...

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...