
Gremlin and Dynatrace announced a strategic integration designed to streamline reliability testing for Kubernetes environments.
This collaboration makes it safe and simple to perform Fault Injection tests, empowering organizations to enhance their reliability programs and maintain their Kubernetes applications in a desired state.
"With this new integration, Gremlin and Dynatrace are simplifying how organizations introduce fault injection into their Kubernetes environments," said Samuel Rossoff, CTO of Gremlin. "Many teams have faced challenges operationalizing reliability testing across complex cloud-native architectures, often requiring multiple manual steps to identify and target the right resources. By combining advanced AI observability and topology insights with Gremlin's fault injection and reliability capabilities, customers can more easily identify, test, optimize, and strengthen critical services at scale."
Starting today, Kubernetes services are automatically discovered within Gremlin, powered by Dynatrace's AI-driven observability and topology mapping. Health checks are then applied to Kubernetes' objects, allowing organizations to efficiently implement standardized reliability testing and gain deeper insights into their environments.
"Kubernetes is the foundation of modern cloud-native infrastructure, supporting a wide spectrum of organizations, from nimble startups to global enterprises," said Wayne Segar,Global Field CTO at Dynatrace. "As AI-driven innovation accelerates, the reliability of Kubernetes becomes mission-critical. Our partnership with Gremlin simplifies chaos engineering, helping teams ensure resilience and performance across complex, distributed systems."
By making Fault Injection testing faster, simpler, and safer, this partnership between Gremlin and Dynatrace underscores their shared mission: helping engineering teams build more resilient systems, reduce risk, and deliver exceptional reliability at scale.
The Latest
Artificial intelligence (AI) has become the dominant force shaping enterprise data strategies. Boards expect progress. Executives expect returns. And data leaders are under pressure to prove that their organizations are "AI-ready" ...
Agentic AI is a major buzzword for 2026. Many tech companies are making bold promises about this technology, but many aren't grounded in reality, at least not yet. This coming year will likely be shaped by reality checks for IT teams, and progress will only come from a focus on strong foundations and disciplined execution ...
AI systems are still prone to hallucinations and misjudgments ... To build the trust needed for adoption, AI must be paired with human-in-the-loop (HITL) oversight, or checkpoints where humans verify, guide, and decide what actions are taken. The balance between autonomy and accountability is what will allow AI to deliver on its promise without sacrificing human trust ...
More data center leaders are reducing their reliance on utility grids by investing in onsite power for rapidly scaling data centers, according to the Data Center Power Report from Bloom Energy ...
In MEAN TIME TO INSIGHT Episode 21, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses AI-driven NetOps ...
Enterprise IT has become increasingly complex and fragmented. Organizations are juggling dozens — sometimes hundreds — of different tools for endpoint management, security, app delivery, and employee experience. Each one needs its own license, its own maintenance, and its own integration. The result is a patchwork of overlapping tools, data stuck in silos, security vulnerabilities, and IT teams are spending more time managing software than actually getting work done ...
2025 was the year everybody finally saw the cracks in the foundation. If you were running production workloads, you probably lived through at least one outage you could not explain to your executives without pulling up a diagram and a whiteboard ...
Data has never been more central to a greater portion of enterprise operations than it is today. From software development to marketing strategy, data has become an essential component for success. But as data use cases multiply, so too does the diversity of the data itself. This shift is pushing organizations toward increasingly complex data infrastructure ...
Enterprises are not stalling because they doubt AI, but because they cannot yet govern, validate, or safely scale autonomous systems, according to The Pulse of Agentic AI 2026, a new report from Dynatrace ...
For most of the cloud era, site reliability engineers (SREs) were measured by their ability to protect availability, maintain performance, and reduce the operational risk of change. Cost management was someone else's responsibility, typically finance, procurement, or a dedicated FinOps team. That separation of duties made sense when infrastructure was relatively static and cloud bills grew in predictable ways. But modern cloud-native systems don't behave that way ...
