Skip to main content

Automated, Flexible and Proactive: 3 Keys to Reducing Toil and Burnout in DevOps

Dan McCall
PagerDuty

Every business is in a constant battle to maximize efficiency, minimize toil, and scale sustainably in a moment of macroeconomic pressure. These goals are challenging in the best of times, but our current environment — continued staffing shortages, hiring freezes, and economic uncertainty — all make it significantly harder.

Because of these pressures, and the increased importance of digital operations to customer experience, teams are under more stress than ever to deliver seamless customer experiences. A recent report found that over 60% of developers are responding to off-hours work alerts on weekly basis and nearly half worked more hours in 2021 than they did in 2020. Companies are working urgently to mature their digital operations, including making incident response strategies more intelligent.

Resiliency at scale requires businesses to become more data-driven than ever before to get ahead of problems before they arise Incident response is essential to digital infrastructure and is at the crux of building a resilient enterprise. Addressing customer issues in real-time means adopting an incident response strategy that is automated, flexible, and proactive.

This next-generation approach enables the automation of repetitive and mundane work, while separating important signals from the flood of noise across all digital services. With this in place, teams can address the most mission-critical incidents when they occur and get ahead of the underlying issues behind attrition and burnout.

By combining the expertise of humans and machines to reduce the manual toil that causes burnout, we allow our teams to have more time to focus on innovation, and mission-critical digital transformation initiatives, instead of firefighting.


1. Leverage machines for automation

First, it's time to recognize that leveraging machines for automation is key to not only achieving key business outcomes, but to reducing burden on the humans that build and maintain digital operations. Beyond automating manual tasks, the right tools can reduce alert fatigue and cut down on system noise by using a mix of data science techniques and machine learning to intelligently group alerts and remove interruptions. In turn, automation empowers teams to balance critical workloads, helping humans to work smarter and reduce the burden. This is paramount when teams are tightly staffed due to attrition, inability to back-fill, or just new team members

2. Adopt a flexible tech stack

Second, technical teams must adopt a flexible tech stack that addresses a multitude of unique business needs at scale. Businesses should look for tools that can easily plug into their existing systems, while maintaining security and compliance. When the market can change at a moment's notice, teams must have the resources at their disposal to react to change as it happens to minimize disruption to their workloads and to operations.

3. Shift from reactivity to proactivity

Finally, we must shift from reactivity to proactivity. The same report as above found only 8% of teams are currently classified as proactive. Proactive businesses often use intelligence to identify root problems to anticipate and prevent disruption down the line. We must help DevOps teams move toward a state of proactivity and prevention to manage and maintain their IT infrastructure's consistency, reliability, and resilience — which will in turn help teams streamline work and free up time.

Get Started

The path to improved incident response depends on where your business falls within the spectrum of operational maturity.

Those still in the manual and reactive stage must start small and stay focused. Put energy into turning manually documented steps into automated steps to enable opportunities for pockets of automation across your organization.

Companies in the responsive stage should work to standardize the incident response process and enable self-service. Standardization helps to build automation that can be reused across teams and services, while self-service empowers more than just your subject matter experts to leverage automation for greater value.

Once you're in the proactive stage, you should be running automation in response to incidents, creating auto-remediation capabilities, and removing some of the real-time burden placed on teams that do critical monitoring and remediation work.

This next phase of incident response will build resilient enterprises in the face of constant challenges. Once we combine the expertise of humans and machines to enable humans to do their most innovative work and embrace an approach that is automated, flexible, and proactive, teams will be able to do their jobs more efficiently and effectively than ever before.

Dan McCall is VP of Product Management, Incident Response, at PagerDuty

The Latest

From growing reliance on FinOps teams to the increasing attention on artificial intelligence (AI), and software licensing, the Flexera 2025 State of the Cloud Report digs into how organizations are improving cloud spend efficiency, while tackling the complexities of emerging technologies ...

Today, organizations are generating and processing more data than ever before. From training AI models to running complex analytics, massive datasets have become the backbone of innovation. However, as businesses embrace the cloud for its scalability and flexibility, a new challenge arises: managing the soaring costs of storing and processing this data ...

Despite the frustrations, every engineer we spoke with ultimately affirmed the value and power of OpenTelemetry. The "sucks" moments are often the flip side of its greatest strengths ... Part 2 of this blog covers the powerful advantages and breakthroughs — the "OTel Rocks" moments ...

OpenTelemetry (OTel) arrived with a grand promise: a unified, vendor-neutral standard for observability data (traces, metrics, logs) that would free engineers from vendor lock-in and provide deeper insights into complex systems ... No powerful technology comes without its challenges, and OpenTelemetry is no exception. The engineers we spoke with were frank about the friction points they've encountered ...

Enterprises are turning to AI-powered software platforms to make IT management more intelligent and ensure their systems and technology meet business needs for efficiency, lowers costs and innovation, according to new research from Information Services Group ...

The power of Kubernetes lies in its ability to orchestrate containerized applications with unparalleled efficiency. Yet, this power comes at a cost: the dynamic, distributed, and ephemeral nature of its architecture creates a monitoring challenge akin to tracking a constantly shifting, interconnected network of fleeting entities ... Due to the dynamic and complex nature of Kubernetes, monitoring poses a substantial challenge for DevOps and platform engineers. Here are the primary obstacles ...

The perception of IT has undergone a remarkable transformation in recent years. What was once viewed primarily as a cost center has transformed into a pivotal force driving business innovation and market leadership ... As someone who has witnessed and helped drive this evolution, it's become clear to me that the most successful organizations share a common thread: they've mastered the art of leveraging IT advancements to achieve measurable business outcomes ...

More than half (51%) of companies are already leveraging AI agents, according to the PagerDuty Agentic AI Survey. Agentic AI adoption is poised to accelerate faster than generative AI (GenAI) while reshaping automation and decision-making across industries ...

Image
Pagerduty

 

Real privacy protection thanks to technology and processes is often portrayed as too hard and too costly to implement. So the most common strategy is to do as little as possible just to conform to formal requirements of current and incoming regulations. This is a missed opportunity ...

The expanding use of AI is driving enterprise interest in data operations (DataOps) to orchestrate data integration and processing and improve data quality and validity, according to a new report from Information Services Group (ISG) ...

Automated, Flexible and Proactive: 3 Keys to Reducing Toil and Burnout in DevOps

Dan McCall
PagerDuty

Every business is in a constant battle to maximize efficiency, minimize toil, and scale sustainably in a moment of macroeconomic pressure. These goals are challenging in the best of times, but our current environment — continued staffing shortages, hiring freezes, and economic uncertainty — all make it significantly harder.

Because of these pressures, and the increased importance of digital operations to customer experience, teams are under more stress than ever to deliver seamless customer experiences. A recent report found that over 60% of developers are responding to off-hours work alerts on weekly basis and nearly half worked more hours in 2021 than they did in 2020. Companies are working urgently to mature their digital operations, including making incident response strategies more intelligent.

Resiliency at scale requires businesses to become more data-driven than ever before to get ahead of problems before they arise Incident response is essential to digital infrastructure and is at the crux of building a resilient enterprise. Addressing customer issues in real-time means adopting an incident response strategy that is automated, flexible, and proactive.

This next-generation approach enables the automation of repetitive and mundane work, while separating important signals from the flood of noise across all digital services. With this in place, teams can address the most mission-critical incidents when they occur and get ahead of the underlying issues behind attrition and burnout.

By combining the expertise of humans and machines to reduce the manual toil that causes burnout, we allow our teams to have more time to focus on innovation, and mission-critical digital transformation initiatives, instead of firefighting.


1. Leverage machines for automation

First, it's time to recognize that leveraging machines for automation is key to not only achieving key business outcomes, but to reducing burden on the humans that build and maintain digital operations. Beyond automating manual tasks, the right tools can reduce alert fatigue and cut down on system noise by using a mix of data science techniques and machine learning to intelligently group alerts and remove interruptions. In turn, automation empowers teams to balance critical workloads, helping humans to work smarter and reduce the burden. This is paramount when teams are tightly staffed due to attrition, inability to back-fill, or just new team members

2. Adopt a flexible tech stack

Second, technical teams must adopt a flexible tech stack that addresses a multitude of unique business needs at scale. Businesses should look for tools that can easily plug into their existing systems, while maintaining security and compliance. When the market can change at a moment's notice, teams must have the resources at their disposal to react to change as it happens to minimize disruption to their workloads and to operations.

3. Shift from reactivity to proactivity

Finally, we must shift from reactivity to proactivity. The same report as above found only 8% of teams are currently classified as proactive. Proactive businesses often use intelligence to identify root problems to anticipate and prevent disruption down the line. We must help DevOps teams move toward a state of proactivity and prevention to manage and maintain their IT infrastructure's consistency, reliability, and resilience — which will in turn help teams streamline work and free up time.

Get Started

The path to improved incident response depends on where your business falls within the spectrum of operational maturity.

Those still in the manual and reactive stage must start small and stay focused. Put energy into turning manually documented steps into automated steps to enable opportunities for pockets of automation across your organization.

Companies in the responsive stage should work to standardize the incident response process and enable self-service. Standardization helps to build automation that can be reused across teams and services, while self-service empowers more than just your subject matter experts to leverage automation for greater value.

Once you're in the proactive stage, you should be running automation in response to incidents, creating auto-remediation capabilities, and removing some of the real-time burden placed on teams that do critical monitoring and remediation work.

This next phase of incident response will build resilient enterprises in the face of constant challenges. Once we combine the expertise of humans and machines to enable humans to do their most innovative work and embrace an approach that is automated, flexible, and proactive, teams will be able to do their jobs more efficiently and effectively than ever before.

Dan McCall is VP of Product Management, Incident Response, at PagerDuty

The Latest

From growing reliance on FinOps teams to the increasing attention on artificial intelligence (AI), and software licensing, the Flexera 2025 State of the Cloud Report digs into how organizations are improving cloud spend efficiency, while tackling the complexities of emerging technologies ...

Today, organizations are generating and processing more data than ever before. From training AI models to running complex analytics, massive datasets have become the backbone of innovation. However, as businesses embrace the cloud for its scalability and flexibility, a new challenge arises: managing the soaring costs of storing and processing this data ...

Despite the frustrations, every engineer we spoke with ultimately affirmed the value and power of OpenTelemetry. The "sucks" moments are often the flip side of its greatest strengths ... Part 2 of this blog covers the powerful advantages and breakthroughs — the "OTel Rocks" moments ...

OpenTelemetry (OTel) arrived with a grand promise: a unified, vendor-neutral standard for observability data (traces, metrics, logs) that would free engineers from vendor lock-in and provide deeper insights into complex systems ... No powerful technology comes without its challenges, and OpenTelemetry is no exception. The engineers we spoke with were frank about the friction points they've encountered ...

Enterprises are turning to AI-powered software platforms to make IT management more intelligent and ensure their systems and technology meet business needs for efficiency, lowers costs and innovation, according to new research from Information Services Group ...

The power of Kubernetes lies in its ability to orchestrate containerized applications with unparalleled efficiency. Yet, this power comes at a cost: the dynamic, distributed, and ephemeral nature of its architecture creates a monitoring challenge akin to tracking a constantly shifting, interconnected network of fleeting entities ... Due to the dynamic and complex nature of Kubernetes, monitoring poses a substantial challenge for DevOps and platform engineers. Here are the primary obstacles ...

The perception of IT has undergone a remarkable transformation in recent years. What was once viewed primarily as a cost center has transformed into a pivotal force driving business innovation and market leadership ... As someone who has witnessed and helped drive this evolution, it's become clear to me that the most successful organizations share a common thread: they've mastered the art of leveraging IT advancements to achieve measurable business outcomes ...

More than half (51%) of companies are already leveraging AI agents, according to the PagerDuty Agentic AI Survey. Agentic AI adoption is poised to accelerate faster than generative AI (GenAI) while reshaping automation and decision-making across industries ...

Image
Pagerduty

 

Real privacy protection thanks to technology and processes is often portrayed as too hard and too costly to implement. So the most common strategy is to do as little as possible just to conform to formal requirements of current and incoming regulations. This is a missed opportunity ...

The expanding use of AI is driving enterprise interest in data operations (DataOps) to orchestrate data integration and processing and improve data quality and validity, according to a new report from Information Services Group (ISG) ...