Skip to main content

The Managed Database Default Is a Resilience Risk

Bennie Grant
Percona

Managed database services, often delivered through database-as-a-service (DBaaS) platforms, have become a pillar of modern engineering workflows. And for many stretched teams, that has been a rational response to pressure. If they need a new environment, they can simply choose a tier, wire up access, and let the provider handle patching, backups, scaling, monitoring, and failover. It removes real operational work and lets developers spend more time building.

The problem is that convenience has a way of becoming architecture. A decision made to move faster early on can harden into the operating model for systems that later become business-critical. By then, the question is not whether the service is useful, but whether the organization can keep operating when the provider, region, API, or management layer it depends on is no longer available.

Control the Control Plane

Most managed database services are designed to survive common infrastructure failures. Data is replicated across availability zones, backups are often automated, and failover may be built into the product. And while those capabilities matter and are important, they can also create a false sense of completion, as if availability inside the provider's architecture automatically translates into resilience for the customer.

Many of the actions organizations need during an incident depend on the control plane. In many DBaaS offerings, that layer is effectively a black box. The logic behind backup orchestration, security patching, scaling, failover, and network changes often sits inside proprietary systems controlled by the provider, not the customer. When that layer is degraded, the underlying database may still exist, but the customer’s ability to act becomes constrained.

AWS's own disaster recovery guidance draws this line clearly. It defines the data plane as the part delivering real-time service and the control plane as the part used to configure the environment, then advises organizations that need maximum resiliency to rely on data-plane operations during failover. So while managed recovery features are valuable day-to-day, customers still have to own their own recovery path in the event the control plane fails.

That recovery path cannot be improvised in the middle of an incident. By the time a control-plane failure exposes how dependent a team is on provider-specific tooling, the architecture has already made many of the important decisions for them.

Portability Should Be Built Long Before an Exit

Vendor lock-in is often discussed as though it begins with a contract renewal or a price increase. In practice, it usually begins much earlier, with small choices that compound over time. It may start with a provider-specific extension, or an integrated monitoring workflow, a backup process tied to the platform, or an identity model that assumes the same cloud will always be there. Each choice may be sensible on its own, but taken together, turns the database from a portable open source engine into a workload that can only operate comfortably inside one provider's ecosystem.

If the organization waits until it has a business reason to leave, whether because of cost, regulation, service quality, geopolitical constraints, or resilience concerns, it may discover that the database can move in theory but not in practice. Instead, portability needs to be treated as an architectural property, rather than a future migration project.

Regulators are starting to look at these types of dependencies more closely. The Digital Operational Resilience Act (DORA) is now applicable to EU financial entities, and says firms must maintain registers of ICT third-party arrangements so supervisors can assess dependency and risk. The EBA's technical standards on critical ICT services also emphasize control of operational risk, information security, and business continuity across the life of those arrangements. NIS2 pushes in a similar direction, including supply chain security and board-level accountability.

As a collective, the regulatory message is clear that outsourcing infrastructure does not outsource accountability. If a database supports an important service, organizations need to explain how they will continue operating when assumptions break.

A Better Default

None of this is an argument against managed services or cloud infrastructure. The cloud is still one of the most effective ways to scale systems, and managed services are a great way to reduce undifferentiated operational work. The mistake is treating managed as synonymous with resilient.

A better default starts with practical design questions:

Can the organization recover data without depending entirely on the provider's primary management interface?

Are backups independently testable and restorable?

Can failover work during a control-plane disruption?

Are encryption keys controlled at the level the data classification requires?

Can the database run on upstream open source software without proprietary hooks that make exit impractical?

For some workloads, the answer may still be a fully managed database. But the safer default is to build around upstream open source software, whether it is delivered through a managed deployment, a self-managed environment, a multi-cloud architecture, or a hybrid model. Proprietary forks and vendor-specific features can create technical cliffs over time, coupling the application to hooks that do not exist in the broader open source community. When an exit becomes necessary, that dependency can turn portability from an architectural option into a costly re-engineering project. The goal is to preserve security, sovereignty, and operational control by making sure the level of dependency matches the criticality of the workload.

Convenience has real value, and, of course, so does speed. But neither should be mistaken for operational control. The organizations that navigate outages, regulatory scrutiny, and geopolitical uncertainty most effectively will not avoid managed services entirely. Instead, they will know what they have delegated, what they have retained, and how they will operate when the easy path is no longer available.

Bennie Grant is COO of Percona

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

The Managed Database Default Is a Resilience Risk

Bennie Grant
Percona

Managed database services, often delivered through database-as-a-service (DBaaS) platforms, have become a pillar of modern engineering workflows. And for many stretched teams, that has been a rational response to pressure. If they need a new environment, they can simply choose a tier, wire up access, and let the provider handle patching, backups, scaling, monitoring, and failover. It removes real operational work and lets developers spend more time building.

The problem is that convenience has a way of becoming architecture. A decision made to move faster early on can harden into the operating model for systems that later become business-critical. By then, the question is not whether the service is useful, but whether the organization can keep operating when the provider, region, API, or management layer it depends on is no longer available.

Control the Control Plane

Most managed database services are designed to survive common infrastructure failures. Data is replicated across availability zones, backups are often automated, and failover may be built into the product. And while those capabilities matter and are important, they can also create a false sense of completion, as if availability inside the provider's architecture automatically translates into resilience for the customer.

Many of the actions organizations need during an incident depend on the control plane. In many DBaaS offerings, that layer is effectively a black box. The logic behind backup orchestration, security patching, scaling, failover, and network changes often sits inside proprietary systems controlled by the provider, not the customer. When that layer is degraded, the underlying database may still exist, but the customer’s ability to act becomes constrained.

AWS's own disaster recovery guidance draws this line clearly. It defines the data plane as the part delivering real-time service and the control plane as the part used to configure the environment, then advises organizations that need maximum resiliency to rely on data-plane operations during failover. So while managed recovery features are valuable day-to-day, customers still have to own their own recovery path in the event the control plane fails.

That recovery path cannot be improvised in the middle of an incident. By the time a control-plane failure exposes how dependent a team is on provider-specific tooling, the architecture has already made many of the important decisions for them.

Portability Should Be Built Long Before an Exit

Vendor lock-in is often discussed as though it begins with a contract renewal or a price increase. In practice, it usually begins much earlier, with small choices that compound over time. It may start with a provider-specific extension, or an integrated monitoring workflow, a backup process tied to the platform, or an identity model that assumes the same cloud will always be there. Each choice may be sensible on its own, but taken together, turns the database from a portable open source engine into a workload that can only operate comfortably inside one provider's ecosystem.

If the organization waits until it has a business reason to leave, whether because of cost, regulation, service quality, geopolitical constraints, or resilience concerns, it may discover that the database can move in theory but not in practice. Instead, portability needs to be treated as an architectural property, rather than a future migration project.

Regulators are starting to look at these types of dependencies more closely. The Digital Operational Resilience Act (DORA) is now applicable to EU financial entities, and says firms must maintain registers of ICT third-party arrangements so supervisors can assess dependency and risk. The EBA's technical standards on critical ICT services also emphasize control of operational risk, information security, and business continuity across the life of those arrangements. NIS2 pushes in a similar direction, including supply chain security and board-level accountability.

As a collective, the regulatory message is clear that outsourcing infrastructure does not outsource accountability. If a database supports an important service, organizations need to explain how they will continue operating when assumptions break.

A Better Default

None of this is an argument against managed services or cloud infrastructure. The cloud is still one of the most effective ways to scale systems, and managed services are a great way to reduce undifferentiated operational work. The mistake is treating managed as synonymous with resilient.

A better default starts with practical design questions:

Can the organization recover data without depending entirely on the provider's primary management interface?

Are backups independently testable and restorable?

Can failover work during a control-plane disruption?

Are encryption keys controlled at the level the data classification requires?

Can the database run on upstream open source software without proprietary hooks that make exit impractical?

For some workloads, the answer may still be a fully managed database. But the safer default is to build around upstream open source software, whether it is delivered through a managed deployment, a self-managed environment, a multi-cloud architecture, or a hybrid model. Proprietary forks and vendor-specific features can create technical cliffs over time, coupling the application to hooks that do not exist in the broader open source community. When an exit becomes necessary, that dependency can turn portability from an architectural option into a costly re-engineering project. The goal is to preserve security, sovereignty, and operational control by making sure the level of dependency matches the criticality of the workload.

Convenience has real value, and, of course, so does speed. But neither should be mistaken for operational control. The organizations that navigate outages, regulatory scrutiny, and geopolitical uncertainty most effectively will not avoid managed services entirely. Instead, they will know what they have delegated, what they have retained, and how they will operate when the easy path is no longer available.

Bennie Grant is COO of Percona

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...