The Amazon S3 Outage - When the Internet's Hard Drive Fails …
March 06, 2017

Denis Goodwin
SmartBear

Share this

Last week, SmartBear observed a sudden and protracted 5X increase in web page timeout errors associated with the failure of Amazon's S3 cloud-based storage service. Looking a bit more closely at our data, we dug up a few more interesting angles on the impact of the failure.

Event Timeline

The issue hit suddenly – we saw an immediate spike in errors at 12:35 p.m. EST, and by 12:45 p.m. EST, error rates were 5X normal. For some specific types of timeout errors, the spike was more than 10X normal. At 3:30 p.m. EST, the error rate began dropping and by 3:50 p.m. EST, rates had returned to normal.


Web vs. API

The issue hit web pages hard, while API monitors were not noticeably impacted by this outage. Web pages and web apps often utilize content storage hosted by cloud services such as Amazon S3.

Common failure scenarios on Tuesday included page elements failing to load, which could cause either the whole web page to time out or specific content on a page might not render. Depending on the design of a given page, this partial content failure could be relatively minor or it could render a critical web journey non-functional. File uploads and downloads that rely on S3 storage endpoints were particularly hard hit.

In order to get a complete picture of application health, it's necessary to monitor your real user's journey through the application. The monitored user journeys that depended heavily on content hosted in S3 failed. Those that didn't have that dependency continued functioning. I personally experienced this with Slack – I was able to use the app, however files could not be uploaded presumably because these files are stored by Slack using S3 as the storage mechanism.

While far less pronounced than the spike in errors, some response time degradation was observed in API monitors that continued running successfully. Given that the issue affected Amazon's storage services rather than their hosting services for applications, this makes sense.

Geographic Impact

The issue was more acutely felt in the United States, but we observed impacts all over the globe. The spike in page errors was seen on websites dependent on Amazon S3, many of which are U.S.-hosted websites that are likely monitored from U.S. locations. Unsurprisingly, error counts spiked by as much as 25X in some U.S. monitoring locations. While not as significant as the U.S. locations, timeout and page error increases were also observed from Canada, Europe and Asia.

Takeaways

Much of the web is built on the backs of cloud providers. Most of the time, these cloud services provide a great user experience. Amazon will learn from the root cause of this issue and likely emerge from this outage more resilient than ever. It's impossible to control all aspects of these shared services – but here are three steps to take that are in your control.

1. Identify your business critical applications

2. Proactively monitor user journeys on these applications

3. Don't rely on your third party provider to tell you when it is down

It is key to utilize independent monitoring services to ensure your applications are up, functioning correctly and fast. Furthermore, missing content can be catastrophic or merely inconvenient to a critical user journey – it's important that your monitoring tool can be configured to know the difference.

Denis Goodwin is Director of Product Management, AlertSite, SmartBear Software.

Share this

The Latest

October 21, 2021

Scaling DevOps and SRE practices is critical to accelerating the release of high-quality digital services. However, siloed teams, manual approaches, and increasingly complex tooling slow innovation and make teams more reactive than proactive, impeding their ability to drive value for the business, according to a new report from Dynatrace, Deep Cloud Observability and Advanced AIOps are Key to Scaling DevOps Practices ...

October 20, 2021

Over three quarters (79%) of database professionals are now using either a paid-for or in-house monitoring tool, according to a new survey from Redgate Software ...

October 19, 2021

Gartner announced the top strategic technology trends that organizations need to explore in 2022. With CEOs and Boards striving to find growth through direct digital connections with customers, CIOs' priorities must reflect the same business imperatives, which run through each of Gartner's top strategic tech trends for 2022 ...

October 18, 2021

Distributed tracing has been growing in popularity as a primary tool for investigating performance issues in microservices systems. Our recent DevOps Pulse survey shows a 38% increase year-over-year in organizations' tracing use. Furthermore, 64% of those respondents who are not yet using tracing indicated plans to adopt it in the next two years ...

October 14, 2021

Businesses are embracing artificial intelligence (AI) technologies to improve network performance and security, according to a new State of AIOps Study, conducted by ZK Research and Masergy ...

October 13, 2021

What may have appeared to be a stopgap solution in the spring of 2020 is now clearly our new workplace reality: It's impossible to walk back so many of the developments in workflow we've seen since then. The question is no longer when we'll all get back to the office, but how the companies that are lagging in their technological ability to facilitate remote work can catch up ...

October 12, 2021

The pandemic accelerated organizations' journey to the cloud to enable agile, on-demand, flexible access to resources, helping them align with a digital business's dynamic needs. We heard from many of our customers at the start of lockdown last year, saying they had to shift to a remote work environment, seemingly overnight, and this effort was heavily cloud-reliant. However, blindly forging ahead can backfire ...

October 07, 2021

SmartBear recently released the results of its 2021 State of Software Quality | Testing survey. I doubt you'll be surprised to hear that a "lack of time" was reported as the number one challenge to doing more testing, especially as release frequencies continue to increase. However, it was disheartening to see that a lack of time was also the number one response when we asked people to identify the biggest blocker to professional development ...

October 06, 2021

The role of the CIO is evolving with an increased focus on unlocking customer connections through service innovation, according to the 2021 Global CIO Survey. The study reveals the shift in the role of the CIO with the majority of CIO respondents stating innovation, operational efficiency, and customer experience as their top priorities ...

October 05, 2021

The perception of IT support has dramatically improved thanks to the successful response of service desks to the pandemic, lockdowns and working from home, according to new research from the Service Desk Institute (SDI), sponsored by Sunrise Software ...