The Amazon S3 Outage - When the Internet's Hard Drive Fails …
March 06, 2017

Denis Goodwin
SmartBear

Share this

Last week, SmartBear observed a sudden and protracted 5X increase in web page timeout errors associated with the failure of Amazon's S3 cloud-based storage service. Looking a bit more closely at our data, we dug up a few more interesting angles on the impact of the failure.

Event Timeline

The issue hit suddenly – we saw an immediate spike in errors at 12:35 p.m. EST, and by 12:45 p.m. EST, error rates were 5X normal. For some specific types of timeout errors, the spike was more than 10X normal. At 3:30 p.m. EST, the error rate began dropping and by 3:50 p.m. EST, rates had returned to normal.


Web vs. API

The issue hit web pages hard, while API monitors were not noticeably impacted by this outage. Web pages and web apps often utilize content storage hosted by cloud services such as Amazon S3.

Common failure scenarios on Tuesday included page elements failing to load, which could cause either the whole web page to time out or specific content on a page might not render. Depending on the design of a given page, this partial content failure could be relatively minor or it could render a critical web journey non-functional. File uploads and downloads that rely on S3 storage endpoints were particularly hard hit.

In order to get a complete picture of application health, it's necessary to monitor your real user's journey through the application. The monitored user journeys that depended heavily on content hosted in S3 failed. Those that didn't have that dependency continued functioning. I personally experienced this with Slack – I was able to use the app, however files could not be uploaded presumably because these files are stored by Slack using S3 as the storage mechanism.

While far less pronounced than the spike in errors, some response time degradation was observed in API monitors that continued running successfully. Given that the issue affected Amazon's storage services rather than their hosting services for applications, this makes sense.

Geographic Impact

The issue was more acutely felt in the United States, but we observed impacts all over the globe. The spike in page errors was seen on websites dependent on Amazon S3, many of which are U.S.-hosted websites that are likely monitored from U.S. locations. Unsurprisingly, error counts spiked by as much as 25X in some U.S. monitoring locations. While not as significant as the U.S. locations, timeout and page error increases were also observed from Canada, Europe and Asia.

Takeaways

Much of the web is built on the backs of cloud providers. Most of the time, these cloud services provide a great user experience. Amazon will learn from the root cause of this issue and likely emerge from this outage more resilient than ever. It's impossible to control all aspects of these shared services – but here are three steps to take that are in your control.

1. Identify your business critical applications

2. Proactively monitor user journeys on these applications

3. Don't rely on your third party provider to tell you when it is down

It is key to utilize independent monitoring services to ensure your applications are up, functioning correctly and fast. Furthermore, missing content can be catastrophic or merely inconvenient to a critical user journey – it's important that your monitoring tool can be configured to know the difference.

Denis Goodwin is Director of Product Management, AlertSite, SmartBear Software.

Share this

The Latest

April 21, 2017

In the spirit of Earth Day, which is Saturday, April 22, we recently asked IT professionals for the tips and tricks they're using to help keep their data centers as green as possible. Here are a few ideas inspired by the responses we got ...

April 20, 2017

Almost One-Third (28 percent) of IT workers surveyed fear that cloud adoption is putting their job at risk, according to a survey conducted by ScienceLogic ...

April 19, 2017

A majority of senior IT leaders and decision-making managers of large companies surveyed around the world indicate their organizations have yet to fully embrace the aspects of IT Transformation needed to remain competitive, according to a new study conducted by Enterprise Strategy Group (ESG) ...

April 18, 2017

The move to cloud-based solutions like Office 365, Google Apps and others is one of the biggest fundamental changes IT professionals will undertake in the history of computing. The cost savings and productivity enhancements available to organizations are huge. But these savings and benefits can't be reaped without careful planning, network assessment, change management and continuous monitoring. Read on for things that you shouldn't do with your network in preparation for a move to one of these cloud providers ...

April 17, 2017

One of the most ubiquitous words in the development and DevOps vocabularies is "Agile." It is that shining, valued, and sometimes elusive goal that all enterprises strive for. But how do you get there? How does your organization become truly Agile? With these questions in mind, DEVOPSdigest asked experts across the industry — including analysts, consultants and vendors — for their opinions on the best way for a development or DevOps team to become more Agile ...

April 12, 2017

Is composable infrastructure the right choice for your IT environment? The following are 5 key questions that can help you begin to explore the capabilities of composable infrastructure and its applicability within your own IT environment ...

April 11, 2017

What is composable infrastructure, and is it the right choice for your IT environment? That's the question on many CIOs' minds today as they work to position their organizations as "digitally driven," delivering better, deeper, faster user experiences and a more agile response to change in whatever vertical market you do business in today ...

April 10, 2017

As companies adopt new hardware and applications, their networks grow larger and become harder to manage. For network engineers and administrators, the continued emergence of integrated technology has forced them to reconfigure and manage networks in a more dynamic way ...

April 07, 2017

The complexity of data in motion is growing and risks undermining the success of the modern data-driven enterprise. A recent survey of data engineers and architects, conducted by StreamSets, sought to bring some perspective to the new reality in the enterprise, leading to some interesting insights about the enterprise data landscape ...

April 06, 2017

The biggest factor driving happiness is the quality of relationships IT professionals have with their coworkers, including users, peers, and managers, according to the 2017 IT Job Satisfaction report from Spiceworks ...