CloudDevelopersFeaturedLet's TalkOpen Source

SLAs And Four Nines Are Not Enough For High Availability In The Cloud


Guest: Dave Bermingham (LinkedIn, Twitter)
Company: SIOS Technology (Twitter)
Show: Let’s Talk

When most people think of high availability, they set four nines or less than five minutes of downtime every month as the baseline. But according to Dave Bermingham, Senior Technical Evangelist at SIOS Technology, high availability is more than that.

“When we look at the big picture of any service that needs to be available, there are many chains in that link. We call it the chain of high availability,” says Bermingham. “So you have to look at the big picture.”

He argues that counting on nines is really a measurement that you might be judged against, but really trying to guarantee a level of nines is almost impossible. Because there’s so many points in that availability chain that can be a single point of failure. Four nines is certainly a great number to be judged against and to strive for, but overall it doesn’t mean a lot to have just four nines for my database server.

Even with Cloud SLAs (Service Level Agreements), one can’t be fully rest assured as most cloud providers offer four nines on compute, which is only one part of the availability chain (along with network, storage, and the hops between). Bermingham warns, “There’s a million points of failure. So, trying to think that my cloud provider offers four nines so I’m covered, you’re kind of fooling yourself there. You have to look at the big picture and do what you can to identify those points of failures, to minimize the potential points of failure and to have a recovery plan, should something happen.”

When considering High Availability/Disaster Recovery (HA/DR), Bermingham believes the thing that causes the most visible downtime is human error. Bermingham also suggests that authorization and access to the system should also be restricted to reduce the point of failure. “You should only give access to those who absolutely need access to it and you should also ensure that they are highly trained and that you have all the things in place to help minimize potential oops.”

Another important tip Bermingham offers is to make sure your storage is highly available. To that, he says, “You’re never going to have more availability than the weakest link in that chain.” Other tips include having the ability to rapidly recover from corruption events and making sure you don’t have nefarious people breaking into your network.

The summary of the show is written by Jack Wallen

Read Transcript
Don't miss out great stories, subscribe to our newsletter.

Login/Sign up