How to keep your mission-critical cloud workloads running

How to keep your mission-critical cloud workloads running

Four things are necessary for application-level resilience and high availability: clustering, data replication, failover, and disaster recovery. Let’s look at each and how to architect for them. Entrusting mission-critical workloads in the cloud puts a lot of pressure on IT operations to build resilience into data infrastructure they don’t even own — to make sure that when an incident happens elsewhere, the lights stay on at home. Fortunately, savvy CIOs are learning they can architect their systems for resilience and high availability to keep mission-critical cloud workloads running when disaster strikes. When designing for resilient, highly available cloud workloads it is critical to understand the details of your cloud provider’s service level agreement (SLA) and recognize that providers operate on a shared responsibility model. That means there are many things the provider is responsible for, but the workloads you’ve got running on their infrastructure are not among them. With that in mind, we’ll need to focus on three core goals essential to building and maintaining highly available, resilient cloud workloads: Identify and eliminate single points of failure Close or mitigate conditions leading to recovery gaps, and Overcome operational complexity. Remember, everything fails Designing for application-level resilience and high availability is predicated on recognizing that everything fails. That’s not pessimism, but realism, and embracing this fact gives you a leg up in architecting resilient IT. Hardware breaks down, disasters occur, mistakes are made, software incompatibilities are common, and limits are exceeded that result in loss of service. There are four categories of failure that need to be taken into consideration when designing for resilience and high availability: Single point of failure: As the term suggests, this describes a non-redundant device, route, or point of connectivity that results in zero fault tolerance in the event of an incident involving that choke point. Excessive load: Exceeding the throughput or memory limits of a device or service can result in latency, service degradation, or failure affecting service level agreements (SLA) and service level objectives (SLO). Misconfigurations and bugs: Perhaps the most common source of failure, these breakdowns often occur during installation, maintenance, and patching leading to the incorrect execution of code. Shared fate: This type of failure is like non-redundant connectivity. When you rely on third-party services and there is an incident beyond your control affecting a partner, it can result in downtime or service degradation. Having identified the primary challenges to maintaining service uptime and continuity, let’s explore an approach to achieving resilience through high availability, disaster recovery, and the process of continuous improvement. Working toward continuous improvement High availability is often thought of as measured by “X Nines,” representing the percentage of time a service is operational over the course of a specified time frame, usually a year. Six nines of availability, for example, means that a service can be expected to run 99.9999% of the time, or 8759.99 hours out of 8760 in a year. That sounds good until a disaster happens. Then we shift to disaster recovery, or the ability to restore a service and its associated data quickly and correctly when things go wrong. Disaster recovery is often measured and evaluated by a recovery point objective (RPO), which is the maximum acceptable amount of data loss, and a recovery time objective (RTO), which is the maximum acceptable amount of time a service can be down. By learning from our mistakes, incidents, and the ongoing process of testing, observing, and evaluating IT operations, we apply new information in the goal of continuous improvement to get peak performance out of our investments in IT. But because we’re addressing resilience and high availability in the context of workloads operating in the cloud, we should understand the shared responsibility model that most cloud service providers use. Details may differ from provider to provider, but the shared responsibility model of cloud services is your vendor’s statement of responsibility for maintaining its infrastructure and ability to deliver storage and compute capacity in a region and availability zone where you have a need. In return, you accept responsibility for the configuration, maintenance, availability and operation of the workloads you choose to run on your vendor’s cloud. SIOS Technology A cloud service provider will typically operate and maintain physical infrastructure in different areas around the world, with multiple availability zones in each area. Each availability zone will have one or more data centers with its own power, cooling, and other operational support. Data centers are typically located far enough from each other to minimize the chance that a natural or regional disaster will affect multiple data centers, but close enough that they can operate as a single, logical data center — thereby maximizing bandwidth and throughput, while minimizing latency. This is redundancy at scale that eliminates a single point of data center failure. Designing for application high availability and disaster recovery To build out or retrofit your infrastructure for application-level resilience and high availability, there are four things you must have: clustering, data replication, failover, and disaster recovery. Let’s look at each and how to architect for them. Clustering Clustering is the practice of connecting like and complementary systems to improve availability and reliability. Think of it as putting into practice the maxim “the whole is greater than the sum of its parts.” Traditionally, clustering was achieved using hardware, specifically storage area networks (SAN), to connect and support systems in a physical data center. However, in the cloud age, hardware-based clustering has proven to be expensive and inflexible. Software-based SANless clustering with products such as SIOS LifeKeeper delivers all the benefits of the traditional approach while also working better in cloud and hybrid environments. SIOS Technology SANless clusters enable the configuration of nodes based on the organization’s operational priorities and work easily for enterprises running geographically distributed data centers or cloud availability zones. Because SIOS LifeKeeper supports single-site, multi-site, cloud, or mixed environments, and affords real-time access to operational data during incidents requiring failover, it is ideal when mission-critical workloads are distributed across multiple providers. LifeKeeper clusters create redundancy and eliminate single points of failure, both of which are necessary for achieving high availability. For enterprises operating hybrid or multi-cloud infrastructure, this enables you to shift to secondary resources should your primary service experience an outage. Data replication IT operations and your mission-critical applications rely on data, so it’s vital to keep data in sync between nodes in real time, whether synchronously or asynchronously, depending on the application. When working to achieve high availability, it’s important to monitor the health of infrastructure elements and the applications you’re running. This ensures that downtime affecting the application and any downstream dependencies is detected so that failover and recovery to a standby node can occur automatically. When everything is in sync and data replication is happening, this ensures complete readiness; there’s no need to rebuild a database or for any manual operations. And with SANless clusters using SIOS DataKeeper, this will happen across availability zones, regions, and to or from on-premises, cloud, or multi-cloud resources in hybrid environments. SIOS Technology What’s important here is that the software you use for SANless clustering be application aware. That’s SIOS LifeKeeper. Because high availability and resilience mean not just failing over to standby infrastructure resources, but for properly failing over mission-critical workloads (think SQL Server, SAP, Postgres, or industry-specific applications like Meditech), the software needs to understand how those applications work so that they can recover and restart cleanly. Otherwise, when failover occurs, there’s a good chance that workload is not going to be available, or it could be out of sync. When your software is application-aware, failover and recovery is fast and predictable with no API dependencies or last-minute provisioning. Failover With SIOS SANless application-aware clustering and replication in place, you can automate real-time failover and achieve the goal of high availability and resilience. Because we’ve chosen a software-based approach, failover occurs in the data plane, so there’s no need to spin up a new instance following an incident. Everything is already running, in sync, and ready. Furthermore, failover is immediate and it’s predictable. That is what you need when dealing with mission-critical workloads. Consider that many of the mission-critical applications you use were not designed to operate in the cloud, so even if your infrastructure is resilient, the application itself represents a single point of failure. Without replication, application monitoring, and real-time failover automation, you risk loss of revenue, security degradation, or worse. SIOS Technology Depending on the size of your enterprise and the industry in which it operates, an hour of downtime can cost $1 million or more, and 90% of organizations report costs of at least $300,000 per hour. Beyond that simple loss of revenue are soft costs associated with reputation damage, lost productivity, missed transactions, and more. Revenue protection is only part of the equation; it is about protecting the business. Disaster recovery Any design for high availability and resilience must also take into consideration the risk of a non-IT disaster that disrupts operations. Natural disasters like fires, earthquakes, and floods, or disasters caused by humans like the accidental severing of a trunkline or an act of sabotage conspire to bring operations to a sudden halt. Clustering, replication, and failover all play into disaster recovery, but there are additional considerations that must be addressed. For disaster recovery (DR), you’re going to want geographic distance between your primary site and your DR site to minimize the risk that a regional event doesn’t affect both. That means you’re going to want asynchronous (or near real-time) data replication. SIOS Technology Asynchronous replication is typically only a second or less behind synchronous replication, so it is de facto real-time, but that delay allows you to replicate across higher-latency networks and across longer distances. And because replication is happening continuously, even if you have a local failover between node one and node two, replication to the DR node doesn’t stop. Instead, when node two comes online it becomes the source of the mirror replicating to the DR site and back to node one once it is restored. Patching and planned maintenance Beyond the simplicity and cost advantages that come from taking a software-based approach to designing your cloud and hybrid infrastructure for high availability there is another big benefit. You are also better prepared to respond to the increased pressure on rapid implementation of software updates and security patches. No one wants to be the person responsible for taking too long to patch for a vulnerability, but every time you touch a system, you’re introducing the risk of error. That risk is greater when dealing with the cloud where you have less control over the underlying infrastructure, more abstraction, and a lot more moving parts. In short, patches can break things as well as fix them. They have thus become a significant contributing factor as a cause of downtime. With SANless clusters you’re able to minimize that risk by enabling rolling updates, meaning you can patch a secondary server, verify that the patch didn’t break anything, and then patch the original server where the primary application is still running. If something does break on the secondary server, the primary is unaffected and you can quickly recover to the previous state while you figure out what went wrong. This simplifies the process of executing ad hoc patches and performing planned maintenance, mitigating the associated risks. An investment in cloud and hybrid infrastructure and services doesn’t have to put IT operations at the mercy of a third party. With the right approach, you can design for high availability and resilience and keep your mission-critical workloads running in the cloud even when things outside of your control go wrong. — New Tech Forum provides a venue for technology leaders—including vendors and other outside contributors—to explore and discuss emerging enterprise technology in unprecedented depth and breadth. The selection is subjective, based on our pick of the technologies we believe to be important and of greatest interest to InfoWorld readers. InfoWorld does not accept marketing collateral for publication and reserves the right to edit all contributed content. Send all inquiries to doug_dineley@foundryco.com.

Original Source

Read the full article at Infoworld →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.