{"id":78916,"date":"2026-10-09T03:21:56","date_gmt":"2026-10-09T03:21:56","guid":{"rendered":"https:\/\/www.devopsschool.com\/blog\/?p=78916"},"modified":"2026-10-09T03:21:57","modified_gmt":"2026-10-09T03:21:57","slug":"why-multi-cloud-doesnt-automatically-mean-high-availability","status":"publish","type":"post","link":"https:\/\/www.devopsschool.com\/blog\/why-multi-cloud-doesnt-automatically-mean-high-availability\/","title":{"rendered":"Why Multi-Cloud Doesn\u2019t Automatically Mean High Availability"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Multi-cloud architecture is often associated with resilience. Spread workloads across multiple cloud providers, the thinking goes, and an outage at one provider no longer has the ability to take an entire application offline.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The reality is more complicated.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Using multiple cloud providers can remove important single points of failure, but simply distributing workloads between AWS, <a href=\"https:\/\/www.devopsschool.com\/blog\/what-is-azure-and-use-cases-of-azure\/\">Microsoft Azure<\/a>, Google Cloud or other platforms does not automatically create a highly available system. Applications still depend on DNS, routing, network transit, physical infrastructure and connectivity between the application and its users.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">True high availability requires looking at the entire delivery path, not just the servers running at the other end of it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Multi-Cloud Solves Only Part of the Problem<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">There are legitimate resilience benefits to multi-cloud infrastructure.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An organization might run its primary environment with one provider while maintaining critical services or replicated data with another. If the primary provider experiences a major regional outage, traffic can potentially be shifted elsewhere.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That reduces dependence on a single cloud platform, but cloud providers are only one layer of modern application infrastructure.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A user attempting to reach an application may depend on a chain that includes a local internet provider, regional network infrastructure, transit providers, DNS services, content delivery networks and physical fibre routes before traffic ever reaches the cloud environment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Redundancy at the compute layer does little to help if another part of that chain fails.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Network Can Still Be a Single Point of Failure<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">One of the easiest mistakes in high-availability planning is assuming that logically separate services are also physically independent.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two cloud environments may use different providers while still depending on some of the same underlying infrastructure. Traffic could traverse common internet exchanges, long-haul fibre corridors or upstream network providers.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Even separate data centres can have hidden dependencies if their connectivity ultimately passes through the same physical routes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cMulti-cloud architecture can protect against a failure within one cloud provider, but it doesn\u2019t eliminate the network as a potential single point of failure,\u201d said Tomas Novosad, broadband analyst and founder of <a href=\"https:\/\/fibrebroadbandnz.co.nz\/\">Fibre Broadband NZ<\/a>. \u201cApplications can be distributed across multiple providers and regions while traffic still depends on shared fibre routes, transit providers or local access infrastructure. High availability has to account for the entire path between the application and the end user, not just where the workload is hosted.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This physical layer is easy to overlook precisely because cloud infrastructure makes computing resources appear abstract. Developers can provision servers, databases and storage across the world in minutes, but the data connecting those resources still has to travel through physical networks.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The expansion of fibre broadband and other high-capacity network infrastructure has made those connections significantly faster and more reliable, but it has not eliminated the need to understand where dependencies exist.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">DNS Deserves the Same Attention as Compute<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">DNS is another dependency that can undermine an otherwise well-designed multi-cloud architecture.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An application may have redundant instances running across several providers, but users still need to resolve its domain before they can reach any of them. If <a href=\"https:\/\/www.devopsschool.com\/blog\/dns-concepts\/\">DNS infrastructure<\/a> becomes unavailable or is incorrectly configured, healthy servers across multiple clouds may effectively become unreachable.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For critical systems, DNS therefore needs its own resilience strategy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That can include geographically distributed authoritative servers, carefully selected providers, appropriate TTL configuration and, for organizations with particularly demanding availability requirements, consideration of whether relying on a single DNS provider creates another unnecessary dependency.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The same principle applies to load balancers, authentication systems, monitoring platforms and other services sitting between users and application infrastructure.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Failover Has to Work Outside a Diagram<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Architecture diagrams make failover look straightforward. Provider A fails, health monitoring detects the outage and traffic moves to Provider B.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Production environments are rarely that clean.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Applications may depend on databases that have not fully replicated. Authentication could still rely on services running in the affected environment. DNS records may take time to update. Network routes can behave differently than expected, and the secondary environment may suddenly receive far more traffic than it normally handles.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A backup environment that exists but has never handled production-scale traffic is not necessarily a reliable failover environment.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This is why resilience testing matters as much as resilience design.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Teams can simulate provider failures, disable individual dependencies and test whether applications continue operating under degraded conditions. Chaos engineering takes this idea further by intentionally introducing failures to uncover dependencies that might otherwise remain invisible until a real incident occurs.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Observability Needs to Extend Beyond the Application<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Traditional application monitoring tends to focus on metrics such as <a href=\"https:\/\/www.devopsschool.com\/blog\/cpu-profiling-tools-uncovering-the-secrets-of-your-computers-performance\/\">CPU utilization<\/a>, memory consumption, response times and error rates.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Those metrics remain important, but distributed applications require visibility further into the network.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Packet loss, DNS resolution times, routing changes, latency between cloud regions and performance from different geographic locations can all reveal problems that server-side monitoring misses.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">An application can appear perfectly healthy from inside a cloud environment while users in a particular region struggle to reach it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Monitoring from multiple external locations provides a better picture of whether a service is actually available from the perspective that matters most: the user.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">High Availability Is an End-to-End Problem<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Multi-cloud can be a valuable part of a high-availability strategy. It can reduce dependence on a single provider, provide additional disaster recovery options and give engineering teams more flexibility when designing resilient systems.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But the number of cloud providers in an architecture is not a measure of its resilience.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A system running across three clouds can still contain a critical DNS dependency, a shared network path or an untested failover process capable of bringing everything down at once.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">High availability comes from identifying those dependencies and designing around them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That means thinking beyond compute instances and cloud regions to include DNS, routing, fibre infrastructure, transit networks, authentication, data replication and the connectivity used by the people ultimately accessing the application.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Multi-cloud removes one potential single point of failure. Good infrastructure engineering asks where the next one is.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Multi-cloud architecture is often associated with resilience. Spread workloads across multiple cloud providers, the thinking goes, and an outage at one provider no longer has the ability&#8230; <\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_joinchat":[],"footnotes":""},"categories":[11138],"tags":[],"class_list":["post-78916","post","type-post","status-publish","format-standard","hentry","category-best-tools"],"_links":{"self":[{"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/posts\/78916","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/comments?post=78916"}],"version-history":[{"count":1,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/posts\/78916\/revisions"}],"predecessor-version":[{"id":78917,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/posts\/78916\/revisions\/78917"}],"wp:attachment":[{"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/media?parent=78916"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/categories?post=78916"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.devopsschool.com\/blog\/wp-json\/wp\/v2\/tags?post=78916"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}