Network Redundancy Design for MSP-Managed Sites
Most redundancy designs fail in the field because they never verify where circuits actually run.

A network goes down because the design assumed two circuits were independent when they shared the same points of failure all along. This piece is about redundancy on paper versus redundancy that survives contact with a backhoe, a bad storm, or a carrier's own internal wiring closet. MSPs sell uptime. What they're often actually selling, without knowing it, is the appearance of uptime, built on assumptions nobody checked.
What genuine path diversity requires, from the street to the rack
Diversity comes in three strengths, and most redundancy designs stop at the weakest one. Provider diversity means two different carrier names on two different invoices, and it sounds like protection but often isn't, since carriers lease last-mile capacity from a small pool of wholesale fiber owners in a given metro. Two circuits from two "different" providers can run through the same duct bank for the last half mile into a building. A single dig crew can take both down at once regardless of what the contracts say.
Technology diversity is a real step up. A terrestrial fiber circuit and a cellular LTE or 5G circuit fail for different reasons under different conditions: one goes down from a cut cable or a fiber cut at a splice point, the other degrades from tower congestion or RF interference. Because the failure mechanisms don't overlap, technology diversity buys something provider diversity can't. Full path diversity is the strongest version, and it means a different provider, a different access technology, a different physical route into the building, and a different entry point on the wall. Getting all four at once is expensive and not always necessary, but for a site where downtime is measured in real dollars per hour, it is the only version worth calling redundant.
Verifying this is a document request. It's a document request. An MSP has to ask each carrier for actual diversity documentation, not a verbal assurance from a sales rep who has no visibility into the wholesale agreements behind their own network. Beyond the paperwork, someone has to physically walk the building's entry points and record where every circuit actually comes in, then trace each one from the street to the demarc room to the rack, flagging every shared conduit and every shared piece of equipment along the way.
A topology diagram earns its keep here in a way a spreadsheet of circuit IDs never will. Put every circuit terminating at a site on the same diagram as the devices they connect to, and shared entry points jump out visually, the kind of thing that stays invisible in a list of account numbers and provider names. And for sites where no amount of fiber diversity solves the underlying problem (both paths still enter through one conduit, one wall, one point a backhoe can reach), point-to-point microwave or licensed wireless offers something no buried cable can match: there's no ground for a backhoe to dig into.
Choosing the right circuit mix: MPLS, DIA, broadband, and cellular for different site tiers
Four circuit types cover most of what an MSP will deploy, and each one trades cost against control in a different way. MPLS is the carrier-managed private WAN, with class-of-service guarantees, strong SLAs, and predictable latency, but it comes with long lead times, multi-year contracts, and a premium per-megabit price, typically $300 to $1,500 a month. DIA, or Dedicated Internet Access, gets you guaranteed bandwidth and BGP peering with SLA options in roughly the same price band, $300 to $1,500 a month, but provisions faster than MPLS usually allows.
Business broadband is the opposite trade: fast to install, cheap at $65 to $500 a month, and generous on raw capacity, but it's best-effort, usually asymmetric, and comes with no SLA. Cellular LTE or 5G is the low end of cost, $50 to $199 a month, and deploys in days without needing a local loop, which makes it genuinely portable. Its ceiling is throughput and latency that shifts with tower load and network conditions on a given day.
The hybrid pattern that makes sense for most sites puts MPLS or DIA under the traffic that can't tolerate degradation, and lets broadband carry everything else. That split gets the cost benefit of broadband without putting a priority application on a best-effort path where it doesn't belong.
Site tiering follows from this pretty directly. Core and headquarters locations warrant full path diversity, meaning DIA as primary with a technology-diverse secondary like cellular or microwave behind it. Mid-tier branch offices do fine on DIA or business broadband as primary with cellular as backup, a pattern that's both defensible and affordable. Small or temporary sites can run on cellular as the primary connection outright, as long as the failure domain gets documented even at that scale, so nobody discovers later that the "backup" shares a tower with the primary.
The math on this isn't subtle. Adding standard redundancy, a DIA circuit plus a broadband backup, runs somewhere in the range of $4,000 to $6,000 a month across a typical site portfolio, or $48,000 to $72,000 a year. A single four-hour outage at a major site, by comparison, runs $200,000 to $500,000 or more. That comparison alone settles the argument for any site where the business actually depends on being online.
Active-active vs. active-passive vs. N+1: choosing the architecture for the site's actual risk tolerance
Three architectures cover almost every redundancy design an MSP will build. Active-active runs both links carrying traffic at the same time, which gives faster failover than active-passive because traffic is already distributed across both paths before any failure occurs. It typically doubles the connectivity bill, but weighed against a single outage avoided at $5,600 to $9,000, the math tends to hold up.
Active-passive is the workhorse: the primary circuit carries everything, the secondary sits idle until it's needed, and the incremental cost is around $100 to $300 a month. It's the pattern most MSP-managed branch sites actually run, because it delivers real protection without doubling the connectivity spend. N+1, where one spare resource covers N active components, is the data center standard rather than a WAN pattern, and it's rarely an MSP's call to make, though it matters wherever a client keeps infrastructure on-premises.
Redundancy in general adds 40 to 80 percent to connectivity costs across a portfolio, so the choice of architecture per site should be a deliberate decision weighed against that number rather than a default applied everywhere because it's easier to standardize.
Topology across a multi-site client follows a similar logic. Hub-and-spoke is the common starting point for organizations with a lot of branches, but concentrating traffic through a central hub means the hub itself requires its own redundancy plan, or the whole design just relocates the risk rather than removing it. Full mesh suits latency-sensitive traffic like real-time voice and video, though cost climbs fast as site count grows. Partial mesh splits the difference for portfolios where full mesh cost is prohibitive, and ring topology offers an alternative for chains of point-to-point links, providing a measure of resilience where a single break would otherwise isolate downstream sites.
Applying one architecture uniformly across a client's whole site portfolio almost guarantees over-spending on the low-criticality locations while under-protecting the ones that actually matter. Tiering by site type, not by convenience of a single template, is the model that holds up.
Where SD-WAN fits: the intelligent overlay that makes failover fast, but doesn't replace diverse transport
SD-WAN adds real value on top of a redundant transport design: faster failover, traffic steering that understands which application is which, bandwidth aggregation across links that aren't the same technology, and continuous measurement of path quality rather than a periodic check. None of that creates diversity where none exists. It manages diversity that's already been built.
Vendor claims of "sub-second failover" describe detection plus decision time under clean lab conditions, and real-world numbers look different. Real-world cellular failover typically runs in the 1 to 10 second range, since probes need enough samples to avoid triggering on a false positive, and encrypted tunnels, whether IPsec, WireGuard, or a vendor's own overlay, need time to reroute flows without breaking them mid-stream. On fiber, SD-WAN failover after an outage can complete in under 30 seconds. Neither number is bad. Neither number is instant, either, and a design built around the marketing figure instead of the real one will disappoint someone during an actual event.
A vendor evaluation should specify more than the failover speed number. Proactive failover detects performance degradation before the path fully drops, rather than waiting for total loss. Granular failover redirects specific application traffic based on priority instead of an all-or-nothing switch of the entire link. Automatic recovery restores the primary path without someone needing to log in and flip it back. Continuity testing checks the backup path's health on an ongoing basis as well as the primary's.
None of this solves the problem of a client with one carrier and no real secondary WAN provider. If a storm takes out a carrier's point of presence and there's no diverse transport underneath the SD-WAN overlay, application-aware steering has nothing left to steer traffic toward. The transport diversity decisions made earlier in the design are what give SD-WAN something to actually work with.
Monitoring probe design: what should trigger a failover
A circuit can report as "up" while the applications riding it have already become unusable, and this gap, commonly called a brownout, is where a lot of redundancy designs quietly fail their users even though every dashboard shows green. A failover policy built only around total path loss will miss this every time. The policy has to treat a degraded-but-alive path as a trigger in its own right.
Good probe design measures more than one target, and does it simultaneously: the SD-WAN hub, a cloud region such as AWS, Azure, or Google Cloud, and a stable public IP address that isn't tied to any of the vendor's own infrastructure. Probing only one target means a single probe failure can hide, or misrepresent, what's actually happening on the path.
Four metrics carry the real signal. Loss is the fastest indicator of LTE or 5G fade or tower congestion, appearing before anything else does. Latency spikes often flag bufferbloat or RF retransmissions well before the path fails. Jitter is the one that breaks VoIP calls, Zoom meetings, and Microsoft Teams sessions first, since loss can stay low while jitter alone makes real-time communication unusable. Brownout detection ties these together: the path reports up, the application performance says otherwise, and the probe design and the failover policy both need to be built to catch that state and act on it, not wait for a total outage that may never come.
Getting this right pays off operationally in a way that's easy to underrate. According to California Telecom, 84.9 percent of organizations that adopted centralized management reported fewer on-site visits and fewer manual interventions, which is the practical payoff of solving what used to be limited visibility across a scattered site portfolio.
Failover testing: why a backup path that has never been tested is a hope, not a plan
A backup path is a plan on the day it's designed. Eighteen months later, after routing changes, firewall policy updates, firmware upgrades, new applications added to the network, and normal capacity growth, it may be something else entirely, and nobody finds out until the day the primary circuit actually goes down.
Drift appears in a handful of predictable ways. Failover mechanisms slip out of the configuration they were built with. UPS batteries the backup depends on degrade quietly, since nobody tests them under real load. Secondary links develop performance problems that go unnoticed because no traffic ever touches them, so the issue becomes visible only once that link is the only path left.
A controlled failover test has to confirm three things, not one. Traffic actually moves to the secondary when the primary is deliberately taken offline. Applications stay reachable when the specific workload the secondary was sized for actually runs on it, not just when a ping gets a reply. And the secondary path can carry that workload under realistic load, since an idle circuit passing a health check tells you almost nothing about what happens when real traffic lands on it.
Quarterly testing is the floor, not a suggestion. Plenty of organizations find out their backup doesn't work at the exact moment they need it most, which is the one moment testing exists to prevent.
Documentation as a design control: why redundancy designs decay silently without it
A redundancy design needs a short, specific record to stay a design at all: which circuit is primary and which is backup, the order in which failover happens, what capacity the backup path is expected to carry, which carrier contact applies to each circuit, and where each circuit physically enters the building. That's five facts, not a binder, and skipping any of them turns the design into something only one person understands, usually the person who isn't there when it matters.
Documentation is a control that keeps the design from decaying without anyone noticing. It's a control that keeps the design from decaying without anyone noticing. A redundancy setup that isn't written down at the point where the circuits are actually managed drifts out of sync with reality the moment the first change gets made to it, whether that's a new firewall rule, a re-terminated circuit, or a technician who reroutes something during an unrelated ticket and has no way of knowing what assumption they just broke.


