The Continuity Brief

Backup and DR Software Evaluation Criteria for MSPs

Ransomware now forces MSPs to prove backups work, not just run.

Correspondent · · 10 min read
Cover illustration for “Backup and DR Software Evaluation Criteria for MSPs”
Tooling & Vendors · September 30, 2026 · 10 min read · 2,350 words

MSP backup has moved from a commodity line item to the operational and commercial center of ransomware resilience. Platforms that looked adequate two years ago now fail on dimensions that regulators, insurers, and clients actively test. Three forces have converged to make that true at the same time. Ransomware campaigns built in 2026 are engineered to break recovery first, before they disrupt daily operations, because destroying a victim's ability to restore data increases the pressure to pay. Regulators have followed with enforcement: HHS OCR settled four ransomware investigations totaling $1,165,000 in April 2026, and underwriting reviews now ask for documented immutable, offline, or air-gapped copies as a baseline control, not an aspirational one.

The Backup-as-a-Service market is growing fast, with Mordor Intelligence's 2026 market report naming ransomware pressure and compliance demand as the main drivers behind that growth. Spend at that scale tells you buyers have stopped treating recovery as a cost center they tolerate and started treating it as the product they're actually buying. Clients, regulators, and cyber insurers no longer take "backups are running" as proof of anything. They want documented evidence that restores work under conditions that resemble an actual incident, and that means the criteria used to evaluate a backup platform have to be rebuilt around proof, not around whether a job completed last night.

Why MSPs evaluate backup differently than internal IT teams

An MSP running backup across dozens of clients faces failure modes that internal IT departments never encounter, because a single architectural weakness doesn't stay contained to one environment. It can expose every client sitting on that platform simultaneously. Internal IT protects one environment under one compliance posture. An MSP might be protecting a law firm's case files, a biopharma client's lab servers, and a five-person retailer's point-of-sale system in the same afternoon, each with its own recovery expectations, retention rules, and regulatory obligations.

That diversity creates a specific structural risk. If the MSP's own management console gets compromised, every client environment reachable from that console becomes reachable to the attacker. Tenant isolation is the architectural answer to a risk that only exists because one console touches many businesses at once.

Scale also changes what counts as a viable workflow. A platform that needs a site visit, or more than an hour of technician labor, just to onboard a new client will not survive contact with MSP economics. Internal IT can absorb friction like that once. An MSP absorbs it every time it signs a new account. The evaluation criteria that matter to an MSP, multi-tenancy, automation, isolation, and margin predictability, are first-order concerns that set whether a platform can be operated profitably.

Multi-tenancy depth as the first architectural gate

Multi-tenancy is where the evaluation has to start, because a platform that lacks genuine tenant isolation, role-based access control, and remote policy management fails at the level of structure before its other features even get tested. Genuine multi-tenancy means more than a dashboard that rolls every client into one consolidated view. NovaBACKUP's 2026 evaluation guide describes it as isolation between tenants, role-based access control, and the ability to remotely create, modify, and delete backup jobs for each client without touching that client's own system.

Platforms that require logging into separate portals for every client are operationally unscalable and increase the risk of missing backup failures. A well-built console solves this with notification tiering. Critical failures need to trigger something immediate, storage nearing capacity can wait for same-day attention, and routine summaries belong in a weekly digest rather than the same inbox as an emergency. That structure keeps the real problems from drowning in noise.

Vendors have built toward this differently. Kaseya Datto SIRIS runs a secure, multi-tenant cloud Datto Partner Portal built for managing many clients from one place, and Acronis Cyber Protect Cloud is built around multitenant design with centralized management, monitoring, alerts, and reporting. Whatever platform is under review, the test in a live demo is specific: can backup policies, retention rules, and reporting be configured independently, client by client, from that single console? Can a technician's access be scoped so they only see the clients assigned to them? A platform that can't answer yes to both questions isn't ready for MSP use, regardless of how strong its recovery engine looks on paper.

RTO and RPO specificity: where SLA commitments either hold or collapse

RTO and RPO numbers only mean something once the restore path behind them, the actual network bandwidth, the recovery method, the infrastructure already staged and waiting, can physically deliver on the promise during a real incident. RTO sets the clock. If a client's accounting server carries a four-hour RTO, every technology and workflow choice made for that server has to fit inside that window, and if the restore path can't hit it, the SLA breach clock starts the moment the incident does.

The math tends to punish optimism. A large VM restored over a WAN connection at its theoretical maximum throughput can already run well past a four-hour window, and real-world network conditions rarely hit theoretical maximums. The restore path has to be engineered around the commitment. Solutions include local restore appliances, image-based recovery run at the hypervisor layer, and pre-staged replicas, each changing the effective RTO for a given workload.

RPO works on a separate axis. Continuous data protection platforms journal every write and can promise RPOs measured in seconds, while snapshot-based tools trade that granularity for lower overhead, and the right answer depends on what a given workload can actually tolerate losing. Kaseya Datto SIRIS, for instance, supports instant virtualization that boots protected systems as VMs directly from backup in seconds to minutes, with an average RTO under six minutes, which is what engineering the restore path to close the RTO gap looks like in practice. Zerto is recognized for delivering the lowest RPOs across complex environments through continuous data protection.

These commitments shape pricing too. A base service tier carrying standard RPO and RTO commitments belongs at a different price point than a premium tier offering shorter recovery windows, immutable copies, and compliance documentation, and the criteria discussed here determine which tier a given workload should sit in. None of it should be taken on faith. RTO claims need to be tested against a real recovery scenario built from a representative workload before that language ever makes it into a client SLA.

Proof of recovery: why verified restores replace backup job completion as the success metric

A failover script that passes staging can still hit a permission error in production that nobody caught, turning what a documented runbook promised would take minutes into four hours of manual reconstruction, with no one owning the gap until the incident review. That scenario is common enough to reshape how the whole category gets measured. A completed backup job was never evidence that a system could actually be recovered. A completed backup job was never proof that a system could actually be recovered; a tested, documented restore is, and that distinction now appears directly in client contracts, insurer questionnaires, and regulatory audits.

Too many MSPs still find out about a gap in their recovery process only when a client urgently needs a restore, and by then the only thing left to deliver is an explanation. The platforms built to prevent that moment share a common feature: automated, scheduled proof. Screenshot Verification, run by Kaseya Datto SIRIS, automatically boots protected systems in an isolated sandbox and captures proof of recoverability as a screenshot, turning a promise into a documented artifact that can be handed to a client.

Non-disruptive testing, recovery drills that don't touch live production systems, has become one of the main things separating mature DR platforms from the rest. Expert Insights' review of the category found Acronis Advanced Disaster Recovery and Arcserve UDP strongest on replication performance and recovery testing accuracy among the platforms tested. Enterprise buyers increasingly won't sign a contract without evidence of tested recovery in hand, which puts pressure on MSPs serving that market to pick a platform that generates that evidence on its own rather than requiring a technician to assemble it manually after the fact. In a demo, that means asking for a scheduled, non-disruptive test and the report it produces, rather than a one-time manual restore of a dataset the vendor picked for the occasion.

Ransomware resilience: the architectural properties that protect backup copies

Modern ransomware campaigns target backup infrastructure early in the attack chain, deliberately, because eliminating an organization's recovery options is what gives an attacker leverage to demand payment. Backup infrastructure sits inside the attack surface now, not outside it.

The instinct to answer that threat with immutability alone is understandable and wrong. Immutable, write-once-read-many storage does protect backup snapshots from encryption or deletion even when an MSP's own management console has been compromised, which matters precisely because that console is a shared operational path across every client the MSP protects. But immutability solves one failure mode, not the whole threat. The strongest objection, that immutability solves ransomware, is false: ransomware-safe recovery also requires end-to-end encryption, MFA-enforced identity and access management, and logical or physical separation, because a single compromised credential should never reach both production data and backup copies.

Separation itself comes in degrees. Air-gapped storage is the most stringent version, physically disconnected and about as hard to reach as backup infrastructure gets. Immutable cloud vaults offer a software-enforced version of the same principle, and the right choice for a given client depends on matching the control to that client's risk profile and whatever their insurer is asking for. Vendors have built toward this from different angles: Kaseya Datto SIRIS pairs immutable backups with machine-learning ransomware detection, and Acronis Cyber Protect Cloud adds anti-malware protection and backup scanning on top of its immutable storage.

All of this sits on top of an older discipline that hasn't gone anywhere. The 3-2-1 rule, three copies of data, on two different media types, with one copy kept offsite, remains the baseline architecture, and a hybrid setup that combines local and cloud backup within a single job is how that rule gets implemented for most SMB clients in practice. A complete architecture stacks a local appliance for fast restore, a replicated cloud vault for off-site protection, and an immutable archive as the last line against a ransomware actor who's already gotten further than expected. Each layer is answering a different way the recovery process can fail.

Workload coverage: where gaps in platform support create unprotected client exposure

A backup platform can cover servers and endpoints well and still leave SaaS workloads, identity systems, or specific hypervisors exposed, and that gap tends to stay invisible until an incident or an insurer's questionnaire forces it into the open. The baseline hasn't changed much: Windows Server, Linux, VMware ESXi, Hyper-V, SQL Server, Exchange, and Active Directory still make up the core physical and virtual workloads any serious platform needs to cover.

What has changed is where the pressure is coming from. SaaS coverage used to be treated as an optional add-on. It's now something underwriters and clients ask about directly, specifically Microsoft 365 and identity backup, not just whether the servers are covered. That expectation collides with a genuine trade-off rather than a settled answer. As of March 2026, Microsoft 365 Backup covers Exchange Online, SharePoint, and OneDrive with point-in-time restore built in, but it doesn't address separation from the Microsoft tenant itself, doesn't offer longer retention, doesn't support delegated partner workflows, and doesn't give an MSP a recovery process it can run consistently across dozens of client tenants. The counter-argument, cost consolidation via Microsoft 365 Backup, may allow MSPs to reduce or replace part of their third-party M365 backup spend, though the decision depends on what the client's contract and insurer actually require.

SaaS coverage has become an insurer and client expectation: underwriters and clients increasingly ask about Microsoft 365 and identity backup specifically, alongside server backup. Salesforce, Entra ID, and Jira are now among the least protected, highest-consequence systems inside most organizations, even as the SaaS backup discussion keeps circling back to M365. On the vendor side, one MSP-focused M365 backup product Expert Insights tested includes built-in malware scanning across 14 attachment types, and it runs its own data center infrastructure with regional offices rather than depending solely on cloud storage from the platform provider. Acronis Cyber Protect Cloud supports physical, virtual, cloud, and SaaS workloads from one console, cutting down on tool sprawl for MSPs juggling mixed client environments. NovaBACKUP goes narrower on purpose, focusing on Windows Server, endpoints, SQL, and Hyper-V for MSPs whose SMB clients live almost entirely on that stack. Whichever direction a platform leans, the workload coverage matrix needs to be checked against a representative client environment before it makes a shortlist. Gaps found late in a sales cycle are far harder to work around than gaps caught during the first pass of evaluation.

Storage flexibility and pricing transparency: the hidden criteria that determine margin predictability

Pricing models that charge per restore, per GB of egress, or per compute hour during a DR transfer push cost uncertainty into a live incident and erode the margins that make backup-as-a-service viable. That kind of pricing looks fine on a normal month. It becomes a problem the moment a client actually needs the recovery the MSP has been billing them for, when the bill for pulling data back out suddenly balloons past what anyone budgeted.

Egress fees are the hidden cost that comes up most often in MSP backup pricing conversations. They're structurally suited to catching an MSP off guard, because the fee is invisible during onboarding, invisible during normal operation, and only becomes visible the first time a large restore actually runs, which is usually the worst possible moment to discover it. A platform's storage pricing needs to be pressure-tested the same way its RTO claims do: not against the pricing page, but against what a full-scale, worst-case restore for a real client would actually cost to execute. Margins built on backup-as-a-service don't survive a pricing model that punishes the exact moment the service is supposed to prove its worth.

Sources

  1. 7 Things MSPs Should Look for in Managed Backup | NovaBACKUP
  2. Best Disaster Recovery (DR) Software Solutions (2026)
  3. Best server backup solutions for MSPs in 2026
  4. MSP Backup Solutions Compared (2026) - Flamingo
  5. Finding the Best MSP Backup Solution for Your Small Business Clients
  6. Meeting disaster recovery SLAs with integrated BCDR | Datto

More in Tooling & Vendors