8 Questions to Ask MSPs About Backup and Recovery

Most backup conversations with an MSP end in the word yes. Yes we back you up. Yes it is offsite, yes it is encrypted. None of that tells you whether your business comes back. These 8 questions sit behind every disaster recovery plan we inherit, and each one has an answer that is either a number, a document, or a shrug, and the shrug is the one worth listening for.

The 8 questions worth asking an MSP about backup cover recovery time, recovery point, retention, immutability, restore testing, Microsoft 365 coverage, runbook ownership and exit terms. Each one has a checkable answer. Vague answers are the finding.

Here is the pattern we see. A buyer asks whether the provider does backups. The provider says yes. Box ticked. Everyone moves on to price, and 2 years later something fails and the room finds out what that yes actually covered, usually at 6am, usually on the worst possible day.

Backup is the easiest thing in managed IT to sell and the hardest thing to verify. An agent installs. A dashboard turns green. The monthly report says 100% success, and every bit of that can be true while your company still sits 4 days away from being usable again after one bad Tuesday morning. Success in a backup report means bytes got copied. It does not mean anything got recovered. Those are different claims.

So here are the questions our engineers ask when they take over somebody else’s stack. There are 8. None of them are trick questions, and no decent provider will mind being asked. They just force a number, a document, or an honest admission. That is the point. If you are running a broader evaluation, our MSP evaluation checklist covers the rest of the vetting. This is the backup half, in detail.

What should you ask an MSP about backup and recovery?

Asking an MSP about backup means testing 3 things. Whether the design survives an attacker who already has valid credentials, whether the recovery speed matches what your business can absorb, and whether anyone has proved either of those recently. Every question below is a version of one of those 3.

Ask them in order. Order matters. The early ones set up the later ones, and by question 5 you will know whether you are talking to somebody who runs restores for a living or somebody who watches a dashboard.

  1. Is that recovery time objective per system, or for our whole environment?
  2. What is our real recovery point objective on the worst night, not the best one?
  3. How long do you keep restore points, and how does that compare with how long an intrusion goes unnoticed?
  4. Which copy can you not delete, even with your own admin credentials?
  5. Can we see the last restore test report?
  6. Who is backing up Microsoft 365, and what happens to a mailbox deleted 100 days ago?
  7. What happens if the one person who knows our environment is unreachable?
  8. When we leave, how do we get our data back and how fast?

One caveat before you start. We are an MSP, so read this knowing we answer these questions for a living and would rather compete on them than on price. Fair warning. That is also why the list is specific enough to use against us.

1. Is that recovery time objective per system, or for our whole environment?

Recovery time objective, or RTO, is how long you are willing to be down. Recovery point objective, or RPO, is how much data you are willing to lose. Providers quote both. Almost nobody says which unit the number applies to. Ask, and watch the pause.

A 4-hour RTO sounds precise. Then you ask what it counts, and it turns out to be per workload, measured from the moment an engineer starts the restore, on a good day, with target hardware already sitting there. Your business does not run on 1 workload. It runs on a domain controller, a file server, a line-of-business database, and whatever the accounting team refuses to move off.

What you are restoringOne at a timeThree in parallelWhat actually decides it
3 workloads at 4 hours each12 hours4 hoursAppliance read speed and how many engineers are free
6 workloads at 4 hours each24 hours8 hoursThe same, plus whether restore order is documented
12 workloads at 4 hours each48 hours16 hoursWhether replacement hardware or cloud capacity exists yet

The arithmetic is not complicated, and that is the point. Nobody publishes it. A per-workload RTO multiplied by your workload count is your real worst case, and parallelism is the only thing that pulls it back down. Order matters just as much. A file server restored ahead of the domain controller is just a pile of files that nobody in the building can authenticate to, so ask for the full sequence in writing before you sign anything. If the line between staying open and getting systems back is fuzzy, our breakdown of business continuity versus disaster recovery separates the two. And once the answers are in, our ranking of seven Texas BDR providers scored on these same questions shows where each one lands.

Ask this. How many restores can you run at the same time, in what order do you run them, and what is the total elapsed time for my full environment rather than for one server?

2. What is our real recovery point objective on the worst night, not the best one?

A nightly backup gives you an RPO of up to 24 hours. Not 24 hours on average. Up to. If the failure lands at 4:30pm and the last good copy was written at 11pm the night before, you have lost a full working day of everything that anyone in the company typed. Spread failures evenly across the cycle and the expected loss is about half the interval, which works out at roughly 12 hours on a nightly job and 2 hours on a 4-hour job.

Then there is the failure nobody plans for, which is the backup job itself failing quietly. Ask what happens when Tuesday’s job errors out and nobody looks until Thursday. Your RPO is no longer 24 hours. It is 72. The number that protects you is not the schedule. It is who gets paged on a failed job, and how fast.

Retention pruning does the same damage from the other direction. A provider taking snapshots every 15 minutes has not given you a 15-minute RPO if those snapshots get pruned after 24 hours and the corruption is not spotted until the third day. Frequency and retention are 2 separate promises. Get both in writing.

Ask this. What is my worst-case data loss if a job fails silently on a Friday, who gets alerted, and how long do the frequent restore points survive before they are pruned?

3. How long do you keep restore points, and how does that compare with how long an intrusion goes unnoticed?

This is the question that changes the answer to every other one. Retention gets sold as a compliance number. It is not. It is a detection number, and your restore history has to reach back further than the intrusion you have not found yet, or the clean copy you actually need is already gone.

Mandiant’s M-Trends 2025 report puts global median dwell time at 11 days, up from 10 the year before, drawn from more than 450,000 hours of investigations across 2024. The median moves a lot depending on who tells you.

How the compromise gets foundMedian days before you knowRetention you need for a clean copy
Your own team detects it10More than 10 days of history
An outside party notifies you26More than 26 days
The attacker announces it5More than 5 days

Read the middle row again. When somebody else is the one who tells you, half of those cases are already past 26 days. A 14-day retention window covers 2 of the 3 rows and fails the one you have least control over. That is the argument for 30 days as a floor, and 90 or more for anything that changes slowly, like finance records and design files. Median also means half of all cases are worse than the number. Plan for the worse half.

Ask this. How many days of restore points do I have for each system, and can you restore me a file from 45 days ago today?

4. Which copy can you not delete, even with your own admin credentials?

Attackers stopped ignoring backups years ago. Sophos research across 2,974 IT and security leaders found that 94% of ransomware victims had attackers try to compromise their backups, and 57% of those attempts worked. More than half.

The gap between those 2 outcomes is the entire business case. Where backups were compromised, the median ransom demand was 2.3 million dollars against 1 million where they were not. 67% of those victims paid, against 36%. Median recovery cost came out 8 times higher, at 3 million dollars against 375,000. And only 26% were fully recovered inside a week, compared with 46% of the organizations whose backups held. Same attack. Different bill.

So the question is not whether backups exist. It is whether any copy sits out of reach of a valid administrator login. The CISA StopRansomware Guide is blunt about it, telling organizations to maintain offline, encrypted backups of critical data, and to back up often, either offline or by leveraging cloud-to-cloud backups. Worth noting that the same guide warns to use immutable storage with caution, because it does not meet compliance criteria for certain regulations and a misconfiguration can impose significant cost. Immutability is not a free win. It is a design decision with a bill attached.

Here is the version of this question that is specific to hiring an MSP, and it is the one we almost never get asked. The credentials that reach your backups might not be yours. In June 2025 CISA published an advisory on ransomware actors exploiting an unpatched flaw in SimpleHelp remote monitoring software to reach a provider’s downstream customers. A provider’s console is a single door into every client behind it. So ask which of your copies survives a compromise of your MSP, not just a compromise of you. If the answer is none, you do not have an offsite copy. You have a second copy inside the same trust boundary. Different building. Same keys. Our walkthrough of what the first 72 hours of a ransomware recovery decide shows what that looks like when it goes wrong.

Two engineers checking which backup copy is locked against deletion during an immutable storage review

Ask this. If your own management console were compromised tomorrow, which copy of my data still exists, and who holds credentials that could delete it?

5. Can we see the last restore test report?

Everything above is a design claim. This one checks whether the design was ever exercised. The Veeam 2025 Ransomware Trends report found that 98% of organizations had a ransomware playbook on file, while fewer than half of them, 44%, actually included backup verification and frequency inside it. The plan exists. The proof usually does not.

There is also a coverage problem hiding inside a reassuring answer. A provider who tests 1 randomly chosen workload every month sounds diligent. If you run 24 protected workloads, that schedule needs 24 months to touch each one once. After a full year you have validated half your environment, and you do not get to pick which half. Nobody does.

A real restore test report names the workload, the date, the restore target, the elapsed time, who ran it, and what failed. That last column is the one that tells you it is genuine. Tests that never fail are not tests. They are screenshots.

If you handle protected health information, this is not optional in the way people assume. Under 45 CFR 164.308(a)(7), the Data Backup Plan and the Disaster Recovery Plan are both Required implementation specifications. Testing and revision procedures are Addressable, which gets widely misread as optional. Addressable means you implement it, or you document why an alternative is reasonable and appropriate for your organization. Nobody has ever written that memo successfully to justify skipping restore tests. Not once.

Engineer timing a restore from a backup appliance and recording elapsed time for a restore test report

Ask this. Send me the last 3 restore test reports with dates, elapsed times and anything that failed. Then tell me how many of my workloads were tested in the last 12 months.

6. Who is backing up Microsoft 365, and what happens to a mailbox deleted 100 days ago?

Microsoft runs the platform. You own the data inside it. That split is the shared responsibility model, and it catches out more businesses than any other item on this list, because replication across data centers looks like a backup and is not one. A file deleted in a replicated system gets deleted everywhere, quickly and correctly. Correctly is the problem.

Microsoft’s own data deletion documentation puts the SharePoint and OneDrive recycle bins at 93 days total before content becomes unrecoverable. Not 93 days each. 93 combined. Ask the 100-day question specifically, because it separates providers who have read the retention documentation from providers who assume the cloud handles it. We laid out the full retention table, row by row, on our data backup services page.

Then ask which workloads are actually covered. Exchange Online, SharePoint, OneDrive and Teams are 4 separate scopes, and Teams in particular scatters its content across 3 of the others. A provider who answers that they back up 365, without naming the workloads, has told you nothing. Politely, but nothing.

Ask this. Which Microsoft 365 workloads are backed up, where does that copy physically live, who can delete it, and can you restore a single mailbox item from 100 days ago?

7. What happens if the one person who knows our environment is unreachable?

Recovery is a people problem wearing a technology costume. The design can be perfect and still stall for 6 hours, because the only engineer who knows how your ERP database is laid out happens to be sitting on a plane.

Ask where the recovery documentation lives. If the answer is the file server, that answer is wrong, because the file server is the thing that is down. Ask whether a break-glass account exists that does not depend on your domain. Ask how they will reach you when email is the encrypted thing. Out-of-band communication sounds like an overreaction right up until the morning you need it. Then it sounds like planning.

This is a bench-depth question too, so ask how many engineers could run your recovery without a handover call. We are a team of 42 across Texas and we have been doing this for more than 25 years, which matters here for exactly one reason. No client’s recovery depends on a single person being reachable. Not one.

Ask this. Name the engineers who could run our recovery without a handover, tell me where the runbook is stored, and show me the out-of-band contact plan.

8. When we leave, how do we get our data back and how fast?

Ask this at the start, not at the end. Nobody negotiates exit terms well while they are leaving, and the answer tells you something useful about a provider’s confidence today. Especially the pause before it.

The trap is portability. Backup history stored as proprietary appliance images is worth very little without the software that reads it, so a provider can hand over your data in full compliance with the contract and still leave you holding files that nobody can open. Get the format agreed. In writing. Ask how many business days it takes, and ask whether restore history transfers, or whether your retention clock restarts at zero on day 1 with the new provider, because that is the usual answer and it is worth knowing before you sign rather than after.

There is a regulatory floor under this as well. Businesses covered by the FTC Safeguards Rule are already required, under 16 CFR 314.4(f), to take reasonable steps to select service providers capable of maintaining appropriate safeguards, to require those safeguards by contract, and to reassess them periodically. Vendor oversight is not a commercial nicety for those firms. It is a written control. Our guide to the MSP RFP process has clause language worth borrowing.

Ask this. On termination, in what format do we get our data, how many business days does it take, does the restore history come with it, and will you certify deletion of your copies?

Two failure patterns we find in almost every stack we inherit

Neither of these is exotic. Both are quiet, both survive years of green dashboards, and we find at least one of them in most environments we take over. Most. Not some.

The backup that copies perfectly and restores into garbage

A database gets backed up while it is running, with no application-aware snapshot to quiesce it first. The copy job succeeds every single night, because copying a file is all it was ever asked to do. What lands in storage is a torn image of a database caught mid-write. It restores. The service starts. Then it fails consistency checks, and nobody finds out until the day they need it, because nobody ever restored it.

The backup account that is a domain admin writing to a share on the same domain

On the architecture diagram there are 3 copies in 2 locations, which reads as textbook. In practice a single stolen credential reaches production and every copy at once, because all of it sits inside one authentication boundary. The storage design is fine. The identity model defeats it.

Both patterns pass a monthly report. Both fail a restore test. That is the whole argument for question 5.

How to score the answers you get

You do not need a weighted matrix for this. You need to notice which answers arrive as numbers and which arrive as adjectives.

QuestionA good answer sounds likeWalk away if you hear
1. Environment RTOA total elapsed time for all workloads, plus a written restore orderRobust, rapid, industry-leading
2. Real RPOA worst-case number, a named alerting path and a retention floorContinuous, real-time
3. Retention against dwell timeA per-system day count, and an offer to restore something from 45 days agoIt depends on the plan
4. ImmutabilityA named copy their own console cannot deleteEverything is encrypted
5. Restore testing3 dated reports, including one where something failedWe monitor them daily
6. Microsoft 365Named workloads, a named backup location and a 100-day restoreMicrosoft handles that
7. Runbook and benchNamed engineers and a documented out-of-band planYour account manager
8. Exit termsA format, a business-day count and a deletion certificateWe will work with you

One more filter. Watch what happens when you ask something they cannot answer on the call. A provider who says they will check and come back with the actual number is showing you exactly how they will behave at 3am on the morning of a real incident. A provider who improvises is showing you that too.

Get a second opinion on the backup you already have

You do not have to switch providers to run these 8 questions. Send us the answers you got and we will tell you which ones are solid, which ones are marketing, and what a realistic recovery time looks like for your environment. If your current provider passes, that is a useful thing to learn for free.

Speak to an IT Expert

Or see how our managed IT services handle backup and recovery day to day. We support businesses across Houston, Dallas-Fort Worth and San Antonio.

Questions buyers ask us after this conversation

How often should an MSP test our backups?

Quarterly for business-critical systems, monthly for a file-level spot check, and once a year for a full environment exercise. Anything less and you are testing during the incident. Which is not a test.

The frequency matters less than the record. A provider who tests twice a year and hands you dated reports with real failures in them is in better shape than one who claims monthly testing and cannot produce a single document. Any significant change to your environment should also trigger an off-schedule test, because a new server that nobody remembered to add to the backup job is the single most common gap we find.

Is cloud-only backup enough for a small business?

Usually not, and the reason is arithmetic rather than opinion. Restoring 10 terabytes over a 100 megabit connection takes days, not hours.

That is why most designs keep a local appliance for speed and an immutable cloud copy for survival. The local copy gets you running. The cloud copy is what is left if the building or the credentials are gone. Either one. If somebody is selling you cloud-only, ask them to compute your full restore time at your actual line speed, then tell you whether your business can sit still that long.

Does our MSP’s SOC 2 report cover our backups?

It covers the provider’s own controls, not your recovery outcomes. A SOC 2 report is evidence about how a company runs its business, not proof that your data comes back.

Read the scope section and the exceptions before you accept it as an answer. Then ask separately for your restore test evidence, because those are 2 different documents answering 2 different questions. Both are worth having. Only one of them is about you.

What is a reasonable RPO for a small business?

For most, 4 hours on transactional systems and 24 hours on file data. Anything that generates revenue by the hour should be lower than that.

Work it backwards instead of picking a number off a menu. Start there. Take one system, ask how many hours of lost work your team could re-key by hand without missing a commitment to a customer, and that is your RPO for that system. It is rarely the same number across the whole environment, and providers who quote a single RPO for everything have not asked the question.

Should the backup vendor be different from the MSP?

Not necessarily different, but the copies should be. What matters is whether one compromised set of credentials can reach production and every backup at the same time.

Separating vendors is one way to get that separation. Immutable storage with a retention lock that the provider’s own console cannot override is another, and it is usually simpler to run. Ask which mechanism they use, then ask them to name the account that could delete your last copy. If nobody can name it, that is your answer. A bad one.

We are fully in Microsoft 365. Do we still need backup?

Yes. Microsoft protects the platform, you protect the data, and the built-in recycle bins run out at 93 days for SharePoint and OneDrive.

Retention policies and legal holds are not backups either, though they get sold as though they are. They are governance tools. They do not give you a point-in-time restore of a mailbox that somebody wiped during a bad week, and they do nothing at all when an integration with delete permissions goes wrong. Assume the deletion will be legitimate, authenticated and irreversible, then decide how much history you want to keep.

About Author

Learn More