How to Build an MSP Scorecard for Your Business

A checklist tells you what to ask a managed IT services provider. A scorecard tells you which one to sign. This is the weighted rubric I use with clients. All 12 criteria, the weight on each, what a 5 out of 5 actually looks like, the 4 gate items that override the total, and the score below which you should walk.

An MSP scorecard is a weighted rubric that scores every managed IT provider on the same criteria, using the same 0 to 5 scale, so the decision comes out as a number instead of a feeling. Weights are set before any provider is scored, and they are published.

Three Texas business leaders scoring managed IT providers against a weighted MSP scorecard at a boardroom table

I sit in a lot of these meetings. Somebody has 3 proposals on the table, everyone in the room agrees the middle one felt best, and nobody can say why in a sentence the CFO will accept. So the decision drifts. It always drifts. It goes to whoever presented last, or whoever came in cheapest, or whoever the office manager already knows, and none of those 3 people have any bearing on whether the business will still be well supported in year 3 of the contract.

A scorecard fixes that, and not because numbers are magic. It fixes it because the weights force an argument you would otherwise have after signing. How much does response time actually matter to us compared to price? Nobody knows. Not until you make them put a percentage on it.

This is the version I hand to clients. It is not a generic procurement template with quality and delivery and cost as the three columns. It is built for buying managed IT specifically. I will defend every weight in it, including the ones that annoy people who think price should count for more.

What is an MSP scorecard?

An MSP scorecard is a scoring instrument with 3 parts. A fixed list of criteria, a weight on each criterion that sums to 100%, and a 0 to 5 rating scale with written definitions so 2 people scoring the same evidence land in the same place. The output is one weighted average per provider.

The word doing the work there is evidence. You score what a provider showed you, not what they said. A provider who says backups are tested monthly scores a 2. A provider who hands you the restore log with a date on it scores a 5. Same claim. Different evidence. Wildly different score.

This is where a scorecard and a checklist part company. Our managed IT services checklist and the longer MSP evaluation checklist both tell you what to look for, item by item, pass or fail. Both are useful. But neither tells you what to do when provider A passes 8 of 10 items and provider B passes 7 completely different ones, which is exactly the situation you land in once the shortlist is down to firms that are all broadly credible. Weighting is the answer. The only answer.

What are the 12 criteria, and how much does each one weigh?

Twelve criteria across 4 domains. Not 30. Procurement research generally lands on 15 to 25 criteria for complex buys, and past roughly 7 rating points evaluators start getting sloppy, which is why the scale here stops at 5. Twelve separates real providers. It is also few enough that a 3-person panel will actually finish scoring, rather than abandoning the exercise halfway and reverting to the conversation they were having before somebody handed them a spreadsheet.

A vCIO writing MSP scorecard categories and percentage weights on a glass wall
CriterionDomainWeight
1. Security stack, and who is watching it at 2amSecurity and risk10%
2. Backup coverage and the date of the last tested restoreSecurity and risk8%
3. Their own security posture and independent proof of itSecurity and risk7%
4. Compliance capability for the framework you answer toSecurity and risk5%
5. Response and resolution targets with remedies attachedService delivery9%
6. Escalation path and first contact resolutionService delivery6%
7. Named team and staffing depth behind that teamService delivery5%
8. A written onboarding plan with dates on itService delivery5%
9. A named vCIO and a business review that actually happensStrategy7%
10. Budget ownership and a 3-year roadmapStrategy6%
11. All-in cost transparencyCommercial and exit20%
12. Contract flexibility, exit terms and data ownershipCommercial and exit12%

Roll those up and the domains land at security and risk 30%, service delivery 25%, commercial and exit 32%, strategy 13%. Copy the table into a spreadsheet, put a provider in each column, and you have the whole instrument. No software to buy. No consultant required.

Why these weights and not the ones everyone else publishes

Three of these numbers get argued with every time, so here is the reasoning behind each.

Price sits at 20%, and that is deliberate

The default in most vendor scorecards is to push price toward 40%, partly because it is the easiest thing to score. Type in a number, compare, done. Everything else needs judgment, and judgment is slow. So price quietly expands to fill the space that judgment is supposed to occupy, and the process reliably selects the cheapest bidder while calling itself a process.

We publish 20% to 25% for price in our MSP RFP guidance, and this rubric holds the line at 20. Add the 12% on exit terms and the commercial domain still carries 32% of the decision, which is more than security. Money matters here. A lot. It just does not get to run the room.

One thing to be careful of on that row. Score the all-in number, not the headline. Across the Texas engagements we have priced, first-year total cost lands somewhere between 1.4 and 1.6 times the quoted per-user rate once onboarding, project work, license pass-through and after-hours coverage are counted. Two providers 15% apart on the headline can be dead level on the real number. Sometimes closer. Our Texas MSP pricing benchmarks give you the current ranges to sanity check against.

Security carries 30% because the downside is not symmetrical

A provider who is 20% slower than the alternative costs you some irritation and some lost hours. A provider whose monitoring gap lets an intrusion sit undetected costs something else entirely. IBM put the average time to identify and contain a breach at 247 days in its 2026 report, reversing 5 straight years of improvement, and the global average cost at $4.99 million. You will not see numbers like that. Not at 55 users. You will see a fraction of them, and a fraction is still the worst quarter your business has ever had.

Weighting is really just a statement about which mistakes you can survive. Slow support is recoverable. Annoying, expensive, recoverable. A 6-month intrusion generally is not.

Strategy only gets 13%, and people find that low

I run vCIO work for a living, so you would expect me to weight it higher. I do not, for an unglamorous reason. Strategy is the easiest thing on this list for a provider to promise convincingly and the hardest thing for you to verify before you sign. Everybody has a slide about roadmaps. So it earns 13%, it gets verified through references rather than proposals, and it moves up in the annual re-score once you have watched them either run 4 quarterly business reviews in a row or quietly cancel 3 of them.

How to score each criterion from 0 to 5

A rating scale without written definitions is a mood ring. Two evaluators will use the same numbers to mean different things, and the averaging hides it from both of them right up until somebody senior asks how a provider with no dated restore test ended up with a 4 on backups. Write the anchors down first. Before anyone scores.

The general shape is simple. 0 means they could not answer. 1 or 2 means they asserted it without evidence. 3 means it meets the market standard. 4 means documented and better than standard. 5 means they handed you proof you did not have to ask for.

Four of the 12 get scored badly more often than the rest, so those get full anchors.

CriterionA score of 1A score of 3A score of 5
Backup and tested restoreBackups run. Nobody has restored from themRestores are tested on a schedule they can describeThey hand you the last restore log with a date inside 90 days and the recovery time it took
Response and resolutionOne blended response time for everythingResponse times defined by priority tierResponse and resolution both defined per tier, with service credits that cost them money when missed
All-in cost transparencyA per-user rate and little elsePer-user rate plus a written list of what falls outside itA first-year total including onboarding, projects, licenses and after-hours, with the assumptions shown
Exit and data ownershipNotice period onlyNotice period plus a general commitment to hand data backNamed export formats, a deadline in days, offboarding cost stated up front, and documentation you own throughout

Score independently, then meet. This matters more than it sounds. The moment your loudest evaluator says a number out loud, everyone else drifts toward it, and procurement guidance is consistent on the point that weights and scores get locked before proposals are compared precisely because it is so easy to rationalise afterward. Three scorers is plenty for a company under 100 users. Keep it odd. Keep it small.

Which failures should disqualify a provider outright?

Some failures should not be averageable. A provider can post an excellent total and still be disqualified, because a weighted average will happily let a great price cancel out a missing control, and that is exactly the trade you must not make.

These 4 are pass or fail. No partial credit. Fail any one and the provider is out, whatever the number says.

  1. Multi-factor authentication is not enforced on their own admin and remote management accounts. They hold privileged access into your entire environment. CISA is unambiguous that MFA belongs on every account that matters, and a provider without it on their own tooling is a shared risk you are buying.
  2. They cannot date a restore test inside the last 90 days. Untested backups are a hypothesis. The federal StopRansomware guidance treats regular restore testing as baseline, not advanced.
  3. They will not name the engineers who will hold your account. Not the sales engineer. The people who will answer at 7am. A provider who will not name them is telling you the answer changes weekly.
  4. There is no written exit clause naming a data return format and a deadline. Everything you build with them lives in their tooling. If getting it back is undefined, the contract has a trapdoor in it.

You can add your own. A healthcare client added a signed business associate agreement as a fifth gate, which is correct and specific to them. Just keep the list short. Every gate you add is a veto, and a scorecard with 9 vetoes is a checklist wearing a costume.

A worked example. The cheaper provider loses by 1.2 points

Two real-shaped bids for a 55-user professional services firm. Provider A quoted $138 per user per month. Provider B quoted $171. That gap is roughly $22,000 a year. Not nothing at this size.

Two managed IT provider proposals laid side by side with scores written in the margins
CriterionWeightA rawB rawA weightedB weighted
1. Security stack and monitoring10%353050
2. Backup and tested restore8%251640
3. Their own posture and proof7%241428
4. Compliance capability5%341520
5. Response and resolution9%342736
6. Escalation and first contact resolution6%341824
7. Named team and depth5%241020
8. Written onboarding plan5%251025
9. Named vCIO and business reviews7%15735
10. Budget and roadmap6%14624
11. All-in cost transparency20%5310060
12. Exit terms and data ownership12%343648
Weighted total100%2.894.10

Provider A won the single heaviest criterion on the card outright, 5 to 3, and still lost by 1.21. That is the entire argument for weighting compressed into one row. The cheapest bid did not lose because price was undervalued. It lost because it was thin everywhere else and one 20% criterion cannot carry a proposal.

Look at row 2 as well. Provider A scored a 2 there, which means no dated restore test, which means A fails a gate item and was disqualified before the total was even computed. We finished scoring anyway. Always do. The number is what you take to the board, not the disqualification.

And B is not a clean win either. B scored a 3 on cost transparency, and that gap is precisely what you negotiate before signing. Which brings us to what the totals mean.

What the total score means, including when to walk away

A weighted total is a 0 to 5 number. Here are the bands I use, and the actions attached to them.

Weighted totalWhat it meansWhat to do
4.2 and aboveStrong across every domain, with documented proof rather than assertionsSign. Uncommon enough that it is worth checking you scored honestly rather than enthusiastically
3.5 to 4.1Solid, with gaps that are known and namedSign only if every gap sits in something you can write into the contract. If it does not, it is not a gap, it is a limit
2.8 to 3.4Meets the market standard and not much moreDo not sign yet. Go back with the 3 lowest-weighted-score rows as specific asks and re-score
Below 2.8Structurally short on things that are hard to fix after signingWalk away
Any gate failureDisqualified regardless of totalWalk away

Two honest cautions about that table. First, 2.8 is a threshold I can defend, not a law of nature. If your whole shortlist scores in the 2s, the shortlist is wrong, not the number, and the fix is finding 2 more providers rather than lowering the bar. That happens most often when the shortlist was built entirely on price and postcode, which tends to produce 4 firms that are all the wrong size for the business and all conveniently available at short notice for a reason nobody has asked about yet.

Second, the total is a comparison tool, not a rating of the company. A provider can score 3.1 against a 55-user firm with compliance obligations and 4.4 against a 20-user firm that just wants tickets answered fast. Same provider. Same week. The scorecard measures fit for you, which is the only thing you actually need it to measure.

Regional context helps too. What a competent bid looks like in Dallas is not identical to San Antonio, and our DFW provider guide and San Antonio buyers guide both cover what a market-standard 3 looks like locally.

How should you adjust the weights for your situation?

The 12 criteria stay. Always. The weights move once, before you score anybody, and here is where I actually change them.

Your situationMove upMove downWhy
Regulated data (HIPAA, GLBA, CMMC, PCI)Compliance capability to 10%Price to 15%A provider who cannot evidence your framework is not cheaper, they are a finding waiting to be written up
You already have internal ITEscalation and named team to 9%Onboarding to 3%You are buying a second-tier bench, not a full transition. Coverage boundaries decide whether this works
Multiple sites or heavy field staffResponse and resolution to 12%Strategy to 9%Travel and remote-hands reality decides the experience more than roadmap quality does
You are switching after a bad experienceExit terms to 16%Price to 16%You already learned what a weak exit clause costs. Do not pay that tuition twice
Cyber insurance renewal is driving thisTheir own posture and proof to 12%Budget and roadmap to 3%Carriers now ask about MFA, endpoint detection and tested backups directly, and a misstated control is how claims get denied

If your internal IT question is still open, the weights are the second decision. In-house versus managed versus co-managed is the first one, and picking the model before the provider will save you re-scoring the whole shortlist.

One rule holds regardless. Change the weights before proposals arrive, then leave them alone. Adjusting weights after the bids arrive is how a scoring process quietly turns into a justification exercise, and everyone in the room knows it happened even when nobody is willing to be the person who says it out loud.

Where the scorecard does not help you

It cannot score whether you will like working with them. That sounds soft. It is not. Not remotely. Over a 3-year contract the difference between a provider who tells you bad news early and one who manages you is worth more than several of these criteria combined, and there is no honest way to put a number on it from a proposal.

Two people running an MSP reference check call using a scored question sheet

References are where you get at it, and only if you ask the right thing. Do not ask whether they are happy. Ask about the last time something went badly wrong and how they found out about it. Ask for one former client, not just 3 current ones. That conversation is consistently the most useful 20 minutes in the whole evaluation, and almost nobody asks for it, because asking feels adversarial in a way that a reference call to a delighted customer never does.

It also cannot score anything you have not made them show you. A provider scored entirely from a written proposal is a provider scored on their marketing team. Do the technical session. Read the actual master services agreement, because SOC 2 attestations and framework references sound identical in a slide deck and differ enormously in a contract. Our breakdown of MSP contract terms worth checking covers the clauses that move criterion 12 by 2 whole points.

And it will not tell you anything useful if you score it in one sitting from memory. Score each provider within a day of meeting them, while the evasive answer is still fresh. It fades fast. And it fades in their favour.

Score your current provider too, not just the new ones

The most useful thing this instrument does has nothing to do with buying. Run it on your incumbent once a year, in the same month, with the same 3 people. You will get a trend line, and a trend line is a far better signal than any single total.

Most relationships do not fail suddenly. They drop half a point a year, which is invisible month to month and obvious across 3 annual scores. The re-score also gives you something concrete for the renewal conversation. Not a complaint. A number. With the 3 rows that moved.

If the score has fallen below 2.8 twice in a row, you already know what the scorecard is telling you. Switching without downtime is a solvable problem, and it is a much smaller problem than staying somewhere that is quietly getting worse.

For the record, I am asking you to point this at us as well. We have been supporting Texas businesses since 1999, our team is 42 people across Houston, San Antonio and Dallas, we triage in under 10 minutes and we back the first 120 days with a satisfaction guarantee you can exit on. Those are criteria 5, 7 and 12. Score them like everybody else. If we come in under 4, ask us why, and we should have an answer better than a brochure.

Want your weights reviewed before you score anyone?

Send me your criteria and weights and I will tell you which rows your shortlist is going to game, and which gate items your situation needs that this list does not include. No pitch attached, and you are welcome to use the result on our competitors. Half the point of publishing weights is that they survive contact with people who did not write them.

Send us your scorecard draft and we will mark it up. We support businesses across Houston, Dallas-Fort Worth and San Antonio.

What Texas buyers ask about scoring providers

How many people should score an MSP?

Three for a company under 100 users. The budget owner, whoever lives with the tickets day to day, and somebody from the part of the business that hurts most during an outage.

Keep it odd-numbered and keep it small. Panels of 7 do not produce better decisions, they produce averaged ones, and averaging is how the blandest safe proposal wins. Everyone scores alone first, then you meet and argue about the rows where you disagree by 2 or more. Those arguments are the actual value of the exercise.

Should we send the scorecard to the providers?

Send the criteria and the weights. Keep your scores private until the decision is made.

Publishing the weights makes proposals better, because bidders write to what you told them you care about instead of guessing. It also removes your own ability to move the goalposts later, which is a discipline worth imposing on yourself. Publishing the live scores is different. That turns the evaluation into a negotiation and you will spend a week defending a 3.

What if a provider refuses to give evidence for a criterion?

Score it a 1 and tell them you did. The response to that is more informative than the original answer.

Some refusals are legitimate. A provider will not send you another client’s SOC 2 report, and they should not. But a redacted summary, an attestation letter, or a screen share of the restore log all exist as options. A firm that offers no alternative at all is telling you something. Write down what they offered instead, because that is what you are actually scoring.

Can we use the same scorecard for a co-managed arrangement?

Yes, with 2 weight changes. Escalation and named team goes up, onboarding goes down.

Co-managed is a different purchase. You are buying bench depth and specialist coverage rather than a full handover, so the criteria that decide success are the boundary ones. Who owns which ticket, how work gets handed to your team and back, and what happens when your own engineer is out for 2 weeks. The 12 criteria still apply. The centre of gravity moves.

How long does scoring a shortlist take?

Budget 3 to 4 hours per scorer across the whole shortlist, plus a 90-minute session together.

That assumes 4 providers and that you have already had the technical sessions. If it is taking dramatically longer, you have too many criteria. If it is taking 40 minutes, you are scoring the proposals rather than the evidence and the total will not survive its first challenge from the CFO.

Does a high score mean a provider is good?

It means they are a good fit for you, on the things you weighted, with what they showed you. Those 3 qualifications are the whole answer.

A provider who scores 4.4 for a 20-person firm that wants fast ticket resolution might score 3.0 for a manufacturer with CMMC obligations, and neither number is wrong. This is also why the annual re-score matters more than the buying-day score. Buying-day is a prediction. The re-score is a measurement.

About Author

Learn More