Why comparisons go wrong
Three products are shown over two weeks by three presenters of different skill. Each demo follows a different route. By the end, the team remembers the last one best and the most confident presenter most warmly, and the comparison is between three performances.
A scorecard exists to stop that. It does it with three rules: criteria fixed in advance, the same test for everyone, and a record of how you know next to every score.
Write the criteria before you see anything
Once you have seen a product, its strengths start to look like requirements. So write the list first, from the work, not from a feature comparison page.
- List what the software must do. Eight to twelve lines, each a task somebody performs: "receive a partial delivery against a purchase order", not "inventory management".
- Mark the deal-breakers. Two or three lines where a failure ends the evaluation regardless of everything else. These are not weighted. They are pass or fail.
- Weight the rest. Three levels are enough: needed daily, needed sometimes, pleasant to have. Finer weighting gives a false sense of precision.
- Add the lines no demo covers. Scale, permissions, integrations, import, export, setup, support and cost. They come from what a demo hides and they are where products differ most.
- Agree it with the people who will use the product. One short meeting now prevents the argument later.
Give every product the same test
Use one set of scenarios and one set of invented records for all of them. In a sandbox, that means running the same checklist. On a call, it means sending the same scenarios to each vendor in advance. If one vendor can only offer a call and another has a sandbox, the comparison is not level, and the scorecard should say so in the evidence column instead of hiding it in the number.
Score the evidence, not the impression
Each line gets two entries: a result and a source. The source is the part that makes the scorecard worth keeping.
| Source | What it means | How far to trust it |
|---|---|---|
| Did it | You performed the task yourself in a sandbox or trial | Settled, for the data you used |
| Watched it | A presenter did it live, on a scenario you supplied | Strong, for that path |
| Watched theirs | A presenter did it on their own scenario | Shows it exists; says little about your case |
| Told | Somebody said it can, with nothing on screen | A claim. Ask for it in writing |
| Documented | You read it in public documentation | Good for exports, APIs and limits |
| Unknown | Not shown, not asked, or no clear answer | An open question, not a zero |
For the result, three values are enough: does it, does it with a workaround, does not. A workaround needs one line saying what the workaround is, because "possible with configuration" covers everything from a checkbox to a month of someone's time.
A worked row
Here is one criterion scored for two unnamed products.
| Criterion | Product A | Product B |
|---|---|---|
| Receive a partial delivery (daily) | Does it. Did it: received 60 of 100, order stayed open with 40 outstanding, history showed my receipt. | Workaround. Watched it: presenter closed the order and raised a new one for the balance. Told that a "split receipt" option exists; not shown. |
Without the source, both cells would read "yes". With it, the difference is obvious, and so is the next step: ask the second vendor to show the option they mentioned.
Keep unknowns as unknowns
The commonest way to spoil a scorecard is to fill every cell. An unknown scored as a zero punishes a product for a question nobody asked; an unknown scored as a pass rewards it. Leave the cell marked unknown and count them. A product with many unknowns has not been evaluated yet, and the count tells you how much work is left before a decision is safe.
Unknowns on deal-breaker lines have to be closed before anything else: with a follow-up call, a reference customer, or a proof of concept if the question is large enough.
Deciding
- Remove any product that failed a deal-breaker. That was the purpose of marking them.
- Compare the daily lines first. A product that is excellent at something done twice a year and clumsy at the daily task will be resented.
- Read the workarounds as a list. Each one is a small permanent tax on somebody's day.
- Do not add the numbers up and pick the largest. A total hides which lines produced it. Use the sheet to see where the products differ, then decide on those lines in the open.
The decision record
When you choose, add half a page to the scorecard: what was chosen, the two or three lines that decided it, what was still unknown at the time, and who agreed. This is not bureaucracy. It is what lets you answer "why did we buy this?" a year later, and it makes the renewal conversation much shorter, because the things you were told and never saw are still written down, each with a name beside it.
If words like tenant, seed data or pilot are being used loosely in your team, the glossary gives each a plain definition, and the formats guide explains which kind of demo can supply which kind of evidence.