{{img:hero}}When you’re comparing two UX options, the hard part usually isn’t creativity—it’s making a decision you can explain, repeat, and refine next time.
This guide gives you a reusable scorecard workflow you can run on Windows in under an hour, whether you’re comparing two onboarding flows, two navigation structures, or two versions of the same screen.
It’s meant to reduce debate heat and increase decision clarity.
The workflow, at a glance (what you’ll do every time)
You’ll run the same sequence no matter what you’re comparing:
- Step 1: Define the two options and the user/job.
- Step 2: Pick a small set of criteria (the scorecard).
- Step 3: Weight the criteria for this decision (not forever).
- Step 4: Score each option with quick evidence notes.
- Step 5: Run a “red flag” check (things a score can hide).
- Step 6: Decide, document, and set a follow-up measurement.
Think of it as “structured taste”: you still use judgment, but you don’t rely on vibes alone.
Step 1: Freeze the comparison (so it doesn’t quietly change mid-discussion)
Before you score anything, write a one-paragraph comparison brief. This prevents the common drift where people start comparing different things than they started with.
Your brief template:
- Option A: (one sentence)
- Option B: (one sentence)
- Primary user: (who is deciding/doing)
- Job to be done: (what they’re trying to accomplish)
- Context: device, time pressure, accessibility needs, environment
- Success moment: “We know it worked when…”
On Windows, it’s easiest to keep this in a single note file alongside screenshots or links so everyone reacts to the same artifacts.
Step 2: Build your UX scorecard (7 criteria that stay useful)
{{img:scorecard}}Use a small, repeatable set of criteria. Seven is enough to cover most UX comparisons without turning the process into homework.
- Clarity: Does the user immediately understand what this is and what to do next?
- Effort: How much work (time, typing, decisions) does it demand?
- Error risk: How likely are misclicks, wrong inputs, or irreversible mistakes?
- Recovery: If something goes wrong, can the user get back on track?
- Accessibility: Does it hold up for keyboard use, contrast, readable sizing, and assistive tech patterns?
- Consistency: Does it match the rest of the product’s patterns and language?
- Trust: Does it feel safe and credible (permissions, privacy cues, confirmations, tone)?
These criteria work for flows, IA/navigation, and even content-heavy screens because they separate “looks good” from “works well.”
Step 3: Weight the criteria for this decision (so “important” actually means important)
{{img:weights}}A scorecard without weights often fails in real teams because every criterion quietly becomes “equally important,” which isn’t how products work.
Pick a total of 100 points and distribute them across criteria based on this specific situation.
- If this is first-time use, increase Clarity and Trust.
- If it’s a high-frequency task, increase Effort and Consistency.
- If mistakes are costly, increase Error risk and Recovery.
- If you have known accessibility requirements, increase Accessibility (and treat failures as blockers).
A practical weighting pattern many teams reuse: give your top 2 criteria 20–25 points each, your next 3 criteria 10–15 each, and keep the rest at 5–10.
Step 4: Score A vs B with “evidence notes” (the part that makes it reusable)
Use a 1–5 scale for each criterion:
- 1: poor / actively confusing
- 3: acceptable / average
- 5: strong / notably better than typical
But the real value is adding a one-line evidence note for each score. Not a paragraph—just enough to explain the rating later.
Evidence note examples:
- Clarity (A=2): Primary action label is abstract; users must read helper text.
- Effort (B=4): One fewer step; auto-fills from earlier selection.
- Error risk (A=3): Two similar options adjacent; needs better separation.
- Recovery (B=2): No clear “back” state after validation error.
On Windows, you can do this quickly in a simple table (even in a plain text doc): criterion → weight → A score + note → B score + note.
Step 5: Run the “scoreboard can lie” checks (red flags and tie-breakers)
{{img:redflags}}Sometimes an option wins on points but loses in reality. Do these checks before you call it.
- Blocker check: If either option fails a must-have (accessibility, compliance, safety), scoring stops. It’s out.
- Hidden complexity check: Did you score the happy path but ignore edge states (empty, loading, error, permissions)?
- Vocabulary check: Are you rewarding “short” labels that are actually vague?
- New-user vs power-user check: Who are you optimizing for in this release?
- Implementation reality check: Is the “winning” option only winning because you assumed perfect execution?
If the total scores are close (say within 5–10%), use a tie-breaker: choose the option that is easier to iterate after launch (instrumentation, modular steps, fewer irreversible commitments).
Step 6: Decide, document, and set one follow-up metric
Finish with a short decision record so you don’t re-litigate the same debate in two weeks.
- Decision: Choose A or B (or “B with these changes”).
- Top 3 reasons: Pull directly from your evidence notes.
- Main risk: What could make this decision wrong?
- One follow-up metric: e.g., completion rate, time-to-first-success, error rate, support contacts.
- Check date: When you’ll review (often 2–4 weeks after shipping).
One sentence is enough for the follow-up metric: “If we don’t see improvement in X by date Y, we revisit.”
Takeaway: a reusable scorecard is a workflow, not a spreadsheet
The scorecard works when it’s paired with a frozen comparison brief, weighted criteria, and short evidence notes.
Run it the same way each time, and you’ll spend less energy persuading—and more energy making the chosen option genuinely better.