· Dave Mathias · Ideas · 28 min read
A Score Is Not a Strategy
How product teams can make better tradeoffs, choose the right prioritization framework, and use AI without outsourcing judgment.

The spreadsheet says Initiative A should win.
Its RICE score is 93.6. Initiative B sits at 81.2. The rows are sorted, the cells are color-coded, and the answer looks objective.
Then someone asks where the impact estimate came from.
Silence.
The number was not discovered. It was negotiated. Reach came from an optimistic forecast. Effort excluded the work required from legal and operations. Confidence was assigned after the team had already fallen in love with the idea. A formula converted those assumptions into a decimal, but it did not make them true.
This is the central mistake in product prioritization: confusing a method for comparing choices with a system for making them.
A prioritization framework should expose judgment, not replace it.
That distinction matters for software, physical goods, and services. Every product team has more possible work than time, money, attention, and organizational patience. The real job is not to produce an impressive ranking. It is to decide where scarce resources can create the most worthwhile change, what will wait, and what will not be done.
AI makes the distinction even more important. It can summarize thousands of customer comments, calculate scores, model scenarios, and produce a polished rationale in seconds. It can also turn weak assumptions into confident-looking recommendations faster than any spreadsheet ever could.
The quality of prioritization still depends on the quality of the choices people make before, during, and after the score.
Prioritization begins before the framework
No framework can answer a question the organization has not framed.
Before comparing ideas, a team needs to settle five things.
1. What outcome are we trying to change?
“Grow the product” is not enough. Neither is “improve customer experience.”
A useful outcome names a customer or business change, a target population, and a relevant period. Reduce the time it takes a new small-business customer to receive the first payment. Increase successful completion of an intake appointment. Reduce warranty claims during the first year of ownership.
Without an outcome, prioritization becomes a contest among features. With one, features become possible ways to create a result.
This is why Teresa Torres argues for prioritizing opportunities rather than solutions.1 A feature list starts too late. The more consequential choice is often which customer problem deserves attention.
2. What level of decision are we making?
Teams often place items of different sizes in the same comparison.
“Enter a new market,” “redesign onboarding,” and “change the button label” do not belong in one RICE sheet. One is a portfolio choice, one is a product initiative, and one is a small solution. They have different evidence, time horizons, and decision owners.
Compare markets with markets, opportunities with opportunities, concepts with concepts, and backlog items with backlog items. If two candidates cannot be described at roughly the same level, the ranking will mislead you.
3. What is not negotiable?
Safety, legal obligations, accessibility, security, contractual commitments, and the minimum integrity of the product should usually be handled as gates or capacity commitments, not forced to compete for points against growth ideas.
This does not make their cost irrelevant. It means the question is different. The team is deciding how to satisfy a constraint wisely, not whether the constraint has enough reach to deserve attention.
4. What evidence counts?
A score is only as credible as the evidence beneath it.
Agree in advance what can support a claim: observed customer behavior, interviews, service data, revenue analysis, technical investigation, market research, experiments, or informed judgment. Then label the difference. A measured reach estimate and an executive hypothesis should not look identical in the final sheet.
5. When will we reconsider?
Prioritization is a temporary decision made with current information. It needs an expiration condition.
That might be a date, a new regulatory ruling, the result of a prototype test, a competitor move, a change in capacity, or evidence that an expected outcome is not appearing. Without a review trigger, yesterday’s assumptions quietly become today’s commitments.
The prioritization stack
A useful framework sits inside a larger decision system. I think of that system as six layers.
- Direction: Name the strategy, target customer, and outcome.
- Eligibility: Remove work that is out of strategy, unsupported, unethical, unsafe, or not ready to compare. Separate true obligations from optional bets.
- Evidence: Assemble customer, business, technical, operational, and risk signals. Record their sources and freshness.
- Comparison: Apply a framework that fits the decision. This is where RICE, Kano, MoSCoW, WSJF, or another method belongs.
- Sequence: Account for dependencies, capacity, learning value, time sensitivity, and the need to maintain a coherent whole product or service.
- Decision: Name the owner, explain any override, state what will not be done, and set the next review trigger.
Most prioritization failures come from asking layer four to perform all six jobs.
RICE cannot supply a strategy. Kano cannot resolve an engineering dependency. MoSCoW cannot tell you whether the underlying customer problem is worth solving. A portfolio matrix cannot tell you what to build next Monday.
The framework is a lens. The prioritization stack is the system.
The five most popular product prioritization frameworks
There is no authoritative global ranking of framework usage. Practitioner guides and research use different samples and definitions. The five below are the most recognizable and repeatedly cited across product practice, and together they cover five distinct questions. Popularity is not proof that any one of them is right for your decision.
1. Impact-Effort Matrix: Where should we begin the conversation?
The Impact-Effort Matrix places each option on two axes: expected impact and required effort. High-impact, low-effort items appear to be quick wins. High-impact, high-effort items are major investments. Low-impact, high-effort items are usually poor candidates.
The Government of Ontario describes it as a way to organize solutions or problem areas by the effort required and the impact on users.2
Best use: An early workshop with a manageable set of opportunities or ideas when the team needs a shared picture before doing deeper analysis.
Why teams like it: It is fast, visual, and easy to discuss. People can see disagreement. The act of placing an item often matters more than the quadrant where it lands.
Where it breaks: Both axes can hide several different concepts. Impact may mean customer value, revenue, strategic fit, risk reduction, or political urgency. Effort may exclude research, rollout, training, manufacturing, service operations, maintenance, and change management. Teams also tend to chase quick wins while postponing the difficult work that actually expresses the strategy. Intercom calls the low-effort, low-impact version of this behavior “snacking.”3
Use it well: Define the axes before placing anything. Add a visible evidence marker to each item. Treat the matrix as a discussion starter, then investigate major bets and suspiciously attractive quick wins.
2. RICE: Which ideas promise the most supported impact for the effort?
Intercom created RICE to compare projects using four inputs: Reach, Impact, Confidence, and Effort.4
RICE score = (Reach × Impact × Confidence) ÷ Effort
In Intercom’s original explanation, reach is measured over a defined period, impact uses a consistent scale, confidence discounts weak estimates, and effort includes the time required from everyone involved.4
Best use: Comparing initiatives of similar size that pursue the same outcome and have enough evidence for rough estimates.
Why teams like it: RICE separates total reach from impact per person. It also makes uncertainty visible instead of letting a bold impact claim pass as fact. The denominator forces some attention onto cost.
Where it breaks: A low-confidence estimate can still be arbitrary. Reach favors broad improvements over narrow but strategically important work. Effort is often reduced to engineering time. Dependencies, trust, regulation, product coherence, and option value are outside the formula. Multiplying ordinal impact ratings can create more precision than the inputs deserve.
Use it well: Use a shared time window and scoring anchors. Record a source beside every input. Estimate total cross-functional effort. Run the scores as ranges, not only point estimates. When a lower-scoring item is selected, write down why. An explicit override is evidence of judgment, not proof the framework failed.
3. MoSCoW: What must fit inside a fixed commitment?
MoSCoW groups requirements into Must Have, Should Have, Could Have, and Won’t Have This Time.5
The distinction that matters is not whether someone strongly wants an item. The Agile Business Consortium suggests asking what happens if a requirement is absent. If the project or release would no longer be viable, legal, or safe, it may be a Must. If a workaround exists, even an inconvenient one, it is probably a Should or Could.5
Best use: Protecting a deadline, defining a minimum usable release, or negotiating scope when time and capacity are constrained.
Why teams like it: The categories are understandable to technical and nontechnical participants. “Won’t Have This Time” gives teams a direct way to set expectations. Used properly, the method creates contingency rather than pretending every planned item is guaranteed.
Where it breaks: Stakeholders quickly learn that calling something a Must is the safest way to protect it. MoSCoW creates categories, not a full ordering within them. It also performs poorly when the release objective is vague or requirements are too large to decompose.
Use it well: Begin with every item outside the commitment and require evidence to move it in. Apply the cancellation test to every proposed Must. The Agile Business Consortium recommends that Musts typically consume no more than 60 percent of the effort, leaving real flexibility in the plan.5 Reapply the labels to the specific release or timebox. A future Must can still be a current Won’t.
4. Kano Model: Which attributes prevent dissatisfaction, improve performance, or create delight?
The Kano Model examines how the presence or absence of an attribute affects customer satisfaction. The American Society for Quality explains three central categories:6
- Must-be needs: Customers expect them. Their absence creates strong dissatisfaction, but their presence rarely earns praise.
- Performance needs: Better performance tends to create more satisfaction.
- Attractive needs: Customers may not expect them. Their presence can delight, while their absence causes little harm.
Kano analysis normally uses paired questions: how a customer feels if an attribute is present and how the customer feels if it is absent.6
Best use: Shaping a product or service concept, balancing basics with differentiators, and understanding why customer satisfaction does not rise evenly with every feature.
Why teams like it: Kano prevents teams from comparing every attribute on one linear value scale. It recognizes that fixing a broken basic is different from adding a surprise. It works across software, manufactured products, and services.
Where it breaks: Classification requires carefully designed research. Results vary by segment and can change over time as yesterday’s delight becomes tomorrow’s expectation. Kano says little about effort, strategic fit, or financial return.
Use it well: Segment the results. Repeat the study as the market changes. Combine Kano with an economic or effort-based method. Do not ask customers to design the solution. Use the research to understand the satisfaction response, then let the product team decide how to address it.
5. WSJF: Which delay-sensitive job should go first?
Weighted Shortest Job First, or WSJF, sequences work according to relative economic urgency.
In the Scaled Agile Framework definition, the numerator combines user and business value, time criticality, and risk reduction or opportunity enablement. That cost of delay is divided by relative job size.7
WSJF = Cost of Delay ÷ Job Size
Best use: Sequencing a set of reasonably understood jobs competing for the same delivery capacity, especially when delay changes their value.
Why teams like it: WSJF recognizes that value has a clock. A smaller item with urgent economic value may deserve to move ahead of a larger item with a bigger total payoff. It is often more useful for sequence than a static return-on-investment ranking.
Where it breaks: Relative numbers can be gamed. Teams may count the same benefit under more than one numerator component. Job size is not always duration. The method can undervalue foundational work whose benefits arrive through many later initiatives.
Use it well: Score all candidates in one session with stable anchors. Compare like-sized decision units. Check whether risk reduction and opportunity enablement are distinct from business value. Recalculate when timing changes. Keep a separate view of dependencies and platform investments.
The comprehensive field guide
No finite list can contain every local scorecard, branded variation, or hybrid. The table below covers the established families and commonly named methods that repeatedly appear in product management, requirements engineering, market research, service improvement, engineering design, and innovation portfolio practice. Closely related aliases are combined rather than presented as separate inventions.
| Framework or method | Best for | How it works | Pros | Cons | Authoritative or original source |
|---|---|---|---|---|---|
| Impact-Effort or Value-Effort Matrix | Fast initial sorting | Place options on impact and effort axes | Simple; visual; exposes disagreement | Subjective axes; favors quick wins; misses dependencies | Government of Ontario2 |
| RICE | Comparable product initiatives | Multiply reach, impact, and confidence; divide by effort | Includes breadth and uncertainty; easy to audit | Weak inputs create false precision; narrow strategic work can lose | Intercom4 |
| ICE | Growth tests and early ideas | Score impact, confidence, and ease | Very fast; useful when evidence is thin | Ease can reward trivial work; scoring anchors vary | Sean Ellis8 |
| MoSCoW | Scope inside a fixed release or timebox | Classify Must, Should, Could, and Won’t Have This Time | Clear language; creates contingency | ”Must” inflation; no ordering within categories | Agile Business Consortium5 |
| Kano Model | Product and service attributes | Classify must-be, performance, attractive, indifferent, and other response types | Shows nonlinear satisfaction; supports differentiated concepts | Needs research; varies by segment and time; ignores cost | American Society for Quality6 |
| Weighted Scoring or Decision Matrix | Decisions with several explicit criteria | Weight criteria, score options, and calculate totals | Flexible; transparent; supports cross-functional discussion | Weights can encode politics; correlated criteria double-count value | American Society for Quality9 |
| WSJF | Sequencing delay-sensitive jobs | Divide relative cost of delay by relative job size | Adds urgency and job size; useful for sequence | Relative scores can be gamed; weak for platform work | Scaled Agile7 |
| Cost of Delay and CD3 | Work whose value changes with time | Estimate value lost per unit of delay; CD3 divides it by duration | Makes time economically visible | Dollar estimates are difficult; duration uncertainty can dominate | Scaled Agile7 |
| Pareto Analysis | Defects, complaints, failure causes, and service demand | Rank categories by frequency or cost to find the vital few | Evidence-based; easy to communicate | Historical frequency is not future value; rare severe issues can disappear | American Society for Quality10 |
| Impact-Urgency Matrix | Incidents and service recovery | Combine breadth or severity of impact with time sensitivity | Fast triage; connects priority to response time | Reactive; requesters may inflate both dimensions | Atlassian11 |
| Numerical Assignment or Priority Groups | Large requirement lists needing a rough first pass | Assign high, medium, low, or numbered classes | Cheap; easy; scales well | Produces ties; criteria are often vague; little tradeoff information | Requirements prioritization review12 |
| Forced Ranking or Pairwise Comparison | A short list that must be ordered | Compare two items at a time until a rank emerges | Forces choices; produces a clear order | Slow at scale; does not show strength of preference | Luke Hohmann13 |
| 100-Point Method or Cumulative Voting | Aggregating stakeholder preferences | Give each person 100 points to distribute among options | Forces tradeoffs; shows intensity; participatory | Strategic stakeholders can dominate; dependencies and costs stay hidden | Requirements prioritization review12 |
| Buy a Feature | Customer or stakeholder workshops | Price features and give participants limited funds to spend, alone or together | Makes scarcity tangible; reveals reasoning through negotiation | Price setup shapes results; group dynamics can bias choices | Luke Hohmann14 |
| Prune the Product Tree | Balancing areas of a growing product or service | Place current and proposed capabilities on branches and releases | Shows product balance and sequencing; invites new needs | Qualitative; facilitation and tree design affect the outcome | Innovation Games15 |
| MaxDiff or Best-Worst Scaling | Measuring relative customer preference across many items | Repeatedly ask respondents to choose the most and least important items in small sets | Strong discrimination; avoids rating-scale inflation | Needs sound survey design and analysis; scores are relative | Sawtooth Software16 |
| Conjoint or Discrete Choice Analysis | Feature, price, package, and product configuration tradeoffs | Ask customers to choose among bundles; estimate part-worth utilities | Models realistic tradeoffs and willingness to pay | Research-heavy; sensitive to attributes, levels, sample, and model | IBM SPSS17 |
| Opportunity Scoring or ODI | Finding underserved customer outcomes | Compare importance with current satisfaction for desired outcomes | Starts with customer jobs and unmet needs; works across product types | Requires disciplined outcome statements and quantitative research | Strategyn18 |
| Importance-Performance Analysis | Service and experience improvements | Plot attribute importance against current performance | Directs resources to important weak areas; intuitive | Stated importance can inflate; quadrant cutoffs can be arbitrary | Martilla and James19 |
| Opportunity Solution Tree | Choosing customer problems before solutions | Connect a desired outcome to opportunities, solutions, and assumption tests | Keeps work tied to outcomes; exposes the opportunity space | Quality depends on discovery; does not by itself calculate priority | Product Talk20 |
| Impact Mapping | Linking delivery choices to behavior and goals | Map goal, actors, behavior changes, and possible deliverables | Makes the causal logic visible; prevents feature shopping lists | Causal links are hypotheses; needs evidence and testing | Gojko Adzic21 |
| Assumption Mapping | Deciding what to learn first | Plot assumptions by importance and strength of evidence | Directs experiments toward decision risk; covers desirability, viability, feasibility, and adaptability | Teams may confuse confidence with evidence; not a delivery ranking | Strategyzer22 |
| User Story Mapping | Creating coherent releases around a user journey | Arrange activities horizontally and stories vertically by priority or sophistication | Preserves end-to-end usability; reveals gaps and dependencies | Can privilege the happy path; not suited to portfolio choices | Agile Alliance23 |
| Now-Next-Later | Communicating sequence under uncertainty | Place work in confidence horizons rather than fixed dates | Honest about uncertainty; simple stakeholder view | Can become three undifferentiated backlogs; not a selection method by itself | ProdPad24 |
| Quality Function Deployment or House of Quality | Translating customer needs into product, service, and engineering requirements | Relate prioritized customer needs to technical characteristics | Connects voice of customer to design; shows tradeoffs and correlations | Data and workshop intensive; matrices can become unwieldy | American Society for Quality25 |
| Pugh Matrix or Concept Selection Matrix | Comparing physical, service, or digital design concepts | Compare concepts against a baseline across defined criteria | Supports concept learning and team discussion; handles mixed criteria | Baseline and criteria shape the answer; simple scores hide magnitude | SAE International26 |
| Analytic Hierarchy Process | Complex choices with many qualitative and quantitative criteria | Build a hierarchy and use pairwise comparisons to derive weights and priorities | Handles intangible criteria; includes a consistency check | Comparison count grows quickly; rank reversal and judgment load are concerns | Project Management Institute27 |
| Cost-Value Approach | Requirements where value and implementation cost both matter | Use pairwise comparisons to estimate relative value and cost, then inspect the value-cost map | Separates benefit from cost; supports release selection | Pairwise work grows rapidly; estimates can become expensive | Requirements prioritization review12 |
| FMEA and Risk Priority Number | Safety, reliability, manufacturing, service, and process risks | Score failure severity, occurrence, and detectability; prioritize mitigation | Proactive; works across designs, processes, products, and services | Multiplication can mask severe failures; scoring scales need care | American Society for Quality28 |
| Stage-Gate | Continuing, pausing, or killing innovation investments | Use evidence and criteria at decision gates between development stages | Limits sunk-cost momentum; fits physical products and high-cost services | Can become bureaucratic; weak gates keep too many projects alive | Stage-Gate International29 |
| NPV, ROI, Payback, and Profitability Index | Capital-intensive product and portfolio investments | Compare expected cash flows, return, or time to recover investment | Uses a common economic language; useful when forecasts are credible | Early product forecasts are fragile; undervalues learning, trust, and strategic options | Project Management Institute30 |
| Real Options | Staged bets under market or technical uncertainty | Value the right, but not the obligation, to expand, delay, switch, or abandon | Rewards learning and flexibility; values enabling investments | Hard to model and explain; assumptions can dominate the result | IBM Research31 |
| BCG Growth-Share Matrix | Allocating resources across products or businesses | Plot relative market share against market growth | Simple portfolio view; forces invest, hold, or exit discussion | Two variables oversimplify advantage; market definitions matter | Boston Consulting Group32 |
| GE-McKinsey Nine-Box Matrix | Portfolio investment across business units or product lines | Assess industry attractiveness and competitive strength | More nuanced than BCG; accommodates several factors | Scores and weights can be subjective; too coarse for feature choices | McKinsey & Company33 |
| Three Horizons | Balancing current products with emerging and future growth | Allocate attention across the core, emerging businesses, and future options | Protects longer-term innovation from near-term work | Often misread as a fixed timeline; does not rank projects within a horizon | McKinsey & Company34 |
| Strategic Buckets | Reserving capacity for different strategic purposes | Allocate resources to categories first, then rank work within each bucket | Prevents maintenance or easy work from consuming the entire portfolio | Bucket sizes encode strategy and politics; can shelter weak projects | Stage-Gate International35 |
| Portfolio Optimization | Selecting the best set under several constraints | Use mathematical programming to maximize value within budget, people, timing, and dependency limits | Chooses a portfolio, not just a ranking; handles constraints explicitly | Needs reliable inputs and specialist skill; can be hard to explain | Project Management Institute30 |
The table is not a menu from which to pick the most sophisticated method. Complexity should be earned.
If five people are comparing eight well-understood ideas, a carefully defined Impact-Effort Matrix may be enough. A manufacturer comparing product concepts might use a Pugh Matrix, drawing on customer and engineering requirements already developed through QFD. A service organization deciding which broken moments to improve might use Importance-Performance Analysis, with complaint patterns and journey research as evidence rather than as additional prioritization exercises. If leadership is allocating millions across product lines, a feature-level framework is the wrong instrument entirely.
Choose the lightest method that makes the important tradeoff visible.
How to choose the right framework
Start with the decision, not the acronym. The methods listed under each question are alternatives. They are not a bundle. For one decision, choose one primary framework and use the other information as evidence, constraints, or context. Do not run several scoring methods against the same candidates simply to make the answer appear more rigorous.
A different framework may be appropriate later because the decision has changed. Choosing a customer problem, selecting an initiative, defining a release, and sequencing delivery are separate decisions. Using one method at each stage is not the same as stacking several methods into one meeting.
If the question is “Which problem should we pursue?”
Choose an Opportunity Solution Tree when the team needs to map the opportunity space, Opportunity Scoring when it has quantitative research on unmet needs, Impact Mapping when it needs to examine the behavior change behind an outcome, or Importance-Performance Analysis when it has measured service or experience attributes. Choose the one that fits the decision and available evidence.
If the question is “Which idea should we test?”
Choose Assumption Mapping when the main question is what must be learned, ICE for a set of comparable growth tests, an Impact-Effort Matrix for a fast workshop, or a small weighted scorecard when the team has several explicit criteria. At this stage, learning value and reversibility often matter more than a forecast of total return.
If the question is “What belongs in this release?”
Choose MoSCoW when a fixed date requires hard scope choices, Story Mapping when the central risk is an incoherent customer journey, Kano when customer research must distinguish basics from differentiators, or Buy a Feature when a facilitated workshop needs to reveal stakeholder or customer tradeoffs.
If the question is “What should go first?”
Choose RICE when reach, impact, confidence, and effort are the useful comparison, WSJF when the economic effect of delay matters, Cost of Delay when that loss can be estimated directly, or a weighted decision matrix when the decision requires custom criteria. Then apply dependencies and capacity as constraints, not as additional frameworks. Priority and sequence are related but not identical. A lower-value prerequisite may need to happen before a higher-value outcome.
If the question is “Where should the organization invest?”
Choose Strategic Buckets when leadership must reserve capacity for different strategic purposes, Stage-Gate when the decision is whether an investment should continue, AHP when several difficult criteria require structured judgment, financial analysis when forecasts are credible, real options when learning and flexibility carry value, or a portfolio matrix or optimization model when the concern is the balance of the full investment set. Portfolio decisions need to consider strategic intent, not only the rank of each project in isolation.
If the question is “Which risk or failure deserves attention?”
Choose FMEA for potential failure modes, Pareto Analysis for recurring defects or service demand, or an Impact-Urgency Matrix for active incidents. Do not let average scores erase catastrophic but infrequent failure modes.
When a second framework earns its place
The default is one. Add a second method only when it answers a separate question that the primary framework cannot answer.
Story Mapping and MoSCoW can work together because the first creates a coherent journey and the second makes scope tradeoffs inside it. Strategic Buckets and a weighted scorecard can work together because the first allocates capacity among types of investment and the second compares candidates within a bucket. In both cases, the methods perform different jobs.
RICE, WSJF, and ICE applied to the same list do not perform different enough jobs to justify the overhead. If a team cannot explain the distinct purpose of the second framework in one sentence, it should not use it.
AI should improve the evidence, not inherit the decision
AI can make prioritization materially better. It can also make bad prioritization look finished.
The safest division of labor is straightforward: let AI increase the breadth, consistency, and testability of the evidence. Keep people accountable for objectives, values, constraints, weights, exceptions, and the final choice.
That boundary is consistent with the NIST AI Risk Management Framework, which emphasizes documented roles, human review, and accountability.36 It also reflects a practical limitation of current models. Language models may agree too readily with the person prompting them. OpenAI has publicly documented sycophantic model behavior,37 and a systematic study of models used as judges found that presentation order can affect comparisons.38
Use AI as an analyst, challenger, simulator, and recorder. Do not make it the unaccountable chair of the product council.
Five useful jobs for AI
1. Build the evidence matrix. Give AI approved research, support data, sales notes, analytics, defect reports, financial assumptions, and technical estimates. Ask it to normalize candidate descriptions, identify duplicate requests, and create a row for every claim with its source, date, segment, and evidence strength.
The source requirement matters. “Customers want this” is not evidence. “Twenty-three of 61 recent support cases from enterprise administrators involved this workflow” can be inspected.
2. Find missing perspectives. Ask what the current rubric ignores. Does it include accessibility, service operations, rollout cost, adoption effort, privacy, reversibility, product coherence, maintenance, and the effect on non-target customers? AI is good at scanning a long decision packet for absent categories.
3. Challenge the inputs. Ask AI to make the strongest case that the leading option is overrated and the losing option is underrated. Have it identify double-counted benefits, inconsistent time windows, unsupported confidence scores, hidden dependencies, and estimates that use different denominators.
This is more useful than asking, “Which option should we choose?” The latter invites a fluent answer. The former asks the model to improve the decision.
4. Run sensitivity and scenario analysis. A priority that changes whenever one weight moves by 10 percent is not a clear winner. Ask AI or an analytical tool to vary uncertain inputs, show which options remain near the top, and identify the assumptions that control the ranking. Run separate scenarios for customer segments, capacity levels, market conditions, and time horizons.
5. Maintain the decision record. After people decide, ask AI to draft a record containing the objective, candidates, evidence, framework version, scores, disagreements, final owner, overrides, what was deferred, and the trigger for review. At the next review, compare expected reach, impact, effort, and risk with what actually happened.
That last step turns a framework from a ritual into a learning system.
Five jobs AI should not own
Do not ask AI to decide:
- which customers the strategy is willing to disappoint;
- how much trust, safety, accessibility, or reputation is worth;
- which stakeholder has decision authority;
- whether confidential customer or company data may be placed in a model;
- whether the organization should accept the consequences of the choice.
These are not calculation problems. They are responsibility problems.
A practical AI-assisted prioritization protocol
Use this sequence for a consequential product decision.
- A human decision owner writes the outcome, scope, constraints, candidates, and review date before asking AI for an answer.
- The team chooses a framework and defines every criterion, scale, time window, and evidence standard before seeing the scores.
- AI compiles the supplied evidence, attaches sources, flags missing data, and separates facts from estimates and opinions.
- Relevant participants score independently before group discussion. This reduces anchoring on the loudest person or the first polished recommendation.
- AI checks the completed sheet for inconsistent units, duplicate criteria, unsupported claims, dependency conflicts, and arithmetic errors.
- The team reviews disagreements, runs sensitivity cases, and makes the decision. The accountable owner may override the ranking, but must record why.
- AI drafts the decision record. A person verifies it. The team reviews actual results when the trigger arrives and updates its scoring anchors.
The protocol is intentionally slower at the points where judgment matters and faster at the work that machines perform well.
A prompt worth reusing
Act as a critical product decision analyst, not the decision owner. Use only the evidence I provide. For each candidate, separate verified facts, estimates, assumptions, and missing information. Apply the supplied framework and scoring anchors consistently. Cite the source for every score. Do not invent values. Flag incomparable candidates, double-counted criteria, dependencies, obligations, and differences in time horizon. Then run sensitivity cases for the uncertain inputs, make the strongest case against the current leader, and list the human judgments required before a final decision. End with a draft decision record, not a recommendation presented as fact.
The prompt cannot repair missing research or a vague strategy. It can make those weaknesses harder to hide.
A 60-minute prioritization reset
Product professionals can use the following exercise before the next roadmap or portfolio discussion.
First 10 minutes: Write the decision
Complete one sentence:
We are deciding which [comparable opportunities, concepts, initiatives, or requirements] to pursue in order to change [specific outcome] for [specific customer or business group] during [time horizon], within [capacity and nonnegotiable constraints].
If the sentence is difficult to write, do not open a scoring sheet yet.
Next 10 minutes: Clean the candidates
Remove duplicates. Split oversized items. Combine fragments that only create value together. Separate obligations from optional bets. Move ideas without enough evidence into discovery rather than delivery.
Next 10 minutes: Choose the question and tool
Use the decision guide above and pick one primary framework. Stop there unless it leaves a separate, named question unanswered. If a complementary method is necessary, write down its distinct job before using it. Story Mapping can establish a coherent journey before MoSCoW trims the release. Strategic Buckets can allocate portfolio capacity before a weighted scorecard ranks work inside each bucket. These pairings normally happen in separate steps, not as competing scores in the same exercise.
Do not run several scoring frameworks against the same candidates in search of certainty. More frameworks do not guarantee more rigor. They often create more ways to count the same preference.
Next 15 minutes: Score independently and show the evidence
Every score needs a short rationale and an evidence label. Use ranges where uncertainty is meaningful. Do not average away disagreement before discussing it. A spread of 2, 5, and 9 contains information that an average of 5.3 conceals.
Next 10 minutes: Stress-test the result
Ask:
- What would have to be true for the top item to be wrong?
- Which input has the least evidence?
- Which option becomes first under a plausible alternative scenario?
- What obligation, dependency, or segment is missing?
- Are we selecting several convenient items instead of one consequential bet?
- What are we explicitly not doing?
Final 5 minutes: Record the choice
Name the decision owner, selected work, deferred work, reason for any override, first evidence expected, and review trigger.
Then stop scoring and make the choice.
The framework is successful when the team learns
The best prioritization process does not always produce the best outcome. Uncertainty does not disappear because a team worked carefully.
The better test is whether the process produced a clear choice from comparable options, used the best available evidence, made tradeoffs visible, protected important constraints, and created a way to learn.
A framework earns trust when people can see how the decision was made and when the organization is willing to revisit its assumptions. It loses trust when the score becomes a costume for authority.
Use the formula. Use the matrix. Use the workshop. Use AI to search, sort, challenge, calculate, and remember.
But keep the responsibility where it belongs.
Prioritization is not the arithmetic of deciding what goes first. It is the discipline of deciding what deserves the organization’s next unit of attention, and accepting what that choice leaves behind.
One Good Question
Where is your current prioritization process using a precise score to avoid an honest strategic choice?
Endnotes
Footnotes
Teresa Torres, “Prioritize Opportunities, Not Solutions,” Product Talk, February 13, 2019. ↩
Government of Ontario, “Impact vs. Effort Matrix,” updated November 26, 2024. ↩ ↩2
Paul Adams, “The First Rule of Prioritization: No Snacking,” Intercom, April 20, 2016. The article uses “snacking” for convenient, low-impact work. ↩
Sean McBride, “RICE: Simple Prioritization for Product Managers,” Intercom, January 5, 2018. This is Intercom’s first-party explanation of the inputs, formula, and intended use. ↩ ↩2 ↩3
Agile Business Consortium, “What Is MoSCoW Prioritization?.” The guidance defines the four categories, the cancellation test, and the recommended capacity allowance for Must Have requirements. ↩ ↩2 ↩3 ↩4
American Society for Quality, “What Is the Kano Model?.” The category descriptions in the article are paraphrased from the method, not quoted passages. ↩ ↩2 ↩3
Scaled Agile, “Weighted Shortest Job First.” This is the first-party SAFe definition of WSJF and its relative cost-of-delay components. ↩ ↩2 ↩3
Sean Ellis, presentation at Growth Marketing Summit 2024, slide 28. The slide presents ICE as Impact, Confidence, and Ease for prioritizing tests. ↩
American Society for Quality, “Decision Matrix.” ↩
American Society for Quality, “Pareto Chart.” ↩
Atlassian, “How Impact and Urgency Are Used to Calculate Priority,” Jira Service Management Cloud documentation. ↩
Sunita Malgaonkar, Sherlock Licorish, and Bastin Tony Roy Savarimuthu, “Understanding Requirements Prioritisation: Literature Survey and Critical Evaluation,” IET Software 14, no. 3 (2020). The review covers numerical assignment, cumulative voting, the cost-value approach, and other requirements-prioritization methods. ↩ ↩2 ↩3
Luke Hohmann, “20/20 Vision.” ↩
Luke Hohmann, “Buy a Feature.” ↩
Luke Hohmann, Innovation Games: Creating Breakthrough Products Through Collaborative Play (Addison-Wesley, 2006). ↩
IBM, “Introduction to Conjoint Analysis Methodology,” IBM SPSS Statistics documentation. ↩
Strategyn, “Outcome-Driven Innovation Process.” ↩
John A. Martilla and John C. James, “Importance-Performance Analysis,” Journal of Marketing 41, no. 1 (1977): 77-79. ↩
Product Talk, “Opportunity Solution Tree.” ↩
Gojko Adzic, “Impact Mapping.” ↩
Strategyzer, “How Assumptions Mapping Can Focus Your Teams on Running Experiments That Matter.” ↩
Agile Alliance, “Story Mapping.” ↩
ProdPad, “Now-Next-Later Roadmap.” ↩
American Society for Quality, “Quality Function Deployment.” ↩
George Nagy, “A Review of the ‘Pugh’ Methodology for Design Concept Selection,” SAE Technical Paper 940887 (1994). ↩
Project Management Institute, “Using the Analytic Hierarchy Process to Select and Prioritize Projects.” ↩
American Society for Quality, “Failure Mode and Effects Analysis.” ↩
Stage-Gate International, “Discovery-to-Launch Process.” ↩
Project Management Institute, “Project Portfolio Selection with Mathematical Programming.” The source discusses financial-selection methods and constrained portfolio optimization. ↩ ↩2
IBM Research, “Prioritizing a Portfolio of Information Technology Investment Projects.” ↩
Boston Consulting Group, “Growth-Share Matrix.” ↩
McKinsey & Company, “Enduring Ideas: The GE-McKinsey Nine-Box Matrix.” ↩
McKinsey & Company, “Enduring Ideas: The Three Horizons of Growth.” ↩
Stage-Gate International, “Product and Technology Strategy.” ↩
National Institute of Standards and Technology, “AI Risk Management Framework Core,” 2023. The Core calls for documented roles, accountability structures, human oversight, and periodic review. ↩
OpenAI, “Sycophancy in GPT-4o: What Happened and What We’re Doing About It,” April 29, 2025. ↩
Lin Shi et al., “Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge,” AACL-IJCNLP 2025. The study reports that position bias varies across models and tasks and is not attributable to random chance in its experiments. ↩
- Product management
- Prioritization
- Decision quality
- Artificial intelligence



