Digital Product Discovery: A Practical Guide
Plenty of teams that “did discovery” still build the wrong product.

Digital Product Discovery Is a Decision System, Not a Phase
Most guides show discovery as four tidy boxes: research, define, ideate, validate. The evidence doesn’t support anything that neat. A 2025 systematic review of product discovery found practices that keep coming up: understanding user problems, testing assumptions, iterating, and mixing qualitative with quantitative evidence. It found no universal sequence. Teams go back to earlier stages as new evidence arrives. The review also admits a hard limit. Most published guidance on discovery comes from blogs and vendor material, so nobody has yet shown which techniques work in which situations.
This matters because it changes how you should treat frameworks. “Interview five customers a week” and “always ship an MVP” are working habits, not proven laws. Adopting a framework doesn’t give you evidence. A better definition of discovery is disciplined judgment under uncertainty. You find out what you don’t know, figure out which decision depends on it, and collect only the evidence that decision needs. If you want the basics first, our explainer on what product discovery means covers the vocabulary, and SigOS has a readable overview of how product discovery creates value for SaaS teams.
The phrase has a second meaning that teams tend to forget: how buyers actually find products, through search, social feeds, marketplaces and, more and more, AI assistants. A product can solve a real problem and still fail because nobody comes across it. Good discovery covers both questions. We come back to the second one later in this guide.
Start With the Decision, Then Pick the Method
Every useful discovery activity begins with a sentence you could be wrong about. “Independent consultants will pay for a portal that cuts the time they spend on client updates” is testable. “Consultants need better tools” isn’t. For each assumption that matters, write down four things before you choose a method:
- Assumption: what has to be true for this product or feature to work.
- Prediction: what people would do if it were true.
- Smallest test: the cheapest activity that could show that behavior.
- Decision rule: which result means build, change, defer or stop.
The method follows from the uncertainty. Before you build a subscription checkout, test a clickable flow and measure how many qualified users complete it. Before you migrate a CMS, sample the real content types and prototype the riskiest editorial task. Before you add user permissions, watch who actually does each action and work out what a mistake would cost. Each of these is a different question, so each gets a different test.
This is also where scope gets cut. When a project management consultant came to us wanting an AI assistant that helped project managers with “everything,” we ran a blueprint process before writing any code. That work narrowed Workform into a focused MVP. It became an assistant that understands a project by pulling together data from Slack, email, Asana and meetings, not a general task generator. The broad idea wasn’t wrong. It just couldn’t be tested, and you can’t make a sound build decision about something you can’t test.
For a wider set of options once you know your question, our guide to product discovery techniques sorts ten methods by the kind of risk each one reduces.
Match the Artifact’s Fidelity to What You Need to Learn
A 2025 Aalto University study of eight companies named four common experiment objects: wireframes, prototypes, proofs of concept and MVPs. Any of them can be thrown away, reused, changed or grown into the real product. A related Aalto thesis found that teams often didn’t design these artifacts around what they wanted to learn. They built what was convenient and then tried to read meaning into the response.
| Artifact | What it can tell you | What it cannot tell you |
|---|---|---|
| Rough concept or sketch | Which problem people care about most | Whether they can use the solution or will adopt it |
| Wireframe | Whether the structure and flow make sense | Whether the product looks credible or is technically feasible |
| Interactive prototype | Whether people can finish the key task without help | Whether they will pay or come back next week |
| Proof of concept | Whether a risky technical piece works at all | Whether anyone wants it |
| Landing page | Whether people take an action when they see the offer | Why they acted, or whether they would keep using it |
| MVP or paid pilot | Whether real use and payment happen | Long-term retention, unless you run it long enough |
A 2026 case from WRI’s Product Studio shows rough concepts used well. The team used AI to generate five concepts for urban-cooling practitioners and had 46 practitioners react to them. All five tested well, which is common and tells you almost nothing. So the team asked people to rank them. One concept got 11 “most valuable” votes and no “least valuable” votes. That’s useful evidence for prioritizing. It isn’t proof of adoption, and WRI says so.
WRI also gave a warning that applies to most teams now. AI tools make polished prototypes cheap, and polish arriving too early causes problems. Feedback shifts to button colors and copy, the team gets attached to one direction before it has asked the right questions, and changing course starts to feel expensive. Keep early artifacts rough on purpose, and write down what each one is supposed to reveal. Our breakdown of proof of concept versus prototype covers the technical side of this choice. These recent MVP examples show what happens when the artifact answers a different question from the one the team needed answered.

Interviews That Produce Evidence Instead of Politeness
The most common failure in discovery interviews is showing the solution too early. Once someone has seen your idea, the conversation turns into a reaction to it, and most people react kindly. One r/ProductManagement commenter compared demoing to customers with showing off family photos: everyone says something nice. Start with the person’s work, their recent experience and what it cost them. “Tell me about the last time you prepared a client report” gets you facts. “Would you use a tool that automated reports?” gets you a courtesy.
A few practices make interviews more reliable:
- Recruit by role or task, not with a pitch written around your idea. A pitch attracts people who already agree with you.
- Ask about past behavior, workarounds and consequences. Current workarounds, like spreadsheets, email threads and manual copying, show you what you’re actually competing against.
- Save solution feedback for usability tests, where you watch people try a task without your help.
- Write down evidence that goes against your preferred answer. If you can’t find any, ask yourself whether you were really listening.
How many interviews is enough? Published advice varies a lot. One business idea validation guide recommends 30 to 50 target-customer interviews before moving to landing pages and pilots. The 2025 review warns that interview counts can become a vanity metric, and suggests the time from question to decision may be a better measure. Stop when new conversations stop changing what you’d decide. Don’t stop just because you hit a number.
Access is often the real problem. The 2024 Maze Future of User Research survey of more than 1,200 product professionals found that time and bandwidth (62%) and recruiting (60%) were the biggest constraints. In B2B, customers may only be available every few months. Practitioners get around this by sitting in on sales and customer success calls, reading support tickets, using in-product invitations tied to a scheduling link, and agreeing with Sales on which accounts can be contacted. These channels are useful but filtered. A gatekeeper who only passes along good feedback will lead you to a product that suits a few customers and misses the wider market.
Synthesis needs structure too, or interview notes turn into opinions within a week. A shared agenda like this product discovery meeting template helps a team turn raw notes into named problems. Our guide to user research methods covers how to pick a study design that fits the decision.

Validate the Problem, the Solution and the Payment Separately
Attention isn’t demand. A sign-up, a download or an excited beta comment shows interest. None of them shows that someone will pay or keep using the product. Treat these as three separate questions, in this order:
- Problem: Do target users have a problem that costs them enough to change how they work? Interviews and observation answer this.
- Solution: Can they understand and use what you’re proposing? Prototypes and usability tests answer this.
- Payment: Will they pay for the result? Pre-sales and paid pilots answer this. A market validation framework makes the same point: a customer paying is stronger evidence than a customer saying yes.
Landing pages sit between the second and third questions. They show whether people act on an offer, but not why. One Indie Hackers member reported waitlist conversion of 30% to 50% from just over 50 visitors. That’s an anecdote, and the sample is too small to compare with your situation. The pattern that practitioners share is still worth copying: run the landing page first, then interview the people who signed up to learn what made them act. For pricier products, many reverse the order and start with interviews, because each sale is worth enough to justify the time.
Willingness to pay is the weakest part of most discovery work, and the published research covers it poorly too. A survey asking what people would pay tells you what they’re willing to say. A paid pilot with ten customers tells you far more. For a step-by-step way to set proof thresholds before you test, Fundl’s evidence-first startup validation playbook is a practical companion. Once these three questions have answers, our MVP development process guide picks up at the scoping stage.
When You Run Experiments, Make Sure They Can Teach You Something
After launch, discovery often moves into A/B tests and feature flags. This is where teams most often fool themselves without noticing. Three published engineering accounts are worth learning from.
Count valid learning, not wins
Spotify reported in 2025 that about 12% of its experiments were “wins,” while about 64% produced valid learning. Much of that learning came from finding out what not to ship. Spotify’s definition is strict, though. A neutral result only counts as learning if every relevant metric had enough data to detect an effect. A null result on an underpowered metric doesn’t mean “no effect.” It means “we don’t know.” Spotify also warns that teams can inflate the learning rate by accepting less precision, so it tracks learning alongside win rate, test volume and precision. These are internal figures, not an industry benchmark. The principle carries over to any team.
Decide the rules before you see the results
Deliveroo’s experimentation team insists that hypotheses, success criteria and rollout rules are written down before a test starts. If you pick the rollout rule after seeing the data, you’ll find a story that fits whatever happened. Deliveroo also shows that not everything needs a test. It brought back restaurant favouriting without one. The feature had done well in consumer research and an employee rollout, rolling it back would have confused users, adoption would take too long for metrics to move quickly, and the feature made future experiments possible. Run a test when what you’d learn is worth more than the test costs.
Check the instrumentation before you trust the dashboard
Many experiments are invalid before anyone looks at the results. Feature-flag documentation spells out the traps. Flagsmith defines exposure as the moment a user is actually served a variation, and says exposure should be recorded when the variation is displayed if rendering happens after the flag is checked. Its event names are case-sensitive, and events that don’t match are silently ignored. OpenTelemetry’s feature-flag convention, still marked as in development, emits an evaluation event every time code checks a flag, even when nothing changes for the user. That’s operational telemetry. It doesn’t prove anyone saw the treatment. LaunchDarkly’s documentation warns that the unit you randomize must match the unit you analyze. If you assign by account and analyze by user, your results are misleading.
In practice, treat tracking as part of experiment design, not as cleanup after launch. Log exposure where the user actually sees the change, use the same identity for exposure and conversion events, and check event names against the spec before you start.
Discovery Also Has to Answer How Buyers Will Find You
Knowing that people want something doesn’t mean you can reach them. Buyers rarely follow a single path. Rithum’s 2023 Consumer Behavior Report found that 83% of global shoppers visited at least two websites before buying. Salesforce’s Connected Shoppers survey of 8,350 shoppers, run at the end of 2024, found that 53% discover products on social media, rising to 76% for Gen Z. Yet when Emplifi and Semrush asked US and UK shoppers in 2026 where they start researching a product, search led at 33%, then marketplaces at 18%, social at 15% and AI tools at 11%.
These figures don’t contradict each other. They answer different questions. Social sparks ideas. Search and marketplaces are where people go to research. AI is growing quickly from a small base: Capgemini found that 58% of consumers say they prefer AI recommendations, but stated preference is not behavior. Be wary of any single statistic that claims to tell you “where customers discover products.”
For your product, the practical move is to validate distribution alongside demand. Find out where your buyers already look for similar products. Read competitor reviews for complaints you could address. Check whether your own site helps people find things. Constructor’s surveys found that 68% of shoppers think retailer site search needs an upgrade, and 66% go to Amazon when it lets them down. These are vendor-sponsored surveys, but the frustration is familiar. Some products also need explaining before anyone will buy. In NudFud’s ecommerce build, the challenge was that a shopper landing on a cracker page needed to see nutrition panels, certifications and variant differences before the product made sense. Discovery for that store meant working out what a first-time visitor needed to understand, not only what the brand wanted to say.
Most Discovery Failures Are Organizational
Read enough practitioner threads and a pattern shows up. Discovery rarely fails because nobody knew a framework. It fails because evidence never reaches the decision. The causes repeat: customer access is blocked or filtered, there are no agreed criteria for choosing between options, incentives favor safe incremental work, and research arrives after the roadmap is already locked.
The numbers support this. ProductPlan’s 2024 State of Product Management survey found that 39% of product managers spend most of their time on delivery, and that senior leadership (31%) shapes strategy more often than customer feedback (27%). The sample is self-selected, but it fits what practitioners describe. If the loudest stakeholder outranks the evidence, extra interviews won’t help.
Two fixes come up again and again. First, agree on decision criteria before arguing for any feature, for example its effect on onboarding complexity, churn and delivery effort. That turns an argument about opinions into a comparison against shared criteria. Doodle took this further with a company-wide prioritization framework tied to its strategic pillars. In a case study published by Atlassian (which sells discovery software, so read it with that in mind), Doodle reported that planning time fell from up to 14 hours a quarter to a 20-minute monthly session. Our product strategy framework guide covers how to set up those criteria.
Second, test operational fit before large commitments. Birmingham City Council’s Oracle project is a public warning. Heavy customization to fit existing systems pushed the cost from an estimated £19m to around £100m, and the forecast savings never appeared. Public reporting doesn’t tell us what research was done beforehand, so this isn’t proof of a discovery failure. It does show how integration and workflow problems found late become the most expensive problems of all. This is also the area that requirements intelligence tries to address: finding hidden requirements before they spread into the build.
Cadence needs the same honesty. Weekly interviews sound great until you’re selling to enterprise buyers who can only be reached every six months. Practitioners say regular “what’s working” conversations with a handful of accounts count as discovery, as do insights gathered on existing account-management calls. Our piece on continuous product discovery sets out a sustainable rhythm after launch.
How Much Discovery Is Enough
Some practitioners call formal discovery “enormously expensive” and say a small company should rely on judgment. Others say it’s the cheapest insurance available. Both are right in different situations. Scale the effort to three things: how uncertain you are, how much is at stake, and how hard the decision would be to reverse. A copy change on a pricing page needs a quick check. A new product line, a platform migration or a pricing model change deserves real time.
There’s no evidence-based duration. One lead designer described using one to two weeks for larger features. An agency case for Carrofina used five weeks of discovery inside a 7.5-month design process. Aim for enough evidence to act with confidence, not certainty. No amount of research would have saved a product like the Humane AI Pin from what users experienced after launch, which is why post-launch signals have to keep feeding decisions.
If you’re starting now, this is the sequence we’d use:
- Write down the assumption that would hurt most if it were wrong, as a sentence someone could disagree with.
- Talk to the people who have the problem about what they did recently, before you show them anything.
- Separate repeated problems from one-off complaints and feature requests.
- Build the roughest artifact that can answer your current question, and say what it is meant to reveal.
- Set your decision rule before you look at results.
- Ask for real commitment, such as a pilot, a deposit or a signed letter, before you fund the full build.
The output of discovery is a decision: build, revise, defer or stop. A finished interview script, journey map or prototype proves nothing unless it changes what you build. For more detail on running each stage, see our walkthrough of the product discovery process.
Discovery won’t remove every risk. What it does is show you which risks to deal with before code, budget and reputation are tied to them, and that’s easiest while the product is still an idea. If you’re facing that decision and want strategy, design and engineering judgment in the same room before anything is built, that early work is what Refact’s product design work is built around.
Building a product and unsure what to scope first? Let’s talk. Free 30-minute call, no pitch.
Parnia Sebti is a project and account manager at Refact, coordinating teams, clients, timelines, and delivery across the studio’s work. She helps keep projects organized from planning through execution, making sure communication stays clear and priorities stay aligned. Her role connects client needs with the internal team’s workflow, helping turn requirements, feedback, and moving parts into structured delivery. At Refact, Parnia also contributes to shaping the internal tools and processes the team uses to manage projects more effectively and keep work moving with clarity.
More from Parnia Sebti


