Idea validation is a decision procedure, not a confidence exercise
Validation is finished when a prewritten rule tells you to build, revise, retest or stop — and you obey it. Everything before that is preparation. This guide gives you the framework, the eleven-field scorecard and the default numbers most validation writing leaves blank.
Published Updated Published by Foundable
Overview
Idea validation converts opinions into observed behaviour. You name the one belief that must be true before a bigger build, state a promise a stranger can act on, put it in front of a reachable audience, and ask for exactly one concrete action. The threshold is set before the test so the result cannot be reinterpreted afterwards. The scorecard below is the artefact you carry out.
Quick answers
What is idea validation?
It is the pre-build process of testing whether a specific audience has a real problem, understands your promise, and takes an observable action — a reply, a call booked, a deposit, a payment. Compliments are not validation; behaviour is.
What is a practical idea validation framework?
Five things get tested before a large build: the audience, the painful problem, the promise, the market-facing test, and the behaviour threshold. Everything else — the deck, the brand, the roadmap — waits.
What is the startup validation process, step by step?
Name audience and assumption, build a test artefact (page, message, interview script, waitlist, demo, paid pilot), ask for one concrete action, set the threshold in advance, then decide from what people did rather than what they said.
What should an idea validation scorecard include?
Eleven fields: audience, assumption, method, call to action, response threshold, urgency, pain, willingness to pay, objections, buyer action, and the build / revise / retest / launch decision. The first five are locked before the test; the last six are filled after.
How many people do I need to send this to?
Defaults we publish rather than hide: 20 qualified people for cold outreach, 10 for warm introductions, roughly 100 visits for a landing-page test. These are planning boundaries, not statistical proof — override them before the test locks, never after.
When should I stop instead of retest?
Pause when 20 delivered messages produce no qualified action, and stop when a second round with a repaired promise and a different audience produces the same nothing. Two clean nulls beat a year of tinkering.
At a glance
| Stage | Artefact | Default threshold | Locked before the test |
|---|---|---|---|
| Assumption | One named belief from seven types | — | Yes |
| Promise | One sentence, four required elements | — | Yes |
| Method | Behaviour-based test plus a single CTA | — | Yes |
| Reach | Cold 20 / warm 10 / landing page ~100 visits | n as chosen | Yes |
| Signal | Seven scored dimensions | ≥1 money commitment to continue | No |
| Decision | Build, revise, retest, launch or stop | 0 qualified actions in 20 delivered → pause | No |
- Assumption
- Artefact
- One named belief from seven types
- Default threshold
- —
- Locked before the test
- Yes
- Promise
- Artefact
- One sentence, four required elements
- Default threshold
- —
- Locked before the test
- Yes
- Method
- Artefact
- Behaviour-based test plus a single CTA
- Default threshold
- —
- Locked before the test
- Yes
- Reach
- Artefact
- Cold 20 / warm 10 / landing page ~100 visits
- Default threshold
- n as chosen
- Locked before the test
- Yes
- Signal
- Artefact
- Seven scored dimensions
- Default threshold
- ≥1 money commitment to continue
- Locked before the test
- No
- Decision
- Artefact
- Build, revise, retest, launch or stop
- Default threshold
- 0 qualified actions in 20 delivered → pause
- Locked before the test
- No
- Defaults are planning boundaries. Change them before you launch; changing them after is how a failed test becomes a story.
Only one belief can be on trial at a time
A test that moves audience, price and channel at once cannot tell you which of them was wrong. Pick the single assumption whose failure would waste the most work, from seven types: audience, pain, promise, channel, price, timing, and willingness to switch. Rank them by impact if wrong multiplied by how uncertain you currently are, and write the winner down before you touch anything else.
- Audience — you are talking to people who do not have this problem often enough to pay for it.
- Pain — the problem is real but tolerable, so nothing changes when you offer relief.
- Promise — people want the outcome but cannot tell from your words that you deliver it.
- Channel — the offer is fine and nobody who needs it ever sees it.
- Price — the outcome is worth something, just not what you are asking.
- Timing — it matters in March and you are asking in September.
- Willingness to switch — the current workaround is bad and free, and free is winning.
Keep the ranking visible and dated. When the test fails, the register tells you which assumption to retire and which to promote — that chain is worth more than any single result.
A promise a stranger cannot act on is not a test
Write one sentence containing four elements: who it is for, what outcome they get, why it matters now, and what they should do next. If any element is missing, the reader supplies their own version and you learn nothing about yours. Read it aloud to someone outside the project; if they ask a clarifying question, the sentence has failed and the test has not started.
- Who: a group specific enough that a member recognises themselves in three words.
- Outcome: a state they can picture, not a feature list.
- Now: the reason this week rather than someday.
- Action: exactly one, and it must be observable by you.
Behaviour-based methods exclude the survey by construction
Nine methods carry real information because each one costs the respondent something: interviews with a clear ask, landing pages, direct outreach, waitlists, demos, referral tests, pilot offers, deposits, and pre-orders. Opinion surveys cost nothing, so they measure politeness. Choose the most expensive method your audience will tolerate this week — the more it costs them, the less you need.
- No audience yet, 40 contacts: outreach with a call ask, or three paid pilot offers.
- A small list or community: waitlist with a qualifying question, or a demo booking.
- Budget for traffic: landing page with a checkout start or deposit as the action.
- Existing clients: a referral test, which prices the promise and the channel at once.
Method availability is a function of reach, not preference. Write down what you can actually reach this week before you choose; a landing-page test with no traffic plan is a decision to learn nothing.
The threshold has to exist before the first message goes out
Decide in advance what counts as enough: replies, calls booked, qualified signups, referrals, checkout starts, deposits, payments, or clearly stated objections. Write the number and the deadline in the same sentence. A threshold set afterwards is a rationalisation with a decimal point.
| Method | Default n | Qualifying action | Continue if |
|---|---|---|---|
| Cold outreach | 20 delivered | Call booked or paid pilot accepted | ≥1 money commitment |
| Warm introductions | 10 delivered | Call booked | ≥2 calls, ≥1 money commitment |
| Landing page | ~100 qualified visits | Checkout start or deposit | ≥1 money commitment |
| Waitlist | 50 signups | Reply to the qualifying question | ≥10 qualified replies |
| Paid pilot | 5 offers made | Invoice accepted | ≥1 accepted |
- Cold outreach
- Default n
- 20 delivered
- Qualifying action
- Call booked or paid pilot accepted
- Continue if
- ≥1 money commitment
- Warm introductions
- Default n
- 10 delivered
- Qualifying action
- Call booked
- Continue if
- ≥2 calls, ≥1 money commitment
- Landing page
- Default n
- ~100 qualified visits
- Qualifying action
- Checkout start or deposit
- Continue if
- ≥1 money commitment
- Waitlist
- Default n
- 50 signups
- Qualifying action
- Reply to the qualifying question
- Continue if
- ≥10 qualified replies
- Paid pilot
- Default n
- 5 offers made
- Qualifying action
- Invoice accepted
- Continue if
- ≥1 accepted
- Planning defaults from ordinary practice, not benchmarks. We have no proprietary dataset behind these numbers and say so.
Read seven dimensions, then obey the rule you wrote
Score urgency, audience fit, problem pain, promise clarity, willingness to pay, objections, and buyer action. Any dimension with too little data is marked insufficient evidence rather than averaged away — a half-filled scorecard must not be able to pass. Then take the decision your prewritten threshold requires: build, revise, retest, launch or stop.
- Build — threshold met and at least one person paid or committed money.
- Revise — interest is real, the promise or price is not; change one variable and retest.
- Retest — the test was contaminated (wrong list, broken link, bad send time); rerun it clean.
- Launch — you already have paying commitments and the constraint is delivery, not proof.
- Stop — two clean rounds, repaired promise, different audience, still nothing qualified.
A completed scorecard is evidence of work, not evidence of demand
The most common failure in this method is treating the artefact as the result. Eleven filled fields with zero money commitments is a documented no. The scorecard's job is to make that no unarguable and dated, so you do not relitigate it in month four when the idea starts feeling good again.
Honesty label: this framework is a decision aid drawn from ordinary practice and public sources, not a validated instrument. It has no predictive accuracy claim and we do not publish one.
The numbers move for СНГ reach, and so do the money instruments
Willingness-to-pay tests need an instrument the buyer already trusts. In Russia and the neighbouring markets that usually means an СБП deposit link or a ЮKassa счёт rather than a card checkout; B2B pilots are structured as a счёт-оферта or a short договор оказания услуг. Audience reach lives in profile Telegram chats, VK communities, Avito rubrics and industry catalogues — and their reply rates differ enough that you should record your own baseline instead of importing ours.
Price fields should carry the currency and the tax regime (НДС / НПД) from the first test, because both change the number the buyer hears.
Do this step for me
Send the six locked fields to Ted and get the test artefact, the outreach list structure and the scored card back. Nothing leaves your browser until you sign in.
- Audience
- Assumption to prove
- Test method
- Call to action
- Threshold (n + deadline)
- Buyer action observed
What you leave with
- One named assumption, ranked against the other six and dated.
- A promise sentence with all four required elements, readable by a stranger.
- A behaviour-based method matched to reach you actually have this week.
- A written threshold with an n and a deadline, locked before launch.
- A scored eleven-field card with explicit insufficient-evidence states.
- A build / revise / retest / launch / stop decision you are committed to.
Related Foundable resources
Workflow
- 01
Rank the assumptions
List all seven, score impact-if-wrong against current uncertainty, and pick one. Ted will rank them for you and show its reasoning; you can reorder before locking.
- 02
Build the test artefact
Turn the promise into the page, message, interview script, waitlist or pilot offer the chosen method needs, with exactly one call to action.
- 03
Lock the threshold, then launch
Write n, the qualifying action and the deadline into the card. Send to the reachable group — 20 cold, 10 warm, or your own traffic plan.
- 04
Score and decide in public
Fill the six post-test fields, mark thin dimensions insufficient, and record the decision with its date so a future you cannot quietly overturn it.
Limitations and corrections
- The default thresholds are planning boundaries drawn from ordinary practice, not statistical power calculations and not benchmarks from a dataset we hold.
- The scorecard has no published predictive accuracy. It structures a decision; it does not forecast an outcome.
- Foundable is a commercial product and this guide routes to our tools. The framework works on paper, and we say so rather than gating it.
- Reply-rate expectations vary enormously by market, channel and list quality. Record your own baseline before trusting any figure on this page, including ours.
- Corrections are dated on this page. Tell us what is wrong at /contact.
Corrected in public: tell us what is wrong at /contact and the change is dated on this page.
Frequently asked questions
Yes — outreach, interviews and referral tests cost time rather than budget. Paid traffic buys speed, not certainty, and it is the wrong purchase before the promise sentence is stable.
That is a completed test with a clear reading: the promise is attractive and the price, the moment or the buyer is wrong. Change one of the three and retest; do not change all three and call it a new idea.
Only if the signup cost something — a qualifying answer, a deposit, a calendar slot. A free email address measures curiosity, and curiosity does not renew.
Start free with Ted.
Bring the messy version of the idea. Ted turns it into a first move, something real, and a way to see what people actually do — free to start, no card.
AI output needs your review. No income is promised or guaranteed.