Synthetic audience testing · built on peer-reviewed methodology
How many of 150 realistic buyers would say yes to your website?
Paste your URL. A fixed panel of 150 realistic buyers from your industry reads your page and reacts in plain words. You see how many would buy — or book — how you rank against 1,100+ scored pages, and exactly what stops the rest — plus a heatmap of where attention actually lands and a UX review of the friction standing in their way.
Report in ~5 minutes · $49 one-time · auto-refund if we can't produce it · see a sample report →
What's a synthetic panel? 150 AI-simulated buyers, each modeled on how real customers in your market think and decide — a method validated against 9,300 real survey answers. Not real people, and we never pretend otherwise.
We read only the public page at the URL you paste — no account, no tracking pixels on your site, deleted on request. How we handle data →
“How likely are you to subscribe to this?”
def. not
unsure
def. yes
“The free tier sounds generous, but I can't tell what the paid plan actually adds. I'd try it and probably never upgrade.”
Built for your industry
The panel that judges your page is from your world.
A coaching site is judged by coaching prospects — ROI skeptics, comparison shoppers, people ready to invest. A SaaS page is judged by SaaS buyers. Same method, your market.
Agencies & studios
Marketing, design, dev. See who would book an intro call — and what stops the rest.
Therapists & counselors
Private practice. See who would book a consultation, in plain words.
Founders & SaaS
Landing pages, app listings, launches. Test variants and competitors before traffic.
Three ways to run it: Diagnostic $49 · Variant ranking $79 — rank 2–4 versions, the method's strongest mode · Competitor scan $99. See pricing ↓
New — in every report
See where their eyes go, what they never see — and what's in their way.
Alongside the panel's verdict, every report now maps your page's predicted attention and reviews its UX friction — so you know why the score is what it is, and what to fix first.
What you get
- An attention heatmap of your page from a model trained on thousands of real human eye-tracking recordings — where attention concentrates, and the cold zones it never reaches
- The fold math: what share of predicted attention lands above the fold — and whether your price, proof, and call-to-action live in the warm zone or the dead zone
- Per-buyer blind spots: for each buyer type on the panel, what they fixate on and what they miss — the skeptic who never sees your proof, the budget buyer who never finds your price
- Reading order: the sequence a first-time visitor's eye actually follows through your page
- A UX friction review: severity-ranked usability findings against published inspection methods (Nielsen & Molich, 1990), a walkthrough of your page's primary task, and accessibility basics — a fix-first list, not a lecture
Predicted attention from an eye-tracking-trained saliency model plus the panel's buyer types — honest label: it's a prediction of where people look, not a recording of your visitors, and it explains the score rather than replacing it.
How it works
A consumer research panel, rebuilt in software.
The same four stages a research agency runs over a month — compressed into minutes, at a price a small business survives.
We read your page like a customer does
Screenshot + copy extraction of your website, landing page, or app listing. The panel reacts to exactly what visitors see — headline, images, pricing, all of it.
You approve the audience
We suggest the industry lane your page should be benchmarked against. Nothing runs until you confirm the fixed panel and baseline.
150 strangers respond in their own words
Each persona reacts in free text — no forced ratings. Research shows direct numeric ratings from AI are unrealistic; natural reactions are where the signal lives.¹
Reactions become research-grade metrics
Semantic Similarity Rating maps every reaction onto a purchase-intent distribution, benchmarked against 1,100+ scored pages — plus the top objections, quoted.
The science
Not vibes. A published method, validated against 9,300 real consumers.
150 Strangers implements Semantic Similarity Rating (SSR) — a technique developed by researchers at PyMC Labs and Colgate-Palmolive and tested against 57 real consumer surveys.
“LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings”
Maier, Aslak, Fiaschi, Rismal, Fletcher, Luhmann, Dow, Pappas & Wiecki (2025). arXiv:2510.08338 · open-source reference implementation on GitHub
- Asking AI for a 1–5 rating directly fails. Models cluster on “safe” middle answers and produce distributions nothing like real consumers. The study measured it: only 0.26 distributional similarity.
- Free-text reactions, mapped by meaning, work. SSR embeds each reaction and measures its semantic distance to calibrated anchor statements — recovering realistic distributions (similarity > 0.85) and product rankings at ~90% of what a repeated human panel achieves.
- Detailed personas are non-negotiable. With rich conditioning the method reached ~90% reliability; without it, signal collapsed to ~50%. That's why every lane uses a fixed, detailed panel before we run anything.
- Comparison is where it shines. The method's strength is ranking — which variant, which competitor, which message wins. We engineered the whole product around that, instead of pretending one number predicts your conversion rate.
Straight answers
What this can and can't tell you.
A research tool you can't trust is worthless. So here is exactly where the method is strong — and where we'll refuse to oversell it.
Reliable for
- Ranking variants: which headline, pricing frame, or screenshot set your audience prefers
- Competitive position: how your page lands next to up to 3 competitors, same panel, same question
- Objection mining: the recurring reasons skeptics say no — quoted, clustered, segmented
- Message clarity: whether your value proposition is even understood at a glance
- Segment fit: which of your audience segments responds — and whether it's the one you expected
Not built for
- Predicting your conversion rate: synthetic intent is directional, not a revenue forecast — anyone claiming otherwise is selling you something
- Replacing real usage data: retention, churn, and pricing elasticity need actual customers
- Truly novel domains: if your market has no footprint of real customer conversation online, we flag the report as low-confidence — visibly
- Fine-grained demographic claims: the research found subgroup fidelity uneven; we report segments by behavior, not by census box
Pricing
Cheaper than one hour of a researcher's time.
One-time payments. No subscription. If the scrape fails or the report can't be produced, it's auto-refunded.
Diagnostic
- 150-stranger panel on one URL
- Intent distribution + rank vs. your industry's scored pages
- Attention heatmap + what each buyer type never sees
- UX friction review, severity-ranked (fix-first list)
- Top 5 objections, quoted & segmented
- Copy fixes suggested per objection
Variant ranking
- 2–4 versions of your page or copy
- Same panel reacts to every variant
- Ranked results — the method's strongest mode
- Per-segment winner breakdown
Competitor scan
- Your page vs. up to 3 competitors
- Where you win, where you bleed
- Objections unique to your page
- Positioning gaps in their copy you can claim
Questions
Asked by people who should be skeptical.
Aren't AI survey respondents just made up?
Naively, yes — ask a model for a 1–5 rating and you get useless, middle-clustered answers. That's exactly what the underlying research demonstrated, and why we don't do it. SSR elicits natural-language reactions and maps them to ratings by semantic meaning. Validated against 9,300 real respondents across 57 surveys, it recovered ~90% of the reliability a repeated human panel achieves. Synthetic panels aren't a replacement for talking to customers — they're a way to arrive at those conversations with a sharper page.
I'm a coach / therapist / agency — is this for me, or just for tech founders?
It's for you — service businesses are now the majority of pages we score. Your page is judged by a fixed panel built around how your buyers decide (booking a call or consultation, not installing an app), and your percentile compares you against scored sites in your own industry. Start from your industry page — coaches, agencies, therapists — and the whole flow speaks your language.
How do you know which audience to use?
We read your page and suggest an industry lane, then stop and show it to you. You can switch lanes before a single persona is generated. The scored panel is fixed for that lane, so your percentile compares against peer pages judged by the same buyer types.
Why won't you give me a predicted conversion rate?
Because the method can't honestly deliver one, and we'd rather be trusted than impressive. The validated strength of SSR is relative measurement — rankings, percentiles against a benchmark corpus, and the qualitative why. Synthetic intent distributions are systematically wider and slightly more critical than human ones, which makes them great discriminators and bad absolute forecasters.
What if my product is in a really niche space?
The method works because models have absorbed enormous amounts of real customer conversation about most consumer domains. If your domain has thin coverage — deep tech, novel B2B categories — synthetic reactions get less trustworthy. We detect this and put a low-confidence banner on the report rather than hiding it. If a report is flagged and you don't find it useful, ask for a refund.
What do you do with my page and data?
We screenshot and extract copy from the URL you give us, run the analysis, and store your report so you can revisit it. We don't train models on your pages, don't resell your data, and only ever test publicly accessible URLs you submit.
Has this been tested on itself?
Constantly. This page's structure, copy, and even its price were chosen by running variants through our own panel — the $29-vs-$49 pricing test came back identical to three decimal places across two independent 150-persona runs. We publish self-tests as we run them: see the 731-site coaching teardown in Guides. A method that can't survive its own scrutiny doesn't deserve yours.
Five minutes from now
You could know what 150 strangers think.
Or you could keep guessing, spend on traffic, and find out from a flat conversion chart three weeks from now.
$49 · auto-refund if we can't produce your report