The Deliberate Experimentation System

A better CRO process starts before the experiment.

Most testing programs jump straight to building. We research first, diagnose the real friction, and turn evidence into a prioritized roadmap before a single experiment goes live.

RESEARCH-LED  ·  EVIDENCE OVER OPINION
01 — The system

Seven steps. One loop that never stops running.

Every engagement runs the same seven-step system, from the first Audit to the tenth cycle of Ongoing Experimentation.

01Research 02Diagnose 03Hypothesize 04Prioritize 05Experiment 06Analyze 07Scale SCALE FEEDS BACK INTO RESEARCH — THE LOOP NEVER STOPS
01

Research

Understand what users are doing and why. Qualitative interviews, session recordings, and a review of your current analytics, before we assume anything.

  • Stakeholder and customer interviews
  • Session recordings and funnel analysis
  • Tracking and data-layer sanity check
02

Diagnose

Find friction and conversion opportunities. The research gets turned into a ranked list of where the real problems, and the real opportunities, actually sit.

  • Funnel and page-level friction points identified
  • Behavioral data reviewed against the qualitative research
  • Opportunities ranked by where the impact actually sits
03

Hypothesize

Turn evidence into testable ideas. Every hypothesis is framed as if/then, with a clearly predicted outcome, and traces back to a specific finding.

  • Framed as if/then, with a predicted outcome
  • Tied directly to a research finding, never a hunch
  • Written so it can actually be proven wrong
04

Prioritize

Use the PROOF Score. Every hypothesis is scored on potential impact, reach, evidence, confidence and effort, so the roadmap reflects what actually moves revenue.

  • Every idea scored on the same weighted framework
  • Roadmap agreed with you before a single test is built
  • Biggest, most evidence-backed bets rise to the top
05

Experiment

Run statistically disciplined experiments where traffic supports them, and evidence-based rollouts where it doesn't. Never faked statistics or a test called early to hit a deadline.

  • Proper sample-size and duration planning
  • QA across your key browsers, devices and segments
  • No peeking, no early calls, no exceptions
06

Analyze

Understand both outcomes and learnings. Results are read for statistical validity, not just a green metric, and reported honestly, wins, losses and flat results alike. (For stores where the sale doesn't close entirely on-screen, we can extend measurement to the actual outcome, delivery, disbursement, in-store visit, as one example of more advanced measurement; it's not what defines the engagement.)

  • Read for statistical validity, not just a green metric
  • Wins, losses and flat results, reported the same honest way
  • Learnings extracted whether the test wins or not
07

Scale

Implement winners and feed learnings into the next research cycle. Nothing here is a one-off project. Scale closes the loop straight back into Research.

  • Winning variants rolled out fully
  • Learnings folded back into the next round of research
  • The loop starts again, not a one-time redesign
02 — The framework

The PROOF Score: your experiment decision engine.

Score any idea the way we do. Five inputs, one weighted number, so the biggest, most evidence-backed bets rise to the top of the roadmap, not whatever's easiest to ship.

3
3
3
3
3
DECISION ENGINE — LIVE PROOF-SCORE.01
27/ 45
Worth testing

Potential impact, evidence and confidence carry the most weight, so the biggest, best-evidenced bets should always rise to the top.

Effort is scored in reverse — less effort earns more points, not fewer.

This is exactly how we prioritise your roadmap
03 — Why most fail

Why most CRO programs fail.

Not because testing doesn't work. Because the process feeding it is broken in one of six predictable ways.

01

Random testing

Ideas picked by gut feel or whatever's trending. Nothing compounds, because nothing was ever connected.

02

Redesign-first thinking

Rebuilding the page before understanding why it's actually underperforming.

03

Insufficient research

Jumping straight to hypotheses with no research behind them to test.

04

Weak measurement

Tracking that can't reliably tell a real win from statistical noise.

05

Low-impact tests

Testing a button color while the real friction sits untested, unmeasured.

06

Win/loss-only thinking

Every test graded pass or fail, so the losses never get to teach anything.

04 — What you get

Every engagement produces the same seven outputs.

Whether it's a single Audit or a running Ongoing Experimentation program, the system guarantees the same things land in your hands.

Every engagement produces07 items
01

Research findings

What we learned about your shoppers and where they actually hesitate.

02

Hypotheses

Specific, evidence-backed ideas for what to test, and why it should work.

03

Prioritized roadmap

Every hypothesis scored and ranked on the PROOF Score, agreed with you upfront.

04

Experiment designs

Fully specified tests: variants, sample size, duration and the primary metric.

05

Test results

Statistically read outcomes, reported honestly: wins, losses and flat results alike.

06

Learning library

Every result captured in one place, so nothing gets re-tested by accident.

07

Next experiments

Learnings folded straight into the next round of research. The loop, not a one-off.

05 — Questions

Honest questions about the method itself.

There are plenty of those already. What's different is that every score is tied straight back to potential impact and the evidence behind it, weighted the heaviest, and the roadmap it produces is agreed with you before a single test is built, not adjusted afterwards to justify what we felt like shipping.

Below roughly 1,000 conversions a month on the page in question, we won't fake statistical significance. We switch to research-led, evidence-based rollouts instead, the same disciplined system, just without pretending a small sample is a controlled experiment.

No. We don't call tests early and we don't hide the ones that lose or come back flat. Every scorecard reports all three outcomes, every cycle, that's the whole reason the system is trustworthy enough to prioritise real budget against.

Because a good idea that takes a month to ship isn't automatically better than an equally good idea that takes a day. The Decision Engine rewards the easier win, all else being equal, so effort pulls the score down instead of up.
Start the conversation

Ready to put a real roadmap through the system?

Start with a discovery call, or the CRO Audit if you're ready to go straight to a prioritized, PROOF-scored roadmap.

Get CRO Audit