How to Choose a CRO Agency: 9 Criteria, Questions and Red Flags
Every CRO agency pitch sounds the same. Data-driven. Results-focused. Structured experimentation. Case studies with impressive percentages.
None of that distinguishes a rigorous partner from an agency that ships design changes and calls it optimisation. What distinguishes them is how they answer specific questions about methodology and most brands don't know which questions expose the difference.
Here are nine criteria for choosing a CRO agency, each with the question that reveals the truth and the red flag that should end the conversation.
1. They Validate Your Analytics Before Testing
Start here, because almost no agency does it and it determines whether everything else matters.
CRO measures success against your conversion data. If your GA4 purchase event fires twice for a share of orders, your baseline conversion rate is inflated — which corrupts every sample size calculation and produces tests that appear significant before they are. If begin_checkout isn't firing on mobile, your funnel shows a collapse that isn't real, and hypothesis prioritisation goes to the wrong step.
Ask: "Will you audit our GA4 tracking before building a test roadmap, or will you work from our existing setup?"
Red flag: "We'll use your existing data." It means they're building on a foundation nobody has verified. When results don't hold in production three months later, this is usually why. Our account of A/B testing on broken GA4 data covers what that failure looks like in practice.
Good answer: They describe a validation step — reconciling GA4 against your ecommerce platform, checking for duplicate events, confirming funnel events fire across devices before the first hypothesis.
2. Research Comes Before the First Test
The single most revealing question in an agency evaluation is asking them to walk you through what happens before test one launches.
Weak agencies describe tools: heatmaps, session recordings, Google Analytics. Strong agencies describe a methodology with named stages and specific deliverables; quantitative funnel analysis to locate where revenue leaks, qualitative research to understand why, then hypothesis development from those findings.
Ask: "Walk me through your research process before the first test launches."
Red flag: A tool list instead of a process. Any competent team can operate VWO or Optimizely. The tool isn't the capability.
3. They Use a Structured Prioritisation Framework
Testing capacity is limited by your traffic. Which tests you run first determines how much value the programme produces in its first two quarters.
Credible agencies use ICE (Impact, Confidence, Ease) or PIE (Potential, Importance, Ease) scoring to rank the backlog. The framework matters less than having one.
Ask: "How do you decide which test runs first?"
Red flag: Intuition, or "whatever looks most interesting." That's not prioritisation, it's preference.
4. Fixed Significance Thresholds and No Peeking
Statistical discipline is where most programmes quietly fail. Tests called early because one variant looks ahead on day six, produce results that don't replicate.
Ask: "What confidence threshold do you use, and how do you calculate sample size before launch?"
Red flag: Running tests for fixed calendar durations without sample size calculations, or checking results daily and stopping when significance appears. Both produce false positives. A rigorous agency calculates the required sample from your baseline conversion rate and minimum detectable effect, commits to it, and doesn't act until it's reached.
5. An Honest Win Rate
Counterintuitive but important: a high win rate is a warning sign, not a selling point.
Industry win rates for rigorous programmes sit at 20–30%. Two or three tests in ten produce significant positive results. An agency claiming 60%+ is usually peeking at results, testing only safe changes that can't lose, or measuring against an inflated baseline.
Ask: "What's your typical win rate, and can I see a test log that includes the losers?"
Red flag: A win rate above 50%, or an unwillingness to show tests that didn't work. Case studies show winners. A full test log shows methodology.
6. They Report Revenue, Not Vanity Metrics
Click-through rate, scroll depth, and time on page have diagnostic value. They don't pay salaries.
A credible agency ties success to metrics that connect to money: conversion to purchase, revenue per visitor, average order value, and for Indian D2C specifically delivered order rate, since a checkout optimisation that increases RTO-prone COD orders can improve reported conversion while reducing realised revenue.
Ask: "What will your work do for our revenue, and how will you measure it?"
Red flag: Success defined by percentage lift in conversion rate without reference to absolute revenue, or an agency that treats revenue metrics as out of scope.
7. Category-Specific Experience
The patterns that convert a ₹800 skincare buyer are different from those that convert a ₹15,000 furniture purchase, and both differ from SaaS or lead generation. Most agencies apply one playbook across every vertical.
For Indian D2C, this extends further: COD versus prepaid buyer psychology, UPI failure handling, tier-2 city purchasing behaviour, and RTO dynamics are market-specific realities that global CRO playbooks don't address. Our post on the Indian D2C analytics problem covers why this matters at the measurement layer too.
Ask: "How many of your test hypotheses come from research on our brand versus patterns from other clients?"
Red flag: A pitch that could have been delivered to any ecommerce brand in any category.
8. Realistic Timeline and Test Velocity Expectations
Test velocity is capped by your traffic, not by the agency's ambition. A store with 20,000 monthly sessions at 1.5% conversion can support roughly three to four tests per quarter, not eight per month.
Ask: "Based on our traffic and baseline conversion rate, how many tests can we realistically complete per quarter?"
Red flag: A number pulled from a package tier rather than calculated from your data. If the answer doesn't involve your actual conversion volume, they haven't done the maths and you'll be paying for testing capacity your traffic can't use. Our post on why CRO results take longer than brands expect covers the underlying arithmetic.
9. You Know Who Actually Does the Work
There's a meaningful difference between an account manager relaying updates and a practitioner with hands-on testing experience running your programme.
Ask: "Who will be our day-to-day analyst, how much testing experience do they have, and how many other accounts do they run?"
Red flag: Senior people in the pitch who disappear after signing. Ask to meet the person who'll actually be doing the work, in the sales process.
Red Flags That Should End the Conversation
Beyond the criteria above, four answers warrant walking away:
Guaranteed uplift percentages. Anyone promising a specific lift before seeing your data either doesn't understand testing or is misrepresenting it. Outcomes depend on traffic volume, baseline, seasonality, and variables invisible on a sales call. The only legitimate guarantee is a process guarantee, a committed number of tests in a defined period.
"We'll redesign your site." A redesign isn't CRO. It's a creative project in optimisation clothing, and it destroys the baseline you'd need to measure whether anything worked. Good agencies test within your existing site and recommend a redesign only after data shows the structure is fundamentally broken.
Refusal to touch checkout or pricing. These are where ecommerce revenue is won and lost, and they're also politically sensitive. Agencies that steer exclusively toward low-risk page elements are protecting themselves, not optimising your store. Our post on signs your PDP or checkout needs a CRO audit covers what should be in scope.
Tool-led pitches. If the conversation is about which platform they use before it's about where your revenue is leaking, the sequence is backwards.
The Scorecard
Weight these unevenly. Analytics validation and statistical rigour matter more than the rest, because they determine whether any result is trustworthy.
Criterion | Weight |
|---|---|
Validates analytics before testing | High |
Fixed significance thresholds, no peeking | High |
Research process before first test | High |
Reports revenue, not vanity metrics | High |
Structured prioritisation framework | Medium |
Honest win rate | Medium |
Realistic velocity expectations | Medium |
Category-specific experience | Medium |
Named, experienced practitioner | Medium |
Score every candidate on all nine. The agency that answers all of them live, without reaching for a deck, is usually the one whose tests still look like wins three months after they ship.
Before You Evaluate Anyone
One check to run first, because it changes which agency you need.
Pull your GA4 purchase event count and your Shopify order count for the same 30-day window. If they differ by more than 10%, your conversion data has a problem and a CRO agency that doesn't fix it will spend your first quarter optimising against a number that isn't real. Our guide to why GA4 and Shopify numbers don't match explains what each gap size means.
Full disclosure: FunnelFreaks is a CRO agency, so score us on this list too. The reason criterion one sits at the top is that it's the phase we built the agency around, we validate and rebuild analytics infrastructure before any hypothesis is written, because every decision after that point inherits whatever errors exist in the data. Every recommendation we make is data-backed, which only means something if the data has been verified first.
If you're comparing options, our breakdown of CRO agencies for D2C brands in India covers who specialises in what and applying these nine criteria to all of them, including us, is the right way to use it.
Want to run this checklist on us? Talk to FunnelFreaks; ask all nine questions, and we'll answer them without a slide deck.