How to Evaluate a Shopify CRO Partner in 30 Minutes: 8 Scoping Signals
You have three CRO agencies on a shortlist and about an hour booked with each. By the second call, the decks have started to blur. Same promise to lift revenue per session, same mention of a rigorous testing process, same two case studies with a green arrow pointing up and to the right. The proposals...
Last updated: 24 Jul 2026
CONTENTS
You have three CRO agencies on a shortlist and about an hour booked with each. By the second call, the decks have started to blur. Same promise to lift revenue per session, same mention of a rigorous testing process, same two case studies with a green arrow pointing up and to the right. The proposals will differ by a few thousand euros a month and very little else. What none of the decks will tell you, and what each was built to smooth over, is whether the people on the call have ever located a real conversion leak or only ever run tests someone else prioritized for them.
That distinction is the whole game, and you have thirty minutes to read it.
Why the first scoping call reveals more than the pitch deck
The scoping conversation is a sharper signal than the case-study deck for one reason: the deck is rehearsed and the scoping is not. A capable CRO partner diagnoses where your store loses money before proposing what to test. An agency that arrives with a backlog of “quick wins” before it has seen your funnel is selling tests, not outcomes, and the first conversation is where that instinct shows.
A deck is a curated artifact. It shows you the engagements that worked, with the losing tests and the stalled programs quietly left out, and it has been polished across dozens of pitches before it reached you. You learn what the agency wants you to see. The scoping call is different because it runs in real time. When you describe your store and the person on the other end has to respond without a script, you see how they actually think about a conversion problem. Do they reach for a diagnosis, or do they reach for a solution they already had in the drawer?
The strongest operators treat the first thirty minutes as a diagnostic, not a sales demo. They ask where your revenue is leaking before they tell you what they would change, because a test backlog is only worth building once you know which funnel stage is actually costing you. This is the same order of operations that separates disciplined CRO from a pile of A/B tests run in the wrong place: read the funnel, isolate the leak, quantify the recoverable revenue, and only then build a backlog. An agency that skips straight to “we’d test your product page hero” has told you it prescribes before it diagnoses, which is exactly the pattern that burns a retainer proving small things in stages that were never the problem.
So the eight signals below are not questions you interrogate an agency with. They are things you listen for while the conversation happens naturally. Most of them surface on their own if you simply describe your store and let the agency respond.

Signals 1 – 4: Do they understand your store before they pitch?
The first four signals all test one thing: whether the agency treats your store as a specific system to be understood or as a generic template to run a playbook against. They cluster in the opening half of the call, usually before anyone has mentioned price.
Signal 1: They ask where you are losing customers before they name a fix.
Listen to the sequence. An operator wants to understand your funnel before proposing anything, so early questions point at your data: where does traffic drop, what does your checkout completion look like, which pages carry the intent. An agency that names a fix in the first ten minutes (“we’d start by testing your product page hero”) has told you it prescribes before it diagnoses. A fix proposed before the leak is located is a guess with good production values.
Signal 2: They do the traffic math out loud.
A credible partner asks about your monthly sessions and order volume early, because that number decides whether testing is even viable for you. A store running a few thousand sessions a month cannot reach statistical significance on most tests in a reasonable window, and an honest operator will say so directly rather than sell you a testing retainer that produces inconclusive data. If they ask about traffic and then reason about test velocity and how long a result would take to read, they understand the constraint they are working inside. If traffic never comes up, they are quoting you a program without knowing whether it can work.
Signal 3: They ask what winning means to your P&L, not just your conversion rate.
Conversion rate on its own is a vanity metric. A discount popup will lift it while quietly eroding your margin, and a test that wins on completion can lose you money per order. The agencies worth hiring ask what actually needs to move: revenue per session, average order value, contribution margin, or repeat rate. If the only number they anchor to is conversion rate, they are optimizing the metric that is easiest to move rather than the one that pays your salaries.
Signal 4: They ask about your Shopify setup specifically.
There is a real difference between an agency that knows CRO and one that knows CRO on Shopify. The Shopify-native ones ask what theme you run, what sits in your app stack, whether you are on Plus, and whether your checkout has moved to extensibility yet. They ask because every recommendation they make has to survive contact with your actual store: a test that breaks a subscription app or conflicts with a theme customization is not a win, it is an incident during a sale period. A generalist who offers best practices without asking what they will be building inside has not accounted for the part of the work that decides whether anything ships.
Four signals in, you already know whether the agency is reasoning about your store or reciting a service. The next four tell you whether the machine behind the pitch actually holds together.
Signals 5 – 8: Is their testing a system you can trust?
The first four signals tell you whether the agency understands your store. The next four tell you whether the engine behind the pitch is a real system or a set of tactics with a testing tool bolted on.
Signal 5: They can explain where a hypothesis comes from.
Ask, or wait for them to volunteer, how they decide what to test. A weak answer describes tools: heatmaps, session recordings, an analytics dashboard. Those are inputs, not a method. A strong answer describes a sequence: quantitative data shows what is underperforming, qualitative research explains why, and a hypothesis is written from both before anything is built. Operators anchor their hypotheses in behavioral evidence and established research, the kind of large-scale checkout findings Baymard Institute has documented across hundreds of sites, rather than a hunch about your hero image. If they cannot describe how a test earns its place in the queue, there is no queue, only a backlog of opinions.
Signal 6: They tell you which tests probably will not win.
This is the intellectual-honesty signal, and it is the fastest one to read. An experienced operator will tell you, unprompted, that most tests do not produce a winner, that a healthy program treats a flat or losing result as information, and that the value compounds across cycles rather than arriving in a single hero test. Anyone promising you a specific win rate or a guaranteed lift is describing a sales target, not a testing reality. The honesty costs them something in the room, which is exactly why it is worth hearing.
Signal 7: They are clear about who ships the change and how it survives your stack.
A recommendation that never reaches production is a slide, not a result. Ask who writes the code, who does the QA, and how a variant gets tested without breaking your theme or conflicting with an app during a live sale. Shopify-native teams have an answer ready because they have shipped inside real stores and know where it goes wrong. A team that hands you a document and expects your developers to interpret it has quietly moved the hardest part of the work onto your side of the table, and that is where most programs stall.
Signal 8: They will tell you when you do not need them yet.
This is the strongest signal, and the rarest. An operator who asks about your traffic, your tracking setup, and your internal capacity will sometimes conclude that a retainer is the wrong move right now: that you are below the traffic a testing program needs, or that your obvious friction points should be fixed before anyone runs a formal experiment. An agency willing to scope itself out of an engagement is telling you it optimizes for the right decision rather than the signed contract. That is the same instinct you want pointed at your store once the work begins.
Eight signals, and none of them requires you to audit a case study or check a reference. They surface in how the agency reasons about your store, live, while you are still deciding whether to book a second call.

How to run the 30 minutes
You do not need to work through eight questions in order. Most of the signals surface from three prompts, as long as you resist the urge to fill silences and let the agency reason out loud.
Open by describing your store and your problem in plain terms, then stop talking. Something like: “We do around 80,000 sessions a month on Shopify Plus, revenue per session has been flat for two quarters, and we are not sure where the drop is.” Then watch what they reach for. An operator responds with questions that locate the leak. A test-monkey responds with a fix. That single exchange covers Signals 1, 2, and 4 at once.
The second prompt is simply: “How would you decide what to test first?” This surfaces Signals 3, 5, and 6 together, because a real answer has to name a metric that matters, a method for generating hypotheses, and an honest account of how often tests win. The third prompt closes it out: “What does the first ninety days look like, and who does what?” That question exposes Signals 7 and 8, since a grounded answer covers implementation ownership and, from the strongest partners, a candid read on whether you are ready at all.
Across those three exchanges, the difference between an operator and a test-runner tends to sound like this:
| What you ask | Operator answer | Test-runner answer |
| Where should we start? | “We would locate the leak before proposing tests.” | “We would start testing your product page.” |
| How do you pick tests? | “Data finds the problem, research explains it, then we write the hypothesis.” | “We have a backlog of proven best practices.” |
| How often do tests win? | “Often they do not. The program compounds over cycles.” | “Our tests average a strong lift.” |
| Who ships the change? | “We build and QA inside your theme and app stack.” | “We hand over recommendations for your team.” |
None of the operator answers require a rehearsed deck. They come from a way of thinking about conversion that either exists or does not, and thirty minutes of unscripted conversation is enough to hear which one you are dealing with. The scoping call is not the preamble to the evaluation. It is the evaluation.

Frequently Asked Questions
How long before a Shopify CRO program shows results?
Directional gains often appear within a few weeks when implementation velocity is high, but meaningful, compounding results usually take two to three months. Small checkout or product-page changes can move a number quickly, while larger experimentation programs need several test cycles to reach statistical significance. Anyone promising a step-change in the first month is describing a sales pitch, not a testing timeline.
How much traffic do I need before hiring a CRO agency?
As a working floor, roughly 10,000 monthly sessions. Below that, most A/B tests take too long to reach significance, and you end up paying agency rates for inconclusive data. A credible partner will ask about your traffic early and tell you honestly if you are not yet at a volume where formal testing pays back. In that case, fixing obvious friction first is the better use of money.
How much does a Shopify CRO agency cost?
Retainers vary widely by scope and store size, commonly ranging from a few thousand to five figures a month. What matters more than the number is what sits inside it: research, hypothesis generation, build and QA, and reporting, versus a thinner audit-and-recommendations model. Ask what is included and who does the implementation, because a lower fee that excludes shipping the work often costs more once your team absorbs the build.
Do I need an ongoing program or just a one-off audit?
A one-off audit tells you where you are losing money. An ongoing program is what recovers it. Audits are useful when you need a diagnosis before committing, but conversion gains come from the test-learn-iterate cycle, which a single deliverable cannot provide. If a store is early or budget-constrained, a diagnostic first, then a program once the leaks are known, is a reasonable sequence.
Should I build CRO in-house or hire an agency?
It depends on volume and internal capacity. High-traffic stores with a dedicated growth team can run experimentation in-house once the discipline is established. Most mid-market brands lack the combined research, analytics, and Shopify development bandwidth to sustain a testing cadence, which is where an agency earns its retainer. The deciding factor is whether someone internal can own implementation and act on findings quickly.
Key Takeaways
The evaluation you actually need happens in the first conversation, not in the deck or the reference calls that follow. A rehearsed presentation shows you curated outcomes; an unscripted scoping call shows you how the team thinks about a conversion problem in real time.
- The core tell is sequence: operators diagnose where you lose money before they prescribe what to test, while test-runners arrive with a backlog before they have seen your funnel.
- The eight signals surface naturally from a plain description of your store. You are listening for how they reason, not interrogating them with a checklist.
- The strongest signal is an agency willing to tell you when you are not ready, or when a fix does not need a retainer at all.
- Three prompts cover most of it: describe your store and stop talking, ask how they would decide what to test first, then ask what the first ninety days look like and who does what.
If you are shortlisting CRO partners and want a vendor-neutral baseline to scope against, Flatline runs a conversion leak diagnostic that locates where revenue is actually draining before any test backlog gets built. It is the same diagnostic-first discipline these signals are built to detect, and it works just as well as a reference point for evaluating anyone else on your list. If that would be useful, get in touch and we will walk through it with you.
POPULAIR ARTICLES
GET IN TOUCH
To speak with us, call (+31) 613 326 179, send us an email, or reach out to us by chat or What’s App.