The CRO engagement that goes wrong rarely fails because nobody can think of a new headline or button color. It fails because the buyer and vendor were buying two different things.

One side thought it was buying a research-and-experimentation system that would diagnose friction, prioritize opportunities, ship changes and measure incremental value. The other side was staffed and priced to deliver audits, design ideas and a monthly slide deck. Both could honestly call the service “conversion rate optimization.” The conflict only becomes visible after the contract is signed.

That is why the best vendor checklist does not start with “How many tests can you run?” It starts with what evidence will change a decision, who does the implementation, what happens when the data is weak, and which guardrails can veto a apparent win.

Baymard’s current checkout research is a useful reminder of why diagnosis matters. Its cross-study cart-abandonment benchmark sits at 70.19%, but a large share of abandonment is natural browsing rather than a defect a vendor can remove. The existence of a big industry number does not tell you what your own bottleneck is. A credible CRO partner should be able to say “we do not know yet” and show how it plans to find out.

Use the questions below before procurement, not after the kickoff.

1. What exactly are you responsible for after the audit?

Start by marking every step the vendor owns:

  • instrumentation review;
  • qualitative research;
  • analytics diagnosis;
  • hypothesis generation;
  • design;
  • copy;
  • engineering;
  • QA;
  • experiment configuration;
  • statistical review;
  • launch;
  • winner implementation;
  • post-test documentation.

If implementation belongs to your team, estimate the hours before signing. A strong answer names handoffs, owners and cycle time; “we provide actionable recommendations” is not enough.

2. Show us a project where the first hypothesis was wrong

A case study that only contains wins tells you almost nothing about the operating model.

Ask for a project where the team believed one cause was important, collected evidence, discovered it was not, and changed direction. CRO is valuable partly because it reduces confidence in bad ideas before the company scales them.

You are looking for evidence that the vendor can abandon its own narrative. If every diagnosis conveniently validates the proposal it already sells, the research function may be decoration.

3. What evidence is required before an idea enters the backlog?

Good programs have an admission rule.

For example, a hypothesis might need two of the following: funnel anomaly, customer interviews, usability evidence, support tickets, search behavior, session evidence, competitive constraint, or a known defect. The exact mix varies by company.

The point is to avoid an idea factory where seniority, enthusiasm or visual novelty becomes prioritization.

Ask to see a real hypothesis brief with the sensitive details removed.

4. How do you distinguish a defect from an experiment?

Broken validation, missing shipping information, inaccessible controls, tracking failures and obvious mobile bugs do not always need A/B tests. Some should simply be repaired.

A vendor paid per test can be economically encouraged to turn every problem into an experiment. Ask for the rule that separates:

repair — known failure with a clear expected behavior;
research — cause is disputed;
experiment — two or more plausible solutions remain.

This one question can save weeks of fake rigor.

5. What happens when traffic is too low for the proposed test?

Do not accept “we can still test it” as a complete answer.

Low eligible traffic, small effect sizes and narrow segments can make a test impractically slow. A serious partner should have alternatives: larger changes, qualitative validation, sequential research, broader eligibility, holdout measurement where appropriate, or simply a decision based on risk and reversibility.

Ask for one example where the vendor recommended not running an experiment.

6. Which metric can declare a win, and which metrics can veto it?

Conversion rate is not enough.

A treatment might increase checkout completion while also increasing returns, discounts, support contacts, involuntary subscriptions or low-quality leads. Before launch, the brief should identify the primary outcome and the guardrails.

A good procurement answer sounds like: “Revenue per eligible session is primary; refund rate, gross-margin contribution, page performance and support complaints are guardrails.”

The exact measures differ, but the principle should be explicit.

7. How will you protect the customer from manipulative optimization?

This is not an abstract ethics question. It is an operating risk.

The U.S. Federal Trade Commission has repeatedly described “dark patterns” that obscure material terms, make cancellation difficult, hide fees or steer people into choices they did not intend. Ask the vendor for its rejection criteria for urgency claims, defaults, add-ons, subscription flows, consent and cancellation.

The safest answer is not “our legal team checks it.” The safest answer is a design and experiment process that flags the issue before launch.

8. What accessibility standard is in the acceptance criteria?

W3C’s WCAG 2.2 provides testable success criteria for web accessibility and recommends using the current WCAG version when developing or updating accessibility policies. CRO work changes forms, focus behavior, error messages, controls and page structure—the same surfaces that can create accessibility regressions.

Ask whether accessibility is:

  • part of design review;
  • part of front-end QA;
  • tested with automated tools only or also manual methods;
  • preserved when experiment variants are injected;
  • documented when a third-party testing or personalization script changes the DOM.

Do not treat a vendor’s statement as a legal compliance certification unless the contract actually defines one.

9. What performance budget must a variant stay inside?

A treatment can “win” in a lab and make the real site slower.

Google’s Core Web Vitals are field-oriented measures of loading, interaction responsiveness and visual stability. Ask which real-user performance signals the team watches and what happens if an experiment degrades them.

The useful question is not “Do you optimize page speed?” It is: what threshold, comparison or rollback condition prevents a variant from buying conversion at the cost of user experience?

10. Who owns the experiment data, event definitions and raw exports?

A program becomes fragile when the vendor owns the only readable version of history. Confirm that you retain event definitions, experiment IDs, result tables, research notes, designs, code and the decision log. If the relationship ends, the company should not lose its institutional memory.

11. How do you handle contradictory tools?

When analytics, the payment processor and research disagree, ask which system defines orders, refunds and time zones, and how bots, staff sessions, duplicates and consent-related gaps are treated. A glossy dashboard is not evidence that the definitions agree.

12. What is your stopping rule for an experiment?

Ask before seeing the platform demo.

You want to hear about minimum runtime, expected sample, decision thresholds, novelty effects, peeking discipline, seasonality and business context. The vendor does not need to worship one statistical school, but it should be able to explain its method without hiding behind “the tool says significant.”

Also ask what happens to an inconclusive test. Mature programs keep the learning; immature programs quietly rename it and run it again.

13. How much implementation capacity do you assume from us each month?

Turn the answer into hours and roles. Put expected engineering, design, analytics and review time beside the agency fee. This exposes the real operating cost and prevents an internal queue from being misdiagnosed as vendor slowness.

14. How are winners shipped permanently?

An experiment winner still running inside a testing platform months later is technical debt.

Ask for the winner-to-production workflow: ticket creation, owner, deadline, code removal, QA, analytics verification and experiment cleanup. The same applies to losers—temporary scripts and abandoned variants should not accumulate forever.

The contract should reward realized changes, not just completed tests.

15. Show us your last five post-test summaries, including losers

Client names can be removed. Look for four things: the original uncertainty, what happened, the next decision, and what can be reused elsewhere. If reports celebrate uplift but rarely explain uncertainty, the program may be optimized for renewal decks rather than learning.

16. What should make us fire you after 90 days?

A credible vendor can name failure conditions: untrustworthy instrumentation, no shipped changes, chronic implementation delays, unresolved data discrepancies or a backlog without evidence. Then ask what it needs from you to prevent them.

17. What will you refuse to do even if we ask?

The answer reveals operating maturity.

A good CRO partner should be willing to refuse deceptive scarcity, hidden fees, misleading defaults, fabricated social proof, inaccessible variants, tests with no interpretable outcome, or measurement tricks that inflate the dashboard while worsening the business.

That refusal is part of the product.

A scorecard that is harder to game

Do not score vendors on presentation quality alone. Use weighted categories:

Category Weight What a high score looks like
Diagnosis and evidence quality 20% Clear research methods and evidence thresholds
Implementation ownership 20% Explicit handoffs, cycle times and capacity
Measurement discipline 20% Primary metric, guardrails, exposure logic, decision rules
Customer and compliance guardrails 15% Dark-pattern, consent, accessibility and claim review
Technical quality 10% Performance budgets, QA, clean experiment lifecycle
Knowledge transfer 10% Exportable data, reusable learnings, decision archive
Commercial fit 5% Pricing matches the actual work and internal capacity

The scorecard matters less than the conversation it forces. A strong vendor can still be wrong for your current stage. Procurement should select the operating model, not the longest tool list.

The contract should contain the answer to four questions

Before signature, make four things unambiguous:

What is the unit of delivery? Research decision, shipped change, experiment, learning cycle, or hours?
Who ships? Name the owner on both sides.
How is success judged? Use an outcome with guardrails, not “number of ideas.”
What can stop the work? Data integrity, customer harm, accessibility or performance regression, legal concern, and insufficient traffic must be legitimate pause conditions.

Clear answers turn the relationship into a learning system instead of a recommendation subscription.

Bottom line

The right CRO partner is not the one that promises the most tests. It is the one that can tell you what it knows, what it does not know, how it will find out, who will ship the change, and what evidence would make it stop.

Buy that operating discipline first. Tools, dashboards and experiment counts come second.

Sources

Related Reading