There is no single thing called “doing CRO.” A retailer can hire a research team to diagnose friction, commission a redesign, build a continuous experimentation program, or add personalization that changes what different visitors see. All four can improve conversion. They differ sharply in how fast they produce evidence, how much they cost to operate, how much control the company keeps, and what can go wrong.

That distinction matters because buyers often compare unlike proposals. A two-week research sprint can look “small” next to a redesign. An experimentation platform can look “technical” next to a UX audit. A personalization product can promise speed because it creates many variants, even though the organization may still lack a clean answer to a simpler question: which problem is worth changing?

The useful choice is not “which CRO approach is best?” It is which approach matches the uncertainty you have right now.

Baymard’s checkout research is a good reminder that the opportunity can be real without being universal. Its current research reports a roughly 70% cart-abandonment rate across the datasets it tracks and says 17% of surveyed U.S. shoppers had abandoned an order because checkout felt too long or complicated. Those figures are research context, not a forecast for any one store. They tell you there are recurring usability problems; they do not tell you whether your own bottleneck is checkout, product fit, price, traffic quality, trust, or something else.

The comparison in one page

Approach Best when Speed to useful evidence Fixed cost Ongoing cost Control Main risk
Research and diagnostic sprint You do not yet know why people hesitate Fast Low to medium Low High Producing a report that never ships
Focused redesign The broken journey is known and needs coherent change Medium Medium to high Low to medium High Changing too much at once and losing causal clarity
Continuous experimentation You have enough traffic, analytics and shipping capacity to learn repeatedly Medium to slow at first, then recurring Medium to high Medium to high High Running many tests without business relevance
Personalization / adaptive optimization Different audiences clearly need different experiences Medium Medium to high Medium to high Medium Complexity, weak holdouts, privacy or “black box” decisions

These are relative operating bands, not vendor quotes. A small in-house team, a specialist agency and an enterprise platform can put very different prices on the same label.

Route 1: research before you build

A research-first sprint combines analytics review with customer interviews, session evidence, usability testing, support-ticket analysis, funnel diagnostics and a prioritized problem list. Its output should not be “forty ideas.” It should be a small number of decisions with evidence behind them.

This route is strongest when the organization is arguing about causes. Marketing thinks shipping cost is the problem; product thinks the mobile flow is confusing; merchandising thinks the assortment is wrong; leadership thinks the button color is wrong. Research is valuable because it changes the conversation from opinions to evidence.

It is also relatively reversible. You can inspect a checkout, prototype a new sequence or interview customers without committing the entire site to a new architecture.

The failure mode is familiar: the team delivers a polished deck, everybody agrees, and nothing gets implemented. Before buying the work, name the person who owns implementation, the engineering capacity available after the study, and the deadline for the first shipped change.

Choose this route when your largest uncertainty is diagnosis.

Route 2: redesign when the experience is structurally wrong

A redesign is appropriate when problems are connected. If product discovery, pricing communication, cart behavior, checkout and post-purchase expectations all conflict, fixing one control at a time can preserve the underlying mess.

A redesign can simplify the information architecture, remove inconsistent patterns, improve accessibility, modernize components and clean up measurement. It can also create a better technical base for later experiments.

But redesigns create a different risk: they bundle many changes. If conversion rises, you may not know which decisions caused it. If conversion falls, debugging becomes harder because navigation, copy, layout, performance and tracking may have moved together.

Performance is a concrete guardrail here. Google’s current Core Web Vitals guidance treats LCP, INP and CLS as field metrics for loading, responsiveness and visual stability, with recommended “good” thresholds of LCP within 2.5 seconds, INP at 200 milliseconds or less, and CLS at 0.1 or less at the 75th percentile. Those numbers are not conversion guarantees, but they are useful release constraints: a visually impressive redesign that makes the real user experience materially slower has created a new problem while solving an old one.

A redesign therefore needs a before-state benchmark: funnel conversion, retained orders or qualified leads, device split, key performance metrics, error rates and support signals. Otherwise the team launches a new site and loses the ability to tell whether it is actually better.

Choose this route when the journey is known to be structurally inconsistent and isolated fixes would leave the system broken.

Route 3: experimentation when the organization can learn continuously

A/B testing is attractive because it can estimate causality more credibly than a simple before/after chart. But the method is not the program.

A durable experimentation system needs:

  • enough eligible traffic for the decisions being tested;
  • instrumentation the team trusts;
  • a hypothesis intake process;
  • statistical discipline and stop rules;
  • engineering or content capacity to build variants;
  • guardrail metrics;
  • a way to ship winners and document losers;
  • a business metric beyond click-through rate.

The hidden cost is organizational throughput. A company can buy an experimentation platform in a day and still be unable to run a meaningful experiment for weeks because events are unreliable, product teams cannot prioritize variants, legal review arrives late, or nobody owns the final implementation.

The hidden benefit is cumulative learning. One test may produce a small result; twenty well-chosen tests can change how a team thinks about shipping, measurement and uncertainty.

This route fits businesses with recurring traffic and recurring decisions. It is less attractive when the site has too little traffic for narrowly segmented tests or when obvious defects should simply be fixed.

The key procurement question is: how many high-value decisions can we resolve per quarter, not how many experiments can the tool launch?

Route 4: personalization when “one experience” is genuinely the wrong assumption

Personalization can mean anything from simple rule-based messages to dynamic ranking, recommendation systems or AI-selected content. The appeal is clear: the site can adapt to visitor context rather than forcing every user through the same sequence.

It earns its complexity when user groups have meaningfully different jobs to do. A returning wholesale buyer, a first-time consumer, a customer with an expiring subscription and a visitor arriving from a product-specific campaign may need different information.

The danger is that personalization can become a layer of untestable exceptions.

A mature program therefore needs:

  1. a clear eligibility rule for each variation;
  2. a default experience that remains coherent;
  3. holdout traffic that sees the default;
  4. a business metric and guardrails;
  5. a record of what data the decision uses;
  6. a fallback when data is missing or wrong;
  7. periodic deletion of rules that no longer add value.

If the team cannot explain why a user saw a particular experience, who owns the rule, and how the effect is measured, the system is gaining complexity faster than it is gaining conversion.

Choose this route when audience differences are proven and you can preserve measurement and governance.

Do not let a conversion tool create a compliance problem

CRO work frequently touches urgency messages, recurring subscriptions, pricing disclosures, add-ons, cancellation flows and consent. Those are not just “UX levers.”

The U.S. Federal Trade Commission’s dark-pattern guidance and enforcement actions have repeatedly focused on designs that obscure material terms, make cancellation difficult, hide fees or steer people into choices they did not intend. That creates a practical boundary for optimization teams: a variation that increases completion by making the true cost or the exit path harder to understand is not a clean win.

A sensible experiment brief includes one sentence that is rarely shown on dashboards:

What would make this treatment unacceptable even if the primary metric improves?

Examples include more refund requests, more involuntary subscriptions, lower consent quality, higher support volume, worse accessibility or a material degradation in page performance.

How to choose without buying the wrong layer

Use this sequence:

If you do not know the problem, buy research.
Do not buy a testing engine to generate evidence about random ideas.

If the problem is structural, redesign a bounded journey.
Avoid “redesign the entire brand” unless the business case actually requires it.

If you repeatedly face uncertain choices and have enough traffic, build experimentation discipline.
The platform is secondary to the operating system around it.

If different audiences consistently need different experiences, add personalization last.
Start with rules you can explain and holdouts you can measure.

A company can also combine these routes. A practical sequence might be research → focused redesign → controlled experiments → selective personalization. What matters is that each layer earns its place.

A procurement scorecard

Before signing a CRO contract or platform order, score each candidate on eight questions from 1 to 5:

Question Why it matters
Does this approach address our current uncertainty? Prevents buying the wrong layer
Can we measure the outcome with data we trust? Avoids false wins
Can we implement what we learn? Research without shipping has zero realized value
How reversible is the change? Limits downside
What ongoing team capacity does it require? Exposes hidden operating cost
Does it preserve performance, accessibility and consent quality? Protects guardrails
Can we isolate or estimate incrementality? Separates causation from correlation
Is there a clear stop/review date? Prevents permanent spend by inertia

The approach with the highest “feature count” should not automatically win. The best route is the one that resolves the most expensive uncertainty with the least unnecessary complexity.

A fifth option is to fix obvious defects without calling it a CRO program

Not every conversion problem deserves a program at all. Broken validation, misleading stock status, missing shipping information, inaccessible controls, duplicate form fields, mobile layout bugs and tracking failures are defects. If evidence is clear, fix them.

This is an important budget discipline because teams sometimes create an experiment merely to avoid making an obvious repair. Testing whether a broken coupon field should remain broken does not create useful learning. The experiment budget is better reserved for choices where reasonable people genuinely do not know which option will produce the better business outcome.

A practical rule is to separate the backlog into three bins: repair, research, and experiment. Repair known failures. Research disputed causes. Experiment when two or more plausible solutions remain. That simple sorting step can reduce the amount of tooling and meeting time required before any customer sees an improvement.

The bottom line

Research, redesign, experimentation and personalization are not competing religions. They are tools for different failure modes.

Use research to learn what is wrong. Use redesign to repair a connected system. Use experimentation to learn which uncertain change causes a better outcome. Use personalization when different users genuinely need different experiences.

And keep the economic test from the previous stage: the winning route is not the one that produces the prettiest conversion chart. It is the one that improves a valuable outcome, survives downstream guardrails, and creates enough incremental contribution to justify the full cost of operating it.

Sources

Related Reading