A DTC founder sees checkout conversion fall and asks for “CRO.” Three proposals arrive.

One agency wants to redesign the product page. An experimentation vendor wants the team to buy testing software. A developer says the real problem is site performance. The paid-media team argues that conversion fell because traffic quality changed. Everyone can be partly right.

That is why conversion rate optimization is easier to buy when you stop treating CRO as a single service. It is a chain of diagnosis, research, prioritization, implementation, experimentation, measurement and operational follow-through. Different sellers occupy different links in that chain. The buyer’s job is to identify the bottleneck before paying someone to optimize the wrong layer.

Here is the market map that matters in practice.

Start with the buyer: what decision are they actually trying to make?

Most DTC teams do not wake up wanting a “CRO program.” They usually have one of five problems:

  • traffic is growing but revenue is not;
  • product-detail pages get visits but weak add-to-cart activity;
  • checkout completion deteriorated;
  • paid acquisition became more expensive, making every conversion point more valuable;
  • the team has many ideas but no reliable way to decide which changes deserve engineering time.

Those are different buying situations.

A low checkout-completion rate may point toward payment, shipping, form or trust friction. Weak product-page performance can come from merchandising, offer clarity, device performance or traffic mismatch. A mature store with decent UX may need a stronger experimentation system rather than another redesign.

Before choosing a vendor, define the decision unit: page, funnel step, audience, offer, checkout, merchandising rule or experiment program. If the problem statement is simply “our conversion rate is low,” the scope is still too vague.

The first seller is often not a seller: analytics and diagnosis

The first layer of the market is measurement.

This includes web analytics, product analytics, session replay, funnel analysis, customer-service data, search logs, returns data and basic financial reporting. The point is not to buy every tool. The point is to establish where the commercial leak appears and whether it is even a UX problem.

A useful diagnostic sequence is chronological:

  1. confirm tracking and revenue integrity;
  2. split performance by device, geography, traffic source, new versus returning customer and major product category;
  3. locate the funnel step where behavior changed;
  4. check whether promotions, inventory, shipping promises, payment methods or site releases changed at the same time;
  5. only then decide what research or experiment is justified.

This layer can be performed internally, by an analytics consultant, by a CRO agency, or by a specialist vendor. Buyers should care less about the label and more about whether the diagnosis can be reproduced.

The research layer converts “something is wrong” into a testable explanation

Once a team knows where the problem sits, it needs evidence about why.

This is where UX researchers, user-testing platforms, survey tools, support-ticket analysis and heuristic reviews enter the chain. Baymard’s checkout research is a useful reminder that checkout friction is not theoretical: its running benchmark continues to show roughly seven in ten carts are abandoned, while its UX studies identify recurring problems such as account creation, form friction and unexpected costs.

But a benchmark is not a diagnosis of your store.

Good research narrows uncertainty. It may reveal that mobile shoppers cannot see delivery timing until late in checkout, that size guidance is inconsistent, that a subscription option is misunderstood, or that visitors arriving from a certain ad expect a different offer than the landing page delivers.

The output should be a prioritized set of hypotheses, not a 90-slide deck of observations.

The solution market splits into four very different types of seller

After diagnosis, the seller landscape branches.

Seller type What they are good at Where buyers get disappointed
CRO/UX agency Research, prioritization, design, experimentation programs When recommendations depend on engineering the agency cannot control
Experimentation platform Traffic allocation, statistics, test governance, feature/web experiments When the buyer lacks enough traffic, hypotheses or implementation discipline
Development/commerce partner Speed, checkout, theme, app and integration changes When technical delivery is mistaken for proof that the change improves conversion
Specialized app/vendor Search, reviews, personalization, payments, upsell, merchandising or another narrow function When a local improvement creates cost, latency, data or UX problems elsewhere

A fifth category is the internal team: product manager, growth lead, designer, engineer, analyst and merchandiser. In many mature businesses, the best CRO “vendor” is a process that coordinates those people rather than outsourcing the whole function.

Implementation is where theoretical lift becomes expensive

A mockup can look persuasive and still be hard to ship.

The implementation market includes theme developers, front-end engineers, checkout specialists, app integrators and platform-native partners. Their cost is not only the invoice. Every change can add technical debt, performance overhead, QA work and maintenance.

This is why performance belongs inside a CRO market map. A 2026 web.dev case study on Nuvemshop/Tiendanube reported that improved Largest Contentful Paint and Core Web Vitals were associated with an 8.9% relative increase in conversion rate for mobile Google organic sessions. That does not mean every speed improvement creates the same lift. It does show that “design” and “technical performance” cannot be treated as separate commercial worlds.

For buyers, the implementation question should be: what will this change cost to build, maintain, monitor and reverse?

Experimentation vendors sell decision infrastructure, not guaranteed winners

A/B testing platforms are valuable because they can make competing experiences measurable. They do not manufacture good hypotheses.

The buyer needs sufficient eligible traffic, stable event tracking, a predeclared decision rule and an implementation that does not contaminate the test. Optimizely’s documentation on sample-ratio mismatch illustrates one reason experimentation has its own operational discipline: unexpected allocation imbalances can signal implementation problems or external interference and should trigger investigation.

An experimentation tool is therefore best viewed as decision infrastructure. It helps answer, “Did treatment outperform control under this design?” It does not answer, “Was this the most important problem to work on?”

That upstream decision still belongs to the business.

The measurement layer decides who gets credit

After a change ships, CRO enters a politically awkward stage: attribution.

Design may claim the win. Engineering may point to faster pages. Paid media may say traffic improved. The promotion calendar may have changed. Inventory may have recovered.

A credible program separates three questions:

  • Did the target behavior move?
  • Did the intervention cause enough of that movement to justify confidence?
  • Did the business outcome improve after returns, discounts and variable costs?

For small stores, a carefully phased rollout and time-based comparison may be more realistic than a formal experiment. For higher-traffic stores, controlled testing becomes more practical. Either way, the evidence hierarchy should be agreed before results arrive.

How money moves through the CRO market

Buyers typically pay in one of four ways: fixed project fee, monthly retainer, software subscription, or internal payroll/engineering allocation. Performance-based fees also exist, but they create a difficult baseline question: what exactly counts as incremental lift, and what happens when seasonality, media mix or pricing changes?

The hidden cost is coordination.

A cheap audit that produces work nobody can ship is expensive. A sophisticated testing platform with no experiment pipeline is expensive. A redesign that improves aesthetics but removes high-performing information is expensive. A developer who implements ten ideas without measurement may create ten new unknowns.

The commercial unit to compare is not hourly rate. It is cost per decision you can trust and operationalize.

A buyer’s routing rule

Use this simple routing rule before requesting proposals:

  • If you do not know where the problem is, start with analytics.
  • If you know where but not why, fund research.
  • If you have evidence-backed hypotheses but cannot ship, buy implementation capacity.
  • If you can ship but cannot separate signal from noise, improve experimentation and measurement.
  • If the process works but moves too slowly, then consider tooling and automation.

That sequence prevents a common procurement mistake: buying a downstream tool to solve an upstream uncertainty.

What changes the answer?

Traffic volume matters. A store with a few dozen purchases a month cannot run the same experimentation program as a retailer with thousands of daily sessions. Platform architecture matters: a heavily customized headless stack has different implementation economics from a standard theme. Category matters because the decision cycle for furniture, skincare and low-cost accessories is not the same. Promotions, inventory constraints and repeat purchase can also distort the headline conversion rate.

CRO is therefore not one market with one “best” provider. It is a workflow with specialized sellers around each decision.

The buyer who maps the workflow first is much less likely to pay for optimization theater.

Sources

Related Reading