A/B test · form step
Hypothesis first
Control
12 form fields
Abandonment spikes at field 7
Variant
6 form fields
KPI: form completion
Ship it
Kill it
Run it longer
01 — Four ways CRO goes wrong
CRO programs don’t fail in one obvious way. They fail in four quieter ways, and most teams running them have lived through at least two.
We’ve worked through all four. The fix is rarely “more tests.” It’s better hypotheses, the right KPIs, and the discipline to kill what should be killed.
02 — What we actually do
Most CRO engagements blend testing with everything around it — funnel analysis, friction work, KPI alignment. The work below is what actually fills our weeks.
01
Funnel diagnostics
Before we test, we map. Where does traffic enter, where does it leak, where does it stall, where does it convert. We use Contentsquare for friction analysis, GA4 for funnel reporting, and the unglamorous data-pulling work to figure out what’s actually broken before deciding what to test. Most teams skip this and test surface-level things while the real friction sits two steps upstream.
02
Hypothesis design
A test deserves a hypothesis. “Reducing form fields from 12 to 6 will increase form completion by X% because the friction analysis shows abandonment spikes at field 7.” That’s a hypothesis. “Let’s test a green button” isn’t. We’re stricter about this than most agencies. It saves traffic, time, and post-test arguments.
03
Test execution
AB Tasty is our primary platform — we’ve built and run hundreds of variants on enterprise sites. Optimizely when the client’s already on it. Component-level test development included — we don’t just configure, we build the variants that run. Senior hands behind every test, not just configuration. What you’re billed for is what you get.
04
Honest analysis
Tests end one of three ways: ship it, kill it, or run it longer. We’re explicit about which result we got and why. Including telling you when a test that “won” probably shouldn’t be shipped because the result was statistically significant but practically meaningless. The honest analysis is the work most agencies skip.
05
Tactical optimization sprints
The Black Friday landing page. The campaign hub for the new product launch. The optimization work that doesn’t need a structured testing program but still needs to perform. We do this work too — same discipline, shorter cycle. Sometimes the right call isn’t “set up a test.” Sometimes it’s “ship the best possible version and learn from what happens.”
03 — Tools and platforms
AB Tasty as primary — the platform we run day-one, with deep component development experience on enterprise sites. Optimizely when the client’s already on it; we work in it competently. Contentsquare for friction analysis — heatmaps, session replay, funnel analytics that go deeper than what GA4 can show.
GA4 and Looker Studio underneath everything for funnel reporting and post-test analysis (more on that on the Analytics & GTM page).
We don’t pitch Microsoft Clarity, Hotjar, or VWO as primary tools. Solid platforms, just not where we’ve put our depth. If you’re already on one, we can work in it. If you’re choosing, we’ll usually steer toward AB Tasty or Contentsquare based on your actual needs.
04 — How we work
Engagement length depends on whether we’re auditing a funnel or running a continuous testing program. Funnel audits run two to three weeks. Quarterly testing programs run as ongoing retainers. Tactical sprints run two to four weeks. Most clients keep us on retainer because testing programs need consistency to compound.
01
Audit
Two to three weeks. We map the funnel, pull the data, identify where conversion actually breaks down. You get a written diagnosis of where the friction is, which areas have the highest testing potential, and which probably aren’t worth testing because the underlying problem is upstream.
02
Hypothesis backlog
We build a prioritized backlog of test hypotheses, scored by expected impact, traffic available, and effort to build. Not 50 ideas. 10 to 15 well-formed ones, ranked. The team picks the order based on business priorities.
03
Test execution
Variants designed, built, QA’d, launched. We monitor for technical issues, traffic anomalies, and early signal that a test is broken (not winning — broken). Results pulled and analyzed when the test reaches significance or runs out of runway.
04
Decision
Ship, kill, or rerun. Documented with reasoning. The backlog updates based on what we learned. New hypotheses get added based on what didn’t work and why. The program compounds — each test makes the next one better.
05 — Engagements
Premium B2C, hospitality, multi-brand. The common thread is testing programs that have to hold up under enterprise scrutiny — and not just deliver a green-button win.
06 — Related services