Glossary · CRO

What is a/B testing?

A/B testing is a controlled experiment comparing variants with a defined audience, allocation rule, and outcome. Random assignment and a planned analysis help assess a proposed change while preserving uncertainty and implementation limits.

Updated

What is A/B testing, and what makes it an experiment?

A/B testing is a controlled comparison of alternatives, usually with eligible people randomly assigned to different versions. The team specifies what changes, who participates, and what outcome matters. It then evaluates the evidence using a planned method, rather than declaring whichever version currently looks better the winner.

  • A website test might compare explanations above an inquiry form.

    Both groups encounter the same service, but the explanation differs. Random allocation provides the basis for comparing responses without deliberately sending the strongest prospects to a preferred version.

  • The NIST description of completely randomized designs explains random assignment in general experimental design.

    Its examples concern engineering processes. Applying that principle to a website still requires decisions about visitor assignment, repeated visits, measurement, and statistical assumptions.

  • A/B testing is one method within conversion rate optimization.

    It can help resolve uncertainty about a proposed change. It does not replace fixing a broken form, understanding customer needs, or checking whether the offered service is actually available.

See the relationships

A visitor experiment starts before the dashboard

Define the allocation and decision before interpreting differences between variants.

A/B testing: related considerationsConceptual connections between A/B testing and Eligible audience, Random allocation, Outcome and guardrails, Planned analysis. Connections group considerations; they do not represent measured effects or mandatory sequence. Explanations follow below.A/B testingEligible audienceRandom allocationOutcome andguardrailsPlanned analysis
  • Eligible audience

    Specify who can encounter the tested task.

  • Random allocation

    Assign eligible participants to control and variant under the planned method.

  • Outcome and guardrails

    Measure the intended action alongside important adverse effects.

  • Planned analysis

    Interpret uncertainty and practical relevance before rollout.

The outcome, guardrails and analysis plan belong to the experiment design.Conceptual illustration informed by 5.3.3.1. Completely randomized designs.

Concurrent variants, before-and-after monitoring and SEO split tests

An A/B test compares assigned groups under a defined experiment. A before-and-after comparison observes different periods, which may also differ in demand, advertising, staffing, or traffic composition. The latter can describe a change, but it needs additional reasoning before the page revision receives credit for the outcome.

  • Suppose a contractor replaces its estimate page when seasonal demand rises.

    More inquiries afterward could reflect the new page, the weather, or both. A dashboard showing the increase does not isolate those explanations.

  • Concurrent assignment avoids deliberately placing all of one version in the earlier period.

    It does not eliminate every implementation or analysis problem. If the allocation rule treats returning visitors inconsistently, the groups may not represent the comparison the team intended.

Concurrent variants, before-and-after monitoring and SEO split tests
Approach Comparison unit What needs care
Visitor A/B test Assigned visitors or sessions Stable allocation and the chosen customer outcome
Before-and-after monitoring Different periods Demand, campaigns and other changes between periods
SEO split test Groups of pages Page comparability and consistent crawler-visible changes

Start with a decision the business actually needs to make

For example, support calls might suggest that customers mistake a callback request for a confirmed appointment. A test could compare two truthful explanations of the callback process. The question concerns expectations and successful requests, rather than whether a more exciting headline attracts attention.

  • Write the mechanism before designing variants: clearer process information may help suitable customers decide whether to request contact.

    Then identify evidence that could challenge this explanation, such as unchanged request completion or increased cancellations after callbacks.

  • The GOV.UK Test and Learn guide discusses testing assumptions and choosing methods for uncertain policy and service decisions.

    That public-sector framework is useful context, not evidence that a particular commercial page change will work.

  • A revised call to action might clarify that the next step is a callback.

    Keeping the surrounding offer and form unchanged supports that specific comparison. Adding a discount, removing fields, and changing the promise at the same time creates a different experiment.

  • A package comparison can still be useful when the business must choose between complete designs.

    Name it accurately. The result may support adopting one package, but it does not establish a universal rule about headline length, button color, or form fields.

Define completion before choosing a conversion metric

The primary outcome should represent the customer action the test is intended to improve, using the same definition for every assigned group. Choose it before launch. Intermediate actions can explain behavior, but a rise in clicks alone may not support a decision about completed inquiries or suitable booked work.

  • For a form explanation test, the primary outcome might be a successfully accepted request per eligible assigned visitor.

    Define whether acceptance means server confirmation, an analytics event, or entry into an inquiry system. These are different observations unless the implementation reliably connects them.

  • Keep the numerator and denominator compatible.

    A conversion rate based on sessions answers a different question from a rate based on assigned visitors. Repeated visits make that distinction especially relevant. Specify the analysis unit instead of selecting whichever denominator presents the stronger result.

  • If qualification happens later, document how requests are matched to their experiment assignment without exposing personal information in public analytics.

    Also define the follow-up window. A request still awaiting review is not automatically unsuitable, and newer groups may have had less time to receive that review.

  • An event named generate_lead does not establish that the inquiry was accepted, within the service area, or useful to the office.

    Inspect its trigger. If it fires on the submit button, it may include requests that fail validation or never reach the destination.

  • The distinction between events and key events helps organize reporting.

    Google’s key-event documentation explains their platform role. It does not certify the business quality of the action you choose to mark.

  • Keep raw implementation evidence alongside business definitions.

    If analytics and the inquiry system disagree, investigate the mismatch. Do not select the system that makes the preferred variant appear stronger without resolving why the counts differ.

Assign an eligible audience consistently

Assignment should follow a documented allocation rule that matches the experiment’s analysis unit. Randomization belongs in the implementation, not in an informal choice by staff. If the question concerns an experience across repeated visits, decide how the same eligible person will retain an assignment and how unrecognized visits will be handled.

  • Test allocation at the boundaries.

    Check a first visit, a return visit, a private browsing session, and any relevant cross-domain step. These checks establish what the software does. They do not prove that separate devices can always be linked to the same person.

  • Use only identifiers and storage permitted by the site’s privacy arrangements.

    Do not invent identity matching to make a report look cleaner. Document where recognition is limited, including consent choices or devices that prevent persistent assignment.

  • A booking explanation for a particular service may belong only on that service’s entry page.

    Including unrelated traffic would answer a broader and possibly unhelpful question. On the other hand, limiting the test to desktop visitors would not establish how mobile customers respond.

  • Separate legitimate operational exclusions from outcome-based editing.

    Staff checking the page may be excluded using a known rule. Removing visitors because they did not submit a request would erase part of the behavior the experiment is supposed to measure.

Protect the customer task while the test runs

A variant could generate more requests while overwhelming a small office with out-of-area inquiries. That result needs a business decision, not an automatic rollout. Record how service eligibility is assessed and apply the same criteria to both groups.

  • Check whether the change affects keyboard access, error recovery, or the clarity of confirmation messages.

    A known accessibility defect should be corrected. Customers should not have to experience a broken control so the team can measure whether a working version performs better.

  • A client-side replacement might briefly show the original wording before inserting a variant.

    A server-side implementation might avoid that replacement but require different caching and routing checks. Neither approach is automatically correct for every site.

Check delivery and measurement with both variants

Use a test request that the office can identify and safely remove from operational reporting. Confirm the request reaches the intended destination. Inspect whether its completion signal occurs after acceptance rather than before the server has responded.

Check delivery and measurement with both variants
Point to considerExplanation and application
GA4 can record events used in an experiment’s measurement plan.The Google Analytics DebugView documentation explains inspecting debug events. DebugView is an implementation check, not a substitute for validating allocation or analyzing the experiment’s uncertainty.
Also test validation errors, network failures, and repeated clicks.A page may behave correctly on its successful route while producing false completions elsewhere. Preserve evidence of these checks before genuine customers are assigned.
Start with counts that describe experiment delivery: eligible visits, assignments, actual exposures, and recorded outcomes.These may differ for legitimate reasons, but unexplained differences need attention. The planned allocation does not establish that the implementation delivered it correctly.
If a serious defect affects only one group, do not present the remaining data as an uncomplicated design comparison.Explain the incident, determine whether analysis remains defensible, and consider restarting a repaired experiment. A longer report cannot compensate for an unidentified allocation failure.
Google documents GA4 integrations for A/B tests: the experimentation tool runs the test, while Analytics helps interpret collected results.Neither a key-event label nor a DebugView screenshot proves that participants were allocated correctly.

Plan sample size, stopping and uncertainty together

Set the analysis plan before examining which version performs better. Specify the primary outcome, analysis unit, evaluation method, and rules for ending or pausing the experiment. Record how uncertainty and implementation failures will be handled, so the final decision is not reconstructed around whichever result happens to look most attractive.

  • A statistical method should fit the data and design.

    Repeated observations from the same person, sparse completed requests, or delayed qualification can complicate a simple comparison. Ask a qualified analyst to check assumptions when the decision requires statistical inference.

  • NIST’s introduction to statistical tests explains hypotheses and error risks in general statistical testing.

    It is a foundation for understanding analysis, not a ready-made configuration for every website experiment.

  • Monitoring for a broken form is different from repeatedly choosing a winner.

    Define operational pause rules separately from the inferential stopping method. If the platform supports a sequential method, follow that method’s requirements rather than treating an ordinary dashboard as permission to stop at any encouraging moment.

  • A small website may not collect enough relevant observations to answer the proposed question within a useful decision window.

    Feasibility depends on the outcome’s frequency, the effect worth detecting, and the chosen method. There is no universal traffic count or fixed duration that makes every service-business experiment informative for its intended decision.

  • Estimate feasibility before building variants.

    Use recent eligible traffic and a consistently measured outcome as planning inputs. If the measurement is unstable, repair it first. A precise-looking sample calculation based on unreliable inputs does not resolve the underlying problem.

  • A sample plan connects the baseline outcome rate, the smallest effect worth detecting, desired error control and the chosen method.

    A rare accepted request needs a different feasibility assessment from a frequent button click. Replacing the primary outcome with a more frequent action may make a test easier to run while changing the question it answers. Make that tradeoff before launch.

  • Statistical significance concerns evidence under a particular model and testing procedure.

    Practical significance concerns whether the change matters enough to justify a business decision. A result can be statistically detectable yet too small to outweigh implementation cost, or potentially useful but too uncertain to support confident adoption from the collected evidence.

  • Compare the estimated effect with the business’s decision threshold, expressed in a meaningful outcome.

    Also examine uncertainty around that estimate. A percentage headline without context may hide a small baseline, an imprecise estimate, or a costly implementation.

  • NIST’s quantitative techniques guidance distinguishes practical from statistical significance and discusses interval estimates.

    Its engineering examples illustrate statistical concepts. They do not provide a universal threshold for deciding whether a contractor should redesign a website.

  • A p-value should not be described as the probability that the preferred design is best.

    Nor does a nonsignificant result prove identical performance. Record the method and its interpretation instead of using a dashboard badge as the entire explanation.

Use research when an experiment cannot answer in time

Consider the cost of waiting. A business facing an obviously confusing explanation may learn more quickly from customer interviews and usability observation. GOV.UK’s guidance on research in live services describes combining analytics, operational records, and user research.

Use research when an experiment cannot answer in time
Point to considerExplanation and application
Those methods answer different questions.Observing confusion can justify correcting it without claiming a measured conversion lift. If the experiment remains worthwhile, narrow it to a meaningful decision and document when insufficient evidence will lead to another research method.
An inconclusive result means the collected evidence does not resolve the planned decision under the chosen method.It does not mean the designs are equivalent. An unfavorable result can challenge the proposed explanation, but first check whether the implementation and measurement actually tested the intended experience without a delivery defect.
Review the estimated outcome and its uncertainty, along with safeguards.Decide whether to retain the existing version, gather further evidence under an appropriate plan, or revise the hypothesis. Do not extend or redefine a test solely to make the result favorable.
Keep unsuccessful tests in the record.They can prevent colleagues from repeating an unsupported idea and show where the business still lacks evidence. A test archive containing only declared winners misrepresents the learning process.

Google’s website-testing guidance addresses how experiments interact with search crawling and indexing. It advises against cloaking and describes canonical links, temporary redirects, and timely cleanup where applicable. These are implementation considerations for search visibility; they do not establish sample requirements or confirm that an experimental result is statistically reliable for the business.

  • Read the Google Search guidance on website testing before choosing a routing method.

    If alternate URLs are used, its canonical advice concerns those test URLs. If a test redirects visitors, its redirect advice calls for temporary rather than permanent redirects.

  • Do not create a special crawler experience to conceal the experiment.

    Also keep ordinary SEO duties in view: the customer page must remain accessible, meaningful, and technically sound. Search-friendly setup is one part of implementation quality.

A callback-explanation experiment, without an invented result

Illustrative example, not client data. A landscaping business finds that callers expect an immediate appointment after submitting a request. Its existing landing page says that the office will respond. A proposed version explains that the office will call to discuss the project and confirm available next steps.

  • Both versions must describe what the office actually does.

    The test should not promise faster contact unless staffing supports it. The team defines its outcome and qualification review before launch, then checks both variants through a real test submission.

  • A rise in button clicks would be a supporting observation.

    The rollout decision concerns accepted suitable requests, customer expectations, and the planned analysis. No outcome is assumed here; the example shows how the experiment would be designed.

  • The test record could therefore name accepted callback requests per assigned eligible visitor as its primary outcome, with failed delivery and out-of-area requests as separate checks.

    The office must apply the same suitability criteria and follow-up window to both groups. A variant that attracts more clicks but creates misleading appointment expectations has not resolved the original problem.

  • A clearer callback explanation might transfer well to pages sharing that process.

    It would not belong on a page that confirms an appointment immediately. Preserve the customer task instead of copying wording solely because its variant label once appeared above another in a report.

Close the test and leave an inspectable decision

After the decision, implement the chosen experience deliberately and remove obsolete experiment machinery. Preserve the design, evidence, limitations, and reasoning in a test record. Continue checking the customer journey after rollout, because a successful experimental route does not ensure that the permanent implementation behaves identically under later operational conditions or changes.

Our SEO ROI calculator can explore assumptions about business value. It cannot randomize visitors, estimate experimental uncertainty, or establish which variant caused a result. Use it after clarifying the outcome, not as the experiment’s statistical analysis tool.

  1. Reconstruct the decision. Locate the observed problem, proposed explanation, variant record, eligible audience, and primary outcome. If these were never recorded, distinguish later interpretation from the original plan.
  2. Follow both experiences. Verify assignment, return visits, successful requests, validation failures, and relevant external steps. Confirm what constitutes exposure and completion.
  3. Reconcile the evidence. Compare assignments, analytics, and operational records. Explain discrepancies and apply documented exclusions consistently.
  4. Review the analysis. Confirm the method, stopping rules, uncertainty, and safeguard review. Identify exploratory observations separately from the planned conclusion.
  5. Close the implementation. Record the decision and its limits. Remove obsolete routes and scripts, then verify the permanent experience.

A form-field experiment and a reassurance-message experiment may affect the same task. Their combined experience deserves deliberate planning. Record other active tests, and involve the experiment designer before overlapping changes rather than assuming separate dashboard labels make their effects independent.

  • Keep the decision scoped to the tested audience, versions and period.

    Record whether the conclusion was planned or exploratory, which implementation incidents mattered, and why the expected business value justified rollout. An honest result can be to retain the original version or repair the measurement before trying again.

  • The primary sources above cover different parts of the work: NIST explains experimental and statistical concepts; GOV.UK discusses evidence gathering in services; Google documents search implementation and analytics behavior. None supplies a universal conversion benchmark or certifies a website test’s business conclusion.

  • Continue through the CRO glossary for connected definitions.

Questions about A/B testing

How is user A/B testing different from SEO split testing?

Visitor testing randomly assigns users to experiences. SEO split testing compares groups of pages because search engines need a consistent page version to evaluate.

SearchPilot: user tests and SEO split tests ↗
Can GA4 run a test by itself?

GA4 interprets results collected through integrations; a third-party experimentation tool runs and manages the test. Confirm the assignment and outcome collection in that tool.

GA4 A/B tests ↗
Why should a testing redirect be temporary?

Google recommends 302 for experiment redirects so the temporary variation does not communicate a permanent page move.

Google Search guidance for website experiments ↗
What makes a completely randomized design randomized?

NIST describes assigning treatments to experimental units at random. Define the units and treatments before assignment so the experiment follows the planned design.

NIST experimental methods ↗

Continue learning

Connect this to your website

Sources

5.3.3.1. Completely randomized designs ↗Accessed October 8, 20267.1.3. What are statistical tests? ↗Accessed October 8, 20261.3.5. Quantitative Techniques ↗Accessed October 8, 2026Test and Learn (HTML) - GOV.UK ↗Accessed October 8, 2026User research in live - Service Manual - GOV.UK ↗Accessed October 8, 2026A/B Testing Best Practices for Search | Google Search Central  |  Documentation  |  Google for Developers ↗Accessed October 8, 2026Monitor events in DebugView - Analytics Help ↗Accessed October 8, 2026Conversions vs. key events in Google Analytics - Analytics Help ↗Accessed October 8, 2026SearchPilot: user tests and SEO split tests ↗Accessed October 8, 2026GA4 A/B tests ↗Accessed October 8, 2026

Published . Definitions and examples link to their supporting sources. Our SEO methodology →

SEO · Content · Local · Web Design

Connect the website work to your business.

We assess the pages, search demand, and customer actions that matter to your business, then explain where to focus the work.