When a programmatic test is inconclusive: Check validity first: both arms ran, budgets in proportion, creative and exclusions aligned.; An interval spanning useful gain and material loss means an inconclusive choice.; Set the decision threshold and analysis rule before reading results.
Image: Programmatic Ad Guide

Measurement

Part of Programmatic experimentation

Deciding when a programmatic test is inconclusive

Identify an inconclusive programmatic test using setup validity, outcome counts, uncertainty and the decision it was meant to support.

Call a programmatic test inconclusive when its evidence cannot support the planned choice between arms. A higher observed rate is insufficient when outcomes are sparse, uncertainty still includes decision-changing results, or the arms did not run as intended. State which limit prevents the choice.

Check validity first

Compare the saved and observed conditions with the original plan. Did both arms run during the intended period with budgets in proportion? Were creative, destination, event definition and required exclusions aligned? Record pauses, approval delays, competing budget demands and unequal edits with their dates.

Keep the planned change isolated and apply shared edits consistently across arms. Compare reporting periods and outcome definitions before treating a discrepancy as a performance signal.

If a material setup fault cannot be bounded to a clean period, report the fault rather than declaring a winner from combined totals.

Validity checks before reading a test result

  1. Compare saved and observed conditions with the original plan
  2. Confirm both arms ran for the intended period with budgets in proportion
  3. Check creative, destination, event definition and required exclusions were aligned
  4. Record pauses, approval delays, competing budget demands and unequal edits with their dates
  5. Keep the planned change isolated and apply shared edits consistently across arms
  6. Bound a material setup fault to a clean period, or report the fault instead of declaring a winner

Ask whether the evidence resolves the decision

Read actual outcome counts, spend and uncertainty around the difference. Assess uncertainty using the analysis rule set for the test.

An interval may include both a gain large enough to adopt the variant and a loss large enough to reject it. That is inconclusive for the planned choice, even when one point estimate is higher.

A result that does not show a statistically distinguishable difference is not proof that the tactics are equivalent: there may be too little information. A detectable difference may also be too small to justify an operational switch. Set the decision threshold and analysis rule before reading results.

FindingDecision supported
One arm barely deliveredDiagnose eligibility; its outcome rate is a weak basis for wider allocation.
Both delivered but completed outcomes are fewKeep the effectiveness question open and set a follow-up window.
The interval spans useful gain and material lossRecord an inconclusive choice and the uncertainty.
One arm had a material mid-test changeAnalyse a clean period only if defensible; otherwise redesign the test.
A clear gain breaches a required spend or suitability guardrailReject the switch under the original rule.

Findings and the decisions they support

  • One arm barely deliveredDiagnose eligibility; its outcome rate is a weak basis for wider allocation.
  • Both delivered but completed outcomes are fewKeep the effectiveness question open and set a follow-up window.
  • The interval spans useful gain and material lossRecord an inconclusive choice and the uncertainty.
  • One arm had a material mid-test changeAnalyse a clean period only if defensible; otherwise redesign the test.
  • A clear gain breaches a required spend or suitability guardrailReject the switch under the original rule.

Decide whether to test again

Allow for the outcome to appear before interpreting results. Compare periods with similar outcome maturity.

A follow-up is useful when a longer eligible run, repaired setup or better-defined outcome could resolve a valuable decision within the remaining budget. Define its question and stopping rule before launch.

If available inventory is unlikely to produce enough evidence, state that limit and make a bounded operating choice from other evidence. Do not relabel an underpowered result as proof of no effect.

Close the test record with the planned question, actual conditions, counts, uncertainty, unresolved cause and next action.

Deciding whether to test again

  1. Allow time for the outcome to appear before interpreting results
  2. Compare periods with similar outcome maturity
  3. Run a follow-up only if a longer eligible run, repaired setup or better-defined outcome could resolve a valuable decision within the remaining budget
  4. Define the follow-up question and stopping rule before launch
  5. If inventory is unlikely to produce enough evidence, state that limit and make a bounded operating choice from other evidence
  6. Close the test record with the planned question, actual conditions, counts, uncertainty, unresolved cause and next action

More from Measurement

Measurement

Comparing inventory segments at a useful sample size

Judge whether inventory-segment delivery and outcomes can support a budget decision while keeping small samples and buying differences visible.