1. What primary metric would you choose for a call-to-action color experiment?
A product changes only the color of a call-to-action button. Define the eligible user, first exposure, randomization unit, outcome window, exact numerator and denominator, and one primary metric tied to the intended action. Compare click, downstream completion, and revenue alternatives; add latency, error, abandonment, and accidental-click guardrails; and address returning-user contamination, a 48-hour attribution window, novelty, sample-ratio checks, minimum detectable effect, duration, and the launch decision.
I would choose 48-hour CTA Completion Rate if the intended outcome is the downstream action. I would divide eligible users who complete that action within 48 hours of first exposure by all eligible users with first exposure. I would also monitor clicks, revenue, guardrails, SRM, and practical significance before launching.
The color change should be judged by whether it helps users complete the intended action, not only whether it attracts more clicks. I would therefore use 48-hour CTA Completion Rate as the primary metric. An eligible user is a logged-in, in-scope, non-bot user who reaches the relevant page and is not excluded by a holdout. Randomize by user ID with sticky assignment. First exposure is the first eligible page view after assignment. Then count downstream completions during the next 48 hours and compare users according to their assigned variants.
- What is the exact intended downstream action that the CTA should cause?
- Is the experiment population limited to logged-in users so assignment can remain stable across sessions and devices?
- Is a 48-hour attribution window appropriate for the normal time users need to complete the action?
- What baseline completion rate and minimum detectable effect should we use for power planning?
- Are there business constraints that should make revenue or another downstream measure part of the launch decision?
Start with the user journey. The only treatment is CTA color. For each eligible user, assign one variant using user ID and keep that assignment sticky. The flow is: eligible user → randomization unit → assigned color variant → first eligible page view after assignment → 48-hour outcome window → primary metric.
Define the primary metric exactly. The numerator is the number of eligible users who complete the intended downstream action within 48 hours of first exposure. The denominator is the number of eligible users with first exposure. Analyze each exposed user according to the variant they were assigned. Do not condition the denominator on clicking, because clicking can itself be changed by the treatment.
I would prefer downstream completion over click-through rate when completion is the real product goal. Click-through rate has more events and often gives a faster, lower-variance signal, but it is only a proxy. A color can increase curiosity or accidental clicks without improving the real outcome. Revenue per user is closer to business value when monetization is the goal, but it can be much noisier and can be affected by price mix, promotions, refunds, and longer latency. I would normally use click-through rate as a diagnostic or early signal and revenue per user as a supporting business metric rather than automatically making either one primary.
Add guardrails so an apparent gain is not caused by a worse experience. Monitor page and CTA interaction latency, including median and p95, for regressions between variants. Track button errors or load failures. Track abandonment, such as leaving before interaction. Monitor rapid or immediately reversed clicks as an accidental-click diagnostic. These metrics do not replace the primary outcome, but they can block a launch if the treatment causes harm.
Protect experiment validity. Returning users must keep the same assignment; do not re-randomize them on later visits. When user identity is available across devices, use it to persist assignment. Compare observed assignment counts with the planned 50/50 split. A statistically significant deviation is a sample-ratio mismatch, or SRM, and should be investigated before trusting the treatment-effect estimate. Also inspect treatment effect over time because a new color can create a novelty response that later stabilizes or changes.
Use the fixed 48-hour attribution window consistently for both variants. An outcome belongs to the experiment when it occurs within 48 hours after that user's first exposure. The final analysis should wait until the included users' 48-hour outcome windows have matured so the groups are compared with complete outcome opportunity.
Plan sample size before running the experiment. Choose a baseline completion rate, significance level, desired power, and minimum detectable effect, or MDE. The MDE is the smallest improvement worth designing the experiment to detect. Using the diagram's illustrative assumptions, a 5.0% baseline, a 10% relative MDE corresponding to 5.5%, two-sided alpha of 0.05, and 80% power gives about 31,200 users per variant, or about 62,400 total under the shown two-proportion approximation. These are planning numbers, not observed experiment results.
Run until the planned sample size is reached and cover at least one full business cycle, such as a weekly cycle when weekly behavior matters. Do not repeatedly stop an ordinary fixed-horizon test as soon as a nominal p-value becomes significant. For the binary completion outcome, a two-proportion comparison is appropriate at the planned analysis point. Report the estimated lift, a 95% confidence interval, and the p-value, then interpret the effect size rather than relying on the p-value alone.
The launch rule combines statistical evidence and product judgment. Launch only when the primary-metric lift is statistically significant, large enough to matter, experiment-validity checks pass, and guardrails are healthy. If the planned sample size or minimum duration has not been reached, keep collecting data. Do not launch when there is no statistically significant lift, the estimated benefit is too small to matter, guardrails degrade, the experiment is invalid, or the overall business impact is negative.
- Define the intended downstream action before choosing the metric.
- Define eligible users and exclusions before exposure.
- Randomize by user ID and keep assignment sticky across sessions and devices when possible.
- Assign each user to the control or treatment color variant according to the planned allocation.
- Define first exposure as the first eligible page view after assignment.
- Start a fixed 48-hour outcome window from first exposure.
- Use 48-hour CTA Completion Rate as the primary metric when downstream completion is the intended outcome: completions within 48 hours divided by eligible users with first exposure.
- Track click-through rate and revenue per user as supporting alternatives, not automatic replacements for the primary metric.
- Monitor latency, errors, abandonment, and accidental-click diagnostics as guardrails.
- Check returning-user consistency, cross-device assignment where possible, novelty over time, and sample-ratio mismatch against the planned allocation.
- Choose baseline rate, MDE, alpha, power, and required sample size before the test.
- Run to the planned sample and minimum duration, allowing the 48-hour outcome windows to mature.
- Estimate lift with uncertainty and check practical significance, validity checks, and guardrails before deciding whether to launch, run longer, or not launch.
The metric itself is simple to compute, but trustworthy measurement has operational costs. A downstream completion metric usually needs more users than a click metric because completions are less frequent. A 48-hour window delays the final answer because late outcomes must mature. Sticky user assignment needs reliable identity, especially across devices. Guardrails and SRM checks add monitoring work but protect against bad launches. Revenue can be valuable but usually has higher variance and more business noise. A smaller MDE needs a larger sample and therefore more traffic or more time. Longer tests cost time, while shorter tests risk missing weekly patterns or overreacting to novelty.
This question tests whether the candidate can turn a simple UI experiment into a trustworthy causal measurement plan. The interviewer wants to see whether the candidate can define the population, randomization unit, exposure, attribution window, numerator, denominator, and primary outcome precisely. It also tests whether the candidate can distinguish an easy-to-move proxy such as clicks from a downstream user outcome, use guardrails, detect experiment-quality problems such as sample-ratio mismatch, plan power and duration, and make a launch decision using both statistical and practical significance.
Common mistakes are choosing click-through rate just because it moves faster; failing to define the exact numerator, denominator, or 48-hour window; randomizing by page view instead of keeping a user in one variant; re-randomizing returning users; conditioning the denominator on users who clicked; ignoring cross-device contamination; using revenue as primary without considering its variance and latency; treating guardrails as optional; ignoring a statistically significant sample-ratio mismatch; looking at significance repeatedly and stopping early; choosing the MDE after seeing results; ending the experiment before outcome windows mature; and launching on statistical significance when the effect is too small to matter or guardrails have degraded.
Lead with one sentence naming the primary metric and why it matches the intended action. Then define eligibility, randomization, first exposure, numerator, denominator, and the 48-hour window precisely. Finish with metric alternatives, guardrails, SRM, MDE and duration, and a clear launch rule based on statistical significance, practical significance, validity, and user impact.









