Imagine a two-sided services marketplace where customers post requests and professionals compete to fulfill them. The ranking model that decides how professionals appear in those search results is about to be replaced. You must design a controlled test that measures the new ranker while keeping marketplace disruption low and avoiding supply cannibalization between treatment arms.
Address the following six areas:
Primary metric and guardrails. Choose one main success metric, such as completed bookings divided by submitted requests. Then choose at least four guardrail metrics, including elapsed time until the first quote, cancellation rate, typical professional response delay, and fairness or earnings spread among professionals. For each metric, give the exact calculation and state what amount of movement would be considered acceptable.
Assignment unit and experiment structure. Choose between randomizing by request, by customer, by geography-level cluster, or through a region-by-hour switchback. Justify your choice as a way to reduce spillovers when the same professional can serve requests from both arms. Explain how you will prevent professionals from preferentially serving one arm.
Power and duration. Suppose baseline booking conversion is , the target relative lift is , the two-sided significance level is , power is , and there are about eligible requests per day. Estimate the required sample size per arm and the runtime in days under 1:1 allocation. Show your formulas and assumptions, including pooled variance for a two-proportion z-test. Explain how clustering or switchback increases variance via a design effect, and revise the runtime using a plausible intraclass correlation.
Bias controls. Specify pre-experiment checks for covariate balance. Describe variance-reduction methods, such as CUPED based on pre-period request conversion or stratification by category and region. Explain how you will handle customers who appear repeatedly and how you will address daylight saving shifts and time-of-day effects.
Monitoring and stopping. Propose a sequential monitoring plan, for example O’Brien–Fleming boundaries or an alpha-spending schedule, together with anomaly triggers. Define what happens if a guardrail is breached while the primary metric improves.
Readout. Detail the difference-in-means estimator, heterogeneity by category, region, and traffic source, and how you would attribute uplift versus cannibalization across supply-limited segments.
Given parameters: