Evaluating the expedition
You might notice that every metric you chose as your winner was solving something useful. But the effect was more ephemeral than anything else. Why so? Let's take a step back-
- Who are the stakeholders here? - Trekkers & businesses (We exclude monastery since it is an independent cultural landmark and not a commercial entity)
- Primary motivation? - Trekkers are seeking a rejuvenating experience- overcoming jitters, pushing through stretches to achieve summit euphoria, and anchoring sacred memories at the monastery, cafe and souvenir shop. Local businesses want to facilitate this experience to the trekkers and earn a sustainable revenue in return.
Whatever value we bring must cater to both these users for the ecosystem to thrive.
- Prima facie, since this trail is remote, improving acquisition is intricate and may not be enough without retention to keep the small businesses alive.
- Secondly, Summiters Expense Rate (%Summiters who purchase at cafe or souvenir shop) might feel more appropriate; It indeed measures immediate value capture, but it fundamentally relies on a steady foot traffic.
So, for Pahadi, driving trekker retention is the solution to a sustainable visitor and business growth at the same time. (PS: The real industry average retention is different, but we are focusing on locals and neighbors!)
The next challenge is - How do you evaluate retention (a delayed metric) through a short experiment? You cannot run the experiment forever to observe actual retention, due to high costs. And temporary signals in isolation may not align with your long-term metric. For ex: We may assume that higher time to climb means higher summit success. But trekkers may actually walk slower, only to wear out sooner and abandon the rest of the trek altogether.
The solution is surrogate metrics: measurable short-term behaviors that predict long-term success.
How it works, in short
Imagine you are working at YouTube, evaluating a new feature designed to increase user retention. User here is the one who streams and consumes content. For this case, let's assume it's a two-week experiment.
You first measure the short-term signals over the course of two weeks: this can be watch hours, active days, subscriptions etc.
But you still need to know: does this two-week lift actually cause long-term retention?
So you look at historical experiments, far enough back to have the actual retention data. (If no history, read Key takeaway 4 for alternatives). An experiment evaluating 12-month retention must have ended at least one year ago to qualify. You don't just consider user habits. You analyze how short two-week gains from past experiments actually drove long-term retention over time. Once you learn this causal translation layer, you feed your new two-week test results into it, and it outputs your expected 12-month retention impact today.
The Advantage: You get an estimate now instead of three, six or twelve months into the future
How are we measuring true treatment induced lift?
We are not simply comparing treatment and control means of the retention estimate here. There is a chance that these differences could include skewness due to power users and variance from lucky treatment assignments. Instead, causal forests solve this by using pre-treatment user attributes, ex: account age, country etc. to dynamically segment your audience into hyper-similar peer groups. These segments include users from both- treatment and control. By calculating localized treatment effects (CATE) within these segments, the model isolates which specific segments experience the highest uplift. One thing to be wary is that CFs act as variance reduction and targeting framework. When identifying segments where the lift is highest, they provide prioritization for phased feature roll outs which should not be misinterpreted as the only winner segments that deserve feature access.
What are my assumptions?
- The biggest one: your long-term metric
(Y) is influenced only through the observed
short-term metrics (S). The treatment assignment (treated-yes/no)
has no effect.
How to validate: regress Y against S and the treatment indicator (W — 1: yes, 0: no). If the coefficient of treatment assignment is large and significant, you have causal effects that haven't been captured by your short-term surrogates.
- All experiment validation tests hold true — no sample ratio mismatch and no SUTVA violations.
How do I know I have the right short-term metrics?
The right way to design metrics is to align with the user outcomes you expect from the treatment, not just clicks. Build aspects that are more orthogonal to each other. For ex: If this is a video recommendation change, you might consider:
- Intensity: # videos watched above 30s, or average video completions.
- Breadth: distinct categories watched.
- Completion: task success, or view-to-watch conversions.
- Friction: skips, "not interested," notification disables.
If you are new to metric design, check out this FRAMEWORK.
Key takeaways
- North Star Alignment
Long-term metrics ensure that new features/experiments positively move the ultimate metrics - Lifetime value, retention or others over local metrics
- Accounting for Novelty Effects
Initial feature releases often experience a temporary "novelty spike" from curious users that decays over time. Long-term metrics filter out this initial noise, measuring settled user behavior after habits have fully formed.
- Crucial but not a panacea
Long-term metrics are not however the automatic choice in every experiment. When addressing urgent or immediate blockers, short-term metrics add better value and take priority.
Example: Flash sales focus purely on short-term conversions and immediate inventory clearance.
- The holdouts alternative
For novel experiments, where historical exposure is unavailable, keep aside a small holdout group (e.g. 1–5% of traffic) that continues to be exposed to treatment. Observe cumulative treatment effects and guard against novelty decay over extended periods.
- Composite & blended scoring
Creating blended metrics that weigh short-term proxy signals with the north star are additionally explored.
Hope you found this interesting! For people not interested in the math, this is where we take different paths. The rest of the article is for data scientists.
Exploring the mathOptional
Let's continue on the YouTube example: we are testing a new feature designed to retain consumers. Suppose we collect the following short-term metrics.
Short-term metrics: one signal per aspect from the list above — 30s+ views (intensity), distinct categories watched (breadth), video completions (completion), and skips (friction). Note that skips run the other way: higher is worse.
Assumption: the treatment affects the long-term outcome ONLY through the observed short-term surrogate metrics.
A good thumb-rule: always include multiple short-term behaviors to capture both engagement and friction signals.
The surrogate index calculation
- Take completed historical experiments which have actual data available on Y. Given Y here is 12m retention, all experiments conducted at least 12 months before are eligible, since we can evaluate the true Y on customers.
- Measure actual observed short-term metrics (S) at the end of experiments and the actual observed Y that is available post 12 months from the experiment.
Regress Y on both X and surrogate vector S.
Y = β0 + β1S + β2X + ε
S: vector of 14-day surrogate
behaviors.
X: pre-experiment user covariates (e.g. country, account
age, historical activity).
The causal forest
Now let's come to our new experiment. We have S at the end of the experiment but no Y.
- Impute each user's estimated Y from the surrogate index of the previous step. This is done at the end of the experiment, predicting Y using available data on short-term metrics.
- Train a Generalized Random Forest with that imputed Y, treatment W and covariates X.
In effect, the forest is spliting users on their covariates until it finds groups whose treated - control gap in predicted retention genuinely differs.
| Optimization objective | Estimate the Individual Treatment Effect (ITE) function: τ(S, X) = E[ Y(1) − Y(0) | S, X ] |
|---|---|
| Operational output | Personalized Causal Scoring Engine: outputs predicted 365-day treatment effect τ̂i for every user in a new 14-day experiment. |
What we start with
| User | Account age | W | 30s views | cats | compl | skips | Ŷ |
|---|---|---|---|---|---|---|---|
| User 1 | New | 1 | 22 | 3 | 8 | 2 | 0.421 |
| User 2 | New | 1 | 19 | 2 | 6 | 3 | 0.405 |
| User 3 | New | 0 | 17 | 2 | 5 | 5 | 0.392 |
| User 4 | Established | 1 | 24 | 3 | 9 | 1 | 0.716 |
| User 5 | Established | 0 | 23 | 3 | 8 | 2 | 0.712 |
Model output
Average the effect across everyone
Each user gets a score Γ = their leaf's effect + a correction for how far their own Ŷ sat from the leaf-arm average. Treatment was 50/50, so the correction divides by 0.5.
Similarly, Γ = +0.005 for User 2, +0.021 for User 3, +0.004 for User 4, and +0.004 for User 5.
The new algorithm lifts 12-month retention — but not evenly.
Validation across historical experiments
To establish empirical trust in the model, backtest it on a portfolio of completed historical A/B tests where true 365-day outcomes (Yactual) have already been observed.
- Extract the first 14 days of behavioral data for each of the past experiments.
- Run S14d through the trained Causal Forest to generate predicted long-term treatment effect (ΔŶ).
- Compute the actual observed 365-day treatment effect (ΔYactual) for each experiment.
- Plot predicted lift (ΔŶ) against actual 365-day lift (ΔYactual) across all tests.
Validation metrics & benchmark targets
| Validation metric | Methodology / formula | Production target |
|---|---|---|
| Treatment effect correlation (R2treatment) | Measure correlation between predicted lift and actual 365-day lift across experiment points. | R2 > 0.80 |
| Directional alignment rate | Percentage of experiments where Sign(predicted lift) == Sign(actual 365-day lift). | > 90% alignment |
| A/A test calibration | Evaluate model on historical A/A tests or non-impactful UI tweaks. | Predicted lift ≈ 0 (statistically indistinguishable) |
Decision framework
At Day 14 of your live experiment, aggregate individual Causal Forest predictions across all treatment and control users to compute the overall Average Treatment Effect (ATE):
τ̂ATE = (1 / N) ⋅ Σ τ̂i(Si, Xi)
Rollout decision rules
| Decision outcome | Condition criteria & thresholds |
|---|---|
| Ship feature |
|
| Abandon feature |
|
| Inconclusive |
|
Key resources
- The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely Athey, Chetty, Imbens & Kang, 2019
- Focusing on the Long-term: It's Good for Users and Business Hohnhold, O'Brien & Tang, Google, KDD 2015
- How Uber Predicts Long-Term User Value in Experiments Uber Engineering Blog