01The constraint
Most product questions can be settled by experiment. Split the audience, ship two versions, read the result. It is the cleanest form of evidence a product owner has, and when it is available you should use it.
Changes to the rules traders are evaluated against are not available for that treatment, for two reasons that have nothing to do with statistics.
Fairness. An experiment on rule structure means two customers buy the same product on the same day under different terms, and one of them is on the arm you expect to perform worse. In a business where those terms decide whether somebody keeps their account, that is not a test design. It is a decision to treat some customers worse in order to learn something.
The terms are public. Rules are published, compared, screenshotted and discussed. A variant does not stay inside the experiment: it becomes a thing people notice, argue about, and route around. By the time you read the result, the result is measuring the reaction to the variant as much as the variant itself.
So the honest position is that the cleanest evidence is unavailable, and the job is to decide well without it rather than to pretend otherwise.
02What replaces the experiment
The substitute is a backtest: replay the proposed rule against how people actually behaved, and see who would have been affected and how much.
This answers a narrow question well. It tells you the size of the population the change touches, which is usually the first thing people get wrong when they argue about a rule. Debates about rule changes tend to be conducted in anecdotes — a loud case, a recent complaint — and a backtest replaces the anecdote with a distribution. Often the most valuable output is simply that the change affects far fewer, or far more, people than the room assumed.
It also makes the argument concrete enough to disagree with. A proposal that cannot be backtested at all is usually a proposal that has not been specified properly, and finding that out early is worth the exercise on its own.
03What a backtest cannot tell you
Two limits matter, and both are worth stating out loud before anyone presents the numbers.
People change what they do when the rule changes. A backtest replays history produced under the old terms. The traders in that history were optimising against the rules they had, not the rules you are about to introduce. The moment the new rule exists, behaviour moves towards it — so the estimate describes a world that stops existing on the day you ship.
A backtest tells you what would have happened to people who never saw the rule.
The data is thinnest exactly where it matters. The traders most affected by a rule change are, by definition, the ones near its boundary — and they are the rarest group in the history. So the estimate is least reliable precisely for the population the change is about. Aggregate confidence hides this; you have to go and look at the edge deliberately.
Neither limit is a reason to skip the backtest. They are reasons to treat the output as a way to size a decision, not as a prediction of it.
04Deciding anyway
If the evidence cannot settle the question, the discipline moves to what gets agreed before launch, while nobody is yet invested in the outcome.
Two things were written down. The success metric: the single number the change was meant to move, named in advance so that afterwards nobody could quietly nominate whichever number happened to look good. Guardrail metrics: the numbers that must not move, which exist to catch the change that wins on its target while damaging something you were not watching.
The value of writing them down is less about measurement than about timing. After launch, everyone has a position. Before launch, the same people will happily agree what would count as a problem, because nothing is at stake yet.
05What happened
The change did what it was meant to do on its target metric. A guardrail moved too, and we had to react.
That is the entire case for guardrails. Without one, the change is a success: the number it was aimed at went the right way, and the effect elsewhere shows up weeks later as a vague sense that something is off, traced back — if at all — long after the decision has been absorbed into everything else that happened that quarter.
With one, the conversation happens in days, about a specific number, with the original reasoning still in everyone's head.
06What was missing
We agreed what to watch. We did not agree in advance what would make us reverse it.
That sounds like a small omission and it is not. Once a guardrail moves, the discussion is no longer "is this the threshold we set" — it is "is this bad enough", conducted by people who have already invested in the change, under pressure, with the number moving while they argue. The absence of a pre-agreed threshold does not make the decision harder to make. It makes it slower, and it makes it easier to talk yourself out of.
A kill criterion costs nothing to write down in the planning meeting. It is expensive only at the moment you need it and do not have it.
07How I would run it now
- Agree kill criteria before launch.Not just what to watch, but the level at which we stop, decided while nobody has a stake in the answer.
- Write the backtest's assumptions down.Specifically what it assumes about behaviour staying constant, so the assumption can be checked afterwards rather than argued about.
- Look at the boundary population separately.The aggregate is reliable and beside the point; the people near the rule are the ones the change is about.
- Stage the rollout wherever fairness allows it.Applying a change to new cohorts first is not an experiment, but it does buy an observation window at a fraction of the exposure.
- Set a review date.Otherwise a change becomes permanent by default, which is a decision nobody actually made.