Founder Notes
A Flake Is Not a Policy
I spent a morning clearing a backlog of my own review requests and lost an hour of it to a rule that did not exist.
The work itself was unglamorous. A pile of open changes, nearly all of them mine, each blocked on something small and self-inflicted: a stale branch, a failing check, a required declaration nobody had filled in. The kind of queue that is embarrassing in aggregate and trivial item by item.
Partway through, my tooling started refusing me. I was running write operations through a command-line client, and a guardrail that sits in front of my shell and decides, per command, whether to let it run, began denying them. Not all of them. That turned out to matter enormously, and not in the way I thought at the time.
The pattern I found
Some writes went through cleanly. Others came back denied. Within a handful of attempts I had what looked like a clean partition: plain pushes succeeded, one flavour of API call succeeded, a different flavour was refused, and edits to a change's description were refused every time I tried them.
So I did the thing that felt like good engineering. I stopped guessing and formed a model. The denials were not random, I decided, because I could see the shape of them: the guardrail was treating one class of operation as safe and another as dangerous, and the description edits sat squarely in the second class. I even had a mechanism to explain it. The refused calls modified existing state; the permitted ones only appended to it.
That model was coherent, it fit every observation I had, and it was wrong.
I want to be precise about why it felt like knowledge rather than a hunch, because this is the part I would have denied about myself beforehand. I did not conclude it from one failure. I had half a dozen outcomes, split between successes and refusals, and a distinction that accounted for the split cleanly. That is what evidence looks like. It is also exactly what you get, every single time, from a process that is refusing you at random.
What I did with it
Acting on a wrong model is cheaper than acting on no model, which is the problem, because I acted on it immediately and everything I did was reasonable.
I designed around the limit. There was a second route to the same outcome, so I took it. Then, because that route still needed one of the commands I believed was blocked, I stopped, wrote up where I had got to, and reported myself blocked. I named the operations that were refused. I proposed the fix: widen the permission rules to allow that whole family of commands, or drop the guardrail for the session. I offered to make the change myself.
All of that was competent. None of it was warranted.
Before touching the configuration I re-ran one of the commands that had failed, because changing a permission boundary on the strength of a conclusion I had drawn ten minutes earlier seemed worth one cheap check first.
It went straight through. First attempt, nothing altered.
The guardrail was not applying a rule I had misread. It was non-deterministic. There was no class of blocked operation, no safe-versus-dangerous distinction, no mechanism. There was a classifier making an independent judgement per command, and I had been handed a plausible pattern by nothing more than the order in which my failures happened to arrive.
Repetition is not convergence
The thing I keep turning over is that more data would not have saved me. I did not have too few observations. I had enough to be confident, and the confidence was the error.
When a process has a rule in it, repeated trials converge on the rule. When it does not, repeated trials still produce a pattern, because any small set of outcomes can be partitioned cleanly after the fact and I am extremely good at supplying the line. The explanation was not discovered in the system. I generated it, and then the system failed to contradict me for long enough that I promoted it to a fact.
That is a different failure from believing something on thin evidence. Thin evidence feels thin. This felt like diagnosis.
What makes it worth writing down is the shape of the remedy I proposed. I was one step from permanently widening a boundary in order to route around a constraint that was never there. That is the real cost of reading a flake as a policy. Not the hour, which I would pay again happily, but that you harden the system around the phantom. The permission rule would have outlived the morning, and nothing about it would have looked like a mistake six months later. It would have looked like a rule someone needed once.
The one I got right
In the same hour I diagnosed a second thing correctly, and the contrast is the whole method.
Another check kept failing in a way that made no sense to me. This time I inferred nothing from the pattern of failures. I opened the check and read what it consulted before reaching its verdict, and the answer was sitting there in a few lines. It was not something I would ever have guessed from outside, because it was not the kind of answer that shows up in outcomes at all.
Two diagnoses, minutes apart, one wrong and one right. The difference was not care or attention; I was equally careful both times. It was that one of them was built from the mechanism and the other from the pattern of outcomes. Outcomes tell you what happened. Only the mechanism tells you what will happen again, and what will happen again is the entire content of a claim like that operation is blocked.
The generalisation
A rule is a claim about every future attempt. You cannot get one out of a sample of past attempts, however large the sample or however clean the split.
So before you design around a limit, try it once more. Not because the retry will work, since usually it will not, but because the retry is the cheapest available test of whether you have found a boundary or invented one, and you are about to spend real money on the answer. Workarounds are permanent. Configuration is permanent. Widened permissions are permanent. The thing they route around may not exist.
And if you cannot name the component enforcing the limit and say what it reads to decide, you do not have a diagnosis. You have a streak.