Founder Notes
The Server Told Me Every Time
Three times in one day I hit the same error, and all three times it was a different bug.
The platform our back office runs on answers a rejected write with one catch-all validation status. No field name, no reason, just a status line meaning roughly "something about that was unacceptable." The first time, I was trying to assign work to an account the system had never heard of. The second, a stored phone number had a human aside typed in after the digits, so it failed validation. The third, a free-text field ran a few characters past its length cap.
Three unrelated causes, one indistinguishable surface. I spent an embarrassing amount of that day bisecting by hand: drop a field, retry, drop another, retry. At one point I was creating throwaway records in the live system to find out which column it was objecting to, which is a thing I would flag if I watched someone else do it.
Then I read the response body, and the server had told me every time.
The reason was in there on all three. The field name, the offending value, the real exception underneath. What was discarding it was a helper I had written myself months earlier: it opens the request, parses the body on success, and on failure lets the HTTP error propagate carrying the status line. The body, which is the part with the diagnosis in it, was never read at all.
So the error was not opaque. I made it opaque, and then spent a day working around the opacity as though it were a property of the platform.
The fix was small enough to be embarrassing. Catch the failure, read the body, raise what it actually says. The next rejection named the field and quoted the value, and the repair was one line. Every one of the three would have been one line with the message in hand. What cost the day was none of the bugs. It was the cost of a failure that could not say what it was, paid three times.
There is a second thing in here, and I think it is the more useful half.
Later the same day the same status came back from an operation that was working perfectly. A task declined to complete because something it depended on was still open: exactly the guardrail I wanted, doing exactly its job. At the surface I had built for myself, a correct refusal and a malformed request were the same event. I could not tell "you did that wrong" from "no, and here is why not" without reading the body I was throwing away.
That is worse than slow. An error channel that collapses bugs and guardrails into one signal teaches you to treat refusals as noise, and then you start retrying through them.
The generalisable form, and the thing I have actually changed my mind about:
An error message's quality is a property of the whole path the error travels, not of the system that raised it. I had been quietly grading the platform on its error design, when the last layer before me was mine and it was the lossy one. Every layer that catches a failure and re-raises something smaller is a layer that can throw the diagnosis away, and by default most of them do, because the status is the easy thing to pass along and the body is work.
The operational test is simple. If the same unhelpful failure has cost you time more than once, stop debugging the cause and go fix the reporting. The second occurrence is the signal. You are not short of information. You are discarding it in transit, and you pay for it again every time.