Writing Architecture

Reliability is architecture, not prompting.

The failure mode is not that the model is bad. It is that the model is good enough to be plausible and not good enough to be relied upon, which is the worst place on the curve. A system that is right ninety percent of the time and gives no indication which ten percent is wrong is, operationally, a system nobody can use without checking all of it.

The fix is structure rather than persuasion. Anything that can be resolved by rule is resolved by rule. Identity, eligibility, routing, arithmetic — these are lookups and calculations, and handing them to a model converts a deterministic answer into a probabilistic one for no gain. What is left for the model is the genuine ambiguity, and that residue is usually much smaller than it first appears.

Then the residue gets measured. Every output is scored against a labeled set, so accuracy is a number you can look at rather than a claim you have to accept. This is the part teams skip, and skipping it is why so many pilots end with an argument about whether the thing works.

The same discipline applies to agents, and more sharply. An agent that chooses its next action freely will eventually choose wrong, and the blast radius of a wrong action is larger than the blast radius of a wrong answer. An agent that proposes an action and has it validated before execution will not. Propose-then-validate is not a safety feature added at the end; it is the shape of the system.

None of this is exotic. It is the ordinary discipline of building something that runs unattended, applied to a component that happens to be probabilistic. The reason it gets skipped is that prompting produces a demo in an afternoon and architecture does not.

Book a call

Tell us where the work piles up.

30 minutes. We tell you whether it's a fit and what the first phase looks like. No deck.

Book a call

Keep reading

More from the same shelf.