Putting Reins on the AI
What sets the ceiling is not how you ask — it is the layer outside the model.
For a long stretch, my way of getting more out of working with AI was to say things better: clearer prompts, fuller background, a sterner or a more courteous tone. It works, and it hits a ceiling fast.
I was stuck at that ceiling for a while, with no sense of which way to go.
The turn came by accident. I asked an AI to help me take stock of the projects in my hands, and it pointed straight at my team’s flagship project as the one whose architectural design carried the most substance. I caught a faint signal in that — on what grounds had it singled out that project? So I went and took it apart myself, compared it against the others item by item, and finally found the decisive clue in a very early commit; following that clue led me to superpowers, a powerful architecture-design tool. The door to a new world opened at exactly that moment.
From there I traced my way to a whole series of tools of the same kind and dissected them one by one — not only their architecture but, more importantly, the methodology by which they had been designed. That is not the same order of learning as scrolling engineering blogs and trying whatever a writer recommends.
The difference is in the order. I have always kept up with new technology and research, and I test anything that interests me on the spot; but this time the pain came first — being stuck at the efficiency ceiling of working with AI — and I went looking for the cause with a target in mind.
In truth, before the AI singled that project out, I already knew its design followed a method; I simply could not see where the method came from. I tried inversion — the thing we geophysicists are best at — and it got me nowhere. Intuition told me only this: something is there, and it is the key to the whole toolset; I did not know its name, and I did not know what words to search with. So all I could do was follow my nose after the faintest of traces, and then, like hunting, bite down hard on that fleeting chance and tear a new world open.
Agent = model + reins
Once the dissection was done, one equation was left in my head: agent = model + harness. The model is the capability; the harness is how that capability is allowed to be used. There are only so many models to switch between, and anyone can switch them; almost all of the difference falls in the second term.
A harness governs four things in particular — this is the skeleton I read out of that architecture:
- Permissions. What the agent may read, what it may change, whether it may touch a real system: decided from outside, not left to its judgment in the moment. Anywhere you do leave to its judgment in the moment, one of those judgments will eventually be wrong.
- Layered memory. Not stuffing everything into the context, but layering it: what stays resident, what is retrieved on demand, what gets written back for the long term. A fuller context is not a smarter one — usually the opposite.
- Evaluation. Without evaluation there is no “it got better,” only “it feels better.” This is the easiest one to skip, because it is the one that least resembles building a product — and it is the only thing that keeps the direction of a change off gut feeling.
- Observability. What it actually called, at which step it took the wrong turn, how much it spent: there has to be a trace you can go back to. The value of a post-mortem depends entirely on whether a trace was kept at the time.
Not one of those four is “better at chatting.”
My own part taught me something else
A word, in passing, about another judgment of mine on that project that was worth something. Besides the bridge-monitoring algorithms, I was responsible for the product design and the business model — and the most consequential call was this: I persuaded my teammates to accept a large change, that the AI assistant should grow inside the product’s existing monitoring interface rather than become a separate assistant page.
That call did not come from technology; it came from the customers I had met in industry. Nobody pays separately for an isolated assistant. They pay to have their own job go better. So the assistant has to grow on the piece of interface they already open every day.
What followed proved me right. Customer feedback was very positive, and the project took us to the championship at the finals of the DFWIT AI Competition 2026, first place with 83.5 — the highest score among the thirteen finalists — along with a recommendation into the organizer’s Venture-Ready accelerator program.

Ever since, whenever I look at an agent product my first question is always the same: does it grow inside the workflow, or does it float alongside it?
What I did with it
Once the dissection was done, I did not go and reproduce a harness. I put what I had learned onto a different axis: I redesigned the skills I had scattered around over the past few years into a systematic architecture toolkit, murDrift — keeping a single agent running at quality, while making sure a long project does not drift across conversations, across workspaces, across months.
Because drift is not a runtime problem: windows retire, contexts fill up, architecture decisions get forgotten, and acceptance slowly turns into self-congratulation. Those four are caught by governance, research, continuity and cheat-proof acceptance respectively. The multi-window part of it I have written about in another article, Human + AI: A Multi-Window Operating System.
A last word on the term “reins.” It reads easily as distrust; my own sense is the exact opposite: it is because the reins are there that I dare let it run a little faster.