The agent harness I ended up staying with

In my post on agent harnesses in September I said I had been looking at Oh My Pi, Tura and Prime Agent. I still poke at the others. Prime Agent is the one I have stayed with, so here is why.

Two of these I have not found anywhere else. The third one turned out to matter more than I expected.

The two are the continual harness and how easy it is to write a skill. The third is that Python is the only tool.

The continual harness

Most harnesses let you write down things you want to remember. Prime Agent has a layer above that. It looks at what happened across sessions and writes entries into four buckets on its own. Prompt notes, which are policies like how commits should be written. Memories, which are facts and gotchas. Skills, which are procedures with a Python entry point. And subagent specs, which are reusable delegation prompts.

The entries are small and they are versioned, and every change comes with a trigger, the evidence behind it, and the outcome it is supposed to produce. Right now I have 7 prompt notes, 50 memories, 4 skills in the harness, and 34 refinement events that produced them.

I like that it only writes something when there is a reason. A refinement happens after something failed, or after the same useful trick showed up twice, or after you correct the agent about something. I have rules in there now that I only had to ask for once, because the second time I asked, it got written down.

There is a split between local and global entries. Local belongs to one session and global follows you. Most of mine are global, because the way I want Jira edits done does not change between Tuesday and Thursday.

Writing skills

A skill is a folder with a SKILL.md in it. There is no schema to fill in and nothing to register. You write the procedure the way you would write a runbook, with the exact commands, the gotchas and the wrong turns that you know of, and the agent reads it when the task matches.

I have 166 of them now. Most of them came out of doing a task once and not wanting to work out the steps again. The Docker dependency ordering one came from a CI run that rebuilt dependencies from scratch because the layer order was wrong. The CDC corruption one came from an investigation where the application logs had already rotated away, and what saved me was a retained Kafka topic. Neither of those is something I would remember correctly at 2am without a note.

Some of them are just a Python module and a contract, so the agent calls them directly instead of reading prose and re-implementing the logic. The Google search one is a plain function, and the one I wrote to query production Postgres through the bastion is too, because that sequence of steps is long enough that I do not trust a paraphrase of it.

I read a skill before I trust it. The agent writes it after the task, but the agent writing it is not the same as it being right. I open it, and then the next run either works or I fix the skill. The Git history on my config repo is the audit trail.

Python as the only tool

This is the one I did not expect to care about. Every other harness I tried gives the model a set of tools, one action per tool call, and the model decides what to call next. Prime Agent gives it a persistent Python REPL and that is the tool. So the agent writes Python, and in that Python it can loop, keep variables between steps, filter a large result down to the twenty lines that matter, and call the next thing based on what the last thing returned.

The effect is that a result does not have to come through the conversation at all. When I ask about a database, the agent runs the query, and then in the same cell it can group the rows, print the two that matter and hold the rest in a variable for the next question. With one-shot tools that becomes four round trips and a lot of text in the transcript that neither of us needs.

It also makes delegation cheap. Child agents are started from the same Python, so the parent can spawn one, keep working, and check in later. I can inspect what a child did without asking the child to summarize it for me.

There is a cost to this. Sometimes the model writes a lot of Python for something a single tool call would have done, and I have seen it go down a rabbit hole inside one cell where I would rather have stopped it. That has happened more than once and it is the main thing I would want a guardrail for.

A lot of what I have in there now is stuff I would not have bothered to write down on purpose. It gets written after something goes wrong, or after the same trick turns out to be useful twice, and then it is there in the next session and every session after that. For me that is the main reason I stayed.

Notes

  • The harness does not replace knowing how to debug. It replaces re-deriving the same procedure for the fourth time.
  • A skill that was never reviewed is just a guess with good formatting. Read the ones you care about.
  • Skill count is not the metric. A hundred skills where half are wrong is worse than ten you trust.
  • If you write Python all day, the single-tool harness is easier to live with than it looks.

Subscribe to Blog

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
[email protected]
Subscribe