DsgnWrks

The random technical musings of Justin Sternberg

How to set up an autonomous agentic engineering pipeline

My days still look a lot like managing coding agents. The difference is that more and more of the work is being funneled through autonomous agents, and more and more I’m interfacing with one overseer (aka orchestrator, aka conductor) instead of managing every worker myself.

It keeps multiple pieces of work moving, delegates each phase to fresh agents, checks what they report back, and pulls me in when something actually needs me. I can see all the work it’s managing, check in once in a while, and handle the “what’s next for you” question (the one line at the end of every overseer update that says what’s actually mine to do).

It’s also pretty fun.

The why

My days had turned into managing a bunch of agents. Some were building code, some were reviewing issues to see if they were viable to work on, some were reviewing PRs, some were addressing PR feedback, and some were investigating code issues so they could create new issues. I manned each of those sessions. Issue creation, then the work, then QA, then PR review. I was babysitting every one of those agents.

On top of that, when Opus 5 came out, trying to talk to it nearly drove me to insanity. Opus 5 was hated by almost everyone for its verbosity and opaqueness. Thankfully we had also gotten Anthropic’s Fable model a few weeks earlier (with some drama), and Fable was the opposite. It seemed to have good judgment and taste. (I wrote the fable-mode skill to try to help other models adopt its stance.)

So I started having Fable be the communicator for me. It would send instructions to Opus 5, then translate the results into an actionable set of requests for me, making the easier “duh” judgment calls along the way. After working that way for a while, the maestro plugin’s conduct skill was born out of it. (conduct is the operating stance for an agent whose job is managing other agents: it delegates the work, verifies what comes back, and tells me only what I need to decide.)

Why not just use Fable for everything? Different models are good at different things, and Fable eats up tokens really fast (and has its own separate weekly quota). I couldn’t afford to use it for the actual work, which is why I wrote fable-delegate pretty early on. Fable can build software, but that’s not where I want to spend it. It’s the judgment layer. Opus is really good at the build layer, and Sonnet follows instructions really well.

From there came an issue-to-PR runner that builds, self-reviews and triages reviews. I kept Fable on top as the orchestrator for the runs in flight. Its job was to manage them, tell me which ones needed me, and watch for the ways the runner itself was failing so it could assign other agents to fix bugs in the system or improve it. After a while I realized the orchestrator role needed better parameters and needed to be easier to run and hand off to the next orchestrator. That became the overseer and its self-healing workflow.

I’ve talked to my team about it a lot, so I moved it into its own repo where they can try it and even contribute. That repo is private. This post is the public version, so you (or more likely your coding agent) can build the same thing from the tools you already use.

What it is

There are two layers to the issue-to-PR pipeline, plus a parallel review mode:

  1. The overseer keeps a queue of runs moving. It starts a boss per issue, watches the bosses, verifies their reports, pings me when something needs a decision, routes new PR review comments back to the right run, and hands off to a successor before its context fills.
  2. The boss takes one issue (or one line of prose) through build → self-review → create PR → review PR → verify findings → address review → update description → hand-off. It coordinates and never edits code. Each phase gets a fresh agent session in its own pane, run state lives in an agent-driven task tracker, and every run ends with a debrief (a write-up of where the pipeline itself went wrong, so the next fix goes into the system).

The swarm reviews several PRs in parallel, vets every finding, publishes only after I approve, and turns what it learns into rules the next agents inherit (compounding our conventions).

The overseer matters most. Without a frontier model with Fable-level judgment as the overseer, many more decisions bubble up to me that really shouldn’t. Lately, I’ve been rocking Opus 5.5 running fable-mode as my daily driver.

This is a blueprint, not a product.

Something weird has happened with software in the AI code-is-cheap era: I don’t necessarily want your finished tool anymore. I want to know how it works.

If you’ve already solved a problem, the useful thing you can give me may not be the package you built, with your language choices, dependencies, and assumptions baked in. Give me the blueprint and I can hand it to my agent and say, “Build me one of these, but for my setup.”

That’s what I’m trying to do with posts like this and Steal This Tool: dirmap + goto. I’m not really publishing software. I’m publishing enough of the spec, design decisions, gotchas, and acceptance criteria that your agent can build you a bespoke version.

The capabilities (the spec)

Build these from whatever you have. My tool for each is in parentheses, and the public ones are linked.

  • Spawning. One agent starts a fresh agent session for each phase, in a pane you can watch, and hands it a self-contained work order. (Mine: herdr, a terminal multiplexer whose panes keep running after you close the laptop, and cmux, a macOS terminal built for running AI coding agents.)
  • Messaging. Sessions exchange messages and replies, and a caller can wait on a reply without polling in its own context. (Mine: hotline, a plugin that lets one Claude Code session call another and wait for its answer.)
  • Durable shared state. Work, phases, and handoffs live in an agent-owned task tracker every agent reads and writes, so a crashed or rotated session loses nothing. (Mine: Beads, an issue tracker built for coding agents, stored in a version-controlled database.)
  • Remote notifications. Agents ping you when they need a decision or they finish. (Mine: the slack plugin.)
  • Succession. A long-lived overseer writes a handoff and spawns its own replacement before its context fills up.
  • Workspace. One root git repo with every project as a .gitignored sub-repo, so agents find any repo by shortname. The root holds the phase skills, the repo-specific instructions each phase follows: your repo’s create-PR skill, its review skill, its address-review skill, and its update-description skill. Mine are tuned to my work repos, but public versions of the last two are in my pr-workflow plugin (address-pr-comments and update-pr-description). A personal monorepo works fine.
  • Watching. A background job notices new PR review comments and routes them to the right agent, with a cap on back-and-forth rounds.

Why succession

There’s only so much context available to any one agent conversation. Rather than risk the compaction that happens when a session hits that limit, it’s better for the agent to be proactive and hand off the exact information it believes will help the next one. Handing off before compaction also leaves the previous session’s uncompacted transcript intact, so the successor can go back and review earlier decisions and what happened.

Mine hands off at around 50% context, because cumulative token consumption accelerates as the context fills up. Every response sends the whole context back, so later turns get increasingly expensive.

Design rules

Keep these, whatever your tools are.

  • Roles stay separate. The boss coordinates and never edits code, a doer (the fresh session the boss starts for a phase) does one phase, and a verifier who didn’t write a finding tries to falsify it. Agent sessions are inherently biased towards their own work, and this works around that limitation.
  • A doer’s report is a claim. The boss doesn’t trust it until it checks reality (git, CI, the PR) for itself.
  • Each phase gets a fresh session, and the next phase is told what was verified, not what was said.
  • Some calls stay with me. I decide which ones. Right now, the pipeline waits for me before merging, resuming a stalled run, or publishing anything on someone else’s PR. It also opens its own PRs as drafts. You might draw those lines somewhere else. The important part is that whatever you reserve for yourself becomes an explicit gate the pipeline waits on.
  • Every run ends with a debrief that says where the pipeline went wrong.

When things go wrong, the system self-corrects:

  • Three builders independently reported tests passing in a way that looked off. The overseer noticed the pattern across their reports and sent a read-only investigator, which found that some worktrees were loading the main checkout’s code. “Green” sometimes meant “main is green.” The verify step now hard-fails on that setup.
  • A fix for one checkout bug quietly introduced another: it dropped the customer’s discount code. Any PR a builder substantially reworks gets a fresh reviewer before it’s called ready, and that reviewer caught it before a human ever saw the PR.
  • An agent ran pkill -P 0 and took down about 200 processes, including live agent sessions and the task tracker’s database server. Every work order now carries a rule to confirm the exact PID first.
  • The overseer gets it wrong too. It once told me to merge a base PR too early, which would have deleted the branch out from under a stacked run’s remaining review phases. It corrected itself one message later, before I acted.

What’s left for me

Two kinds of work. The first is meta: working with a secondary agent on the system itself. Recently that was making the test suite faster so it didn’t take so long to run. The second is validating the actual choices made on specific PRs. Even when a PR is supposedly ready to go, I start a new review session and have an agent walk me through the work history (with my walk-through-work-history skill) and do real QA testing in the browser (with my qa-walkthrough-pr skill).

The pipeline took away the mind-numbing, repeated workflow of getting work from an issue all the way to a reviewed PR, along with answering benign agent questions, checking their assumptions, and validating their claims. That doesn’t mean I’m spending less time managing agents. Mostly it means I can funnel more work through them, and spend more of my time at the judgment and orchestration layer instead of babysitting each individual step.

And the whole system is meant to compound. Ideally, the more I run it, the more of my own judgment calls I can offload, because every correction we make can become a rule the next agent starts with. For more information about this methodology, see Compound Engineering.

Build order

Smallest working loop first. Stop after each step and see it working before you go on.

  1. Messaging between two sessions
  2. One issue through the pipeline, with durable state
  3. Notifications
  4. Overseer + succession
  5. PR watching
  6. Parallel PR review

If you only steal one piece, steal the overseer: the orchestrator middleman, operating with maestro’s conduct skill.

Acceptance checklist

  • ☐ Two agent sessions can message each other, and the caller waits on the reply without polling in its own context (get hotline plugin/skills)
  • ☐ One issue goes through every phase, each in a fresh session, with state in your tracker (get herdr, Beads)
  • ☐ Kill a phase’s session mid-run and the run resumes from the tracker without losing anything (get Beads)
  • ☐ A verifier that didn’t write a finding gets a chance to falsify it before anyone acts on it
  • ☐ The PR lands the way your normal flow expects (draft or live, reviewers, CI), and the calls you kept for yourself wait on your OK
  • ☐ You get a ping when a run needs a decision and when it finishes (get slack plugin/skills)
  • ☐ An overseer watches several runs at once, verifies their reports, and brings you only the calls that need you (get maestro plugin/skills)
  • ☐ The overseer hands off to its successor before its context fills, and the successor picks up the queue (get handoff plugin/skills)
  • ☐ A new review comment on a PR gets routed back to the run that owns it (start from the watch-pr-then-action skill, which watches a PR and runs an action; routing to the owning run is yours to add)
  • ☐ Every run ends with a debrief that says where the pipeline went wrong
  • ☐ You hand one small issue to the boss, walk away, and come back to a PR, a verified review, a notification and a debrief, with no manual steps in between

Steal this system

Hand this to your own agent:

Read this post: https://dsgnwrks.pro/tools/autonomous-agentic-engineering-pipeline/
It describes one engineer's system for running coding agents unattended: an overseer
keeps a queue of GitHub issues moving, a boss agent takes each issue through build →
self-review → PR → review → verify findings → address review → update description →
handoff, and a swarm mode reviews several PRs in parallel.

I want the same system, built from the tools I already use. Where my setup has gaps,
I'm open to adopting the tools the post recommends.

Before you build anything:
1. Read the post's capabilities and design rules. The linked plugins (hotline, maestro,
   slack, fable) are public reference implementations if you need to see one working.
2. Interview me, one question at a time, about my stack: terminal/multiplexer, task
   tracker, notification channel, code host, agent CLI, OS. For each capability, propose
   what I already have that could do the job, and only what's missing to install.
3. Write a plan mapping each capability to my tool, and wait for my OK.

Then build it in the post's build order, smallest working loop first. Stop after each
step, show me it working, and only then go on.

Done when every item on the post's acceptance checklist passes.Code language: Markdown (markdown)

Spec derived from my private agentic-dev repo. Rebuild it with whatever tools your agent likes.

Leave a Reply

Your email address will not be published. Required fields are marked *