An AI agent that can open PRs but never merge them
Every production app has a pile of small work nobody wants. A Sentry error that fires twice a day. An alarm with a threshold someone guessed at. A dependency with a security advisory. None of it is hard. All of it waits, because there’s always something more important.
At work I run the team that owns that pile. In September I built an agent to take it on. We call it Groundskeeper. It runs in our AWS account, it picks up that kind of work for one repository, and it opens a pull request for each item.
In its first week it opened 41 PRs. 32 were merged. 2 were closed. The rest were still in review when I wrote this.
The part worth writing about isn’t that an AI can fix a bug. It’s how much of the design went into what it can’t do.
What it does
The loop is simple:
- Something happens. A Sentry alert, a CloudWatch alarm, a ticket moved to a column in Asana, or a review comment on one of its own PRs.
- That becomes a GitHub issue, and the issue goes on a queue.
- When a slot is free, a container starts on ECS with Claude Code inside it.
- The agent writes a test that fails, makes the fix, and runs the repo’s
make verify. - If verify passes, it opens a PR into the integration branch.
- A person reviews it and decides.
It runs one item at a time. That’s on purpose. The goal was never speed. It was a steady stream of small PRs that a human can actually review.
What it can’t do
This list is the real product.
- It can’t merge, approve, or deploy. Its changes reach production the same way anyone else’s do: through a person.
- It never holds a long-lived GitHub key. A separate token broker holds the GitHub App’s private key. For each run it mints a token that lasts one hour and works on one repo.
- It can’t touch certain paths at all. The repo’s
CLAUDE.mdhas a “Never” list. A guard hook checks every shell command the agent runs against it before the command executes. - It can’t use raw
docker,curl, oraws. If it needs logs or Sentry context, it gets them through small purpose-built tools that only return what that run needs.
We handle health data, so this mattered beyond good taste. Every one of these rules maps to a SOC 2 control we already had to satisfy. The agent is just another actor that has to follow them.
A signal is never the thing you fix
This was the rule I underestimated.
Ask a model to make an error go away and there are a lot of easy ways to do it. Lower the log level. Catch the exception and return a default. Raise the alarm threshold. Skip the test. Each one makes the alert stop. None of them fix anything. Worse, they leave production unable to report the real problem the next time it happens.
So the agent’s rules say it plainly: a signal is never the thing you fix. And because rules in a prompt are suggestions, there’s also a diff check. When the agent opens a PR, the check looks for exactly those patterns. If it finds one, the PR gets a needs-human label and the agent has to justify each finding in the PR body.
The same thinking applies to slow pages. A bigger instance or a longer timeout is never the fix. It has to find the N+1 query or the missing index, and prove it with a test.
It kept writing the same code twice
The first few days surfaced a problem I didn’t expect. The agent’s fixes were correct, but it kept adding helpers that already existed somewhere else in the repo. Each PR looked fine on its own. Together they were making the codebase worse.
The fix was a deterministic duplicate-code check. Before the agent can open a PR, the guard compares what the branch added against the rest of the repo. If the new code duplicates something, the PR is refused until it reuses the existing version.
Then I pointed the same check at the whole repo. Every duplication it found became an issue with a proposed consolidation. A good chunk of that first week’s PRs came from this: shared hooks, shared components, shared helpers pulled out of three or four copies.
One bug in the guard
A small one I liked. The guard parses every shell command before it runs. The agent writes PR bodies with heredocs, and the guard was reading the heredoc body as if it were shell. A PR description that mentioned a blocked command would get the whole command refused.
The fix was one sentence in the commit message: heredoc bodies are data, not shell. It’s the same idea as the rest of the project. Know which part of the input is instructions and which part is just text.
How it’s built
| Piece | Job |
|---|---|
| Webhook Lambda | Verifies GitHub signatures and queues work |
| SQS + dispatcher Lambda | Holds the queue, starts one ECS task when a slot is free |
| ECS on EC2 | Runs the agent container: Claude Code, gh, and the guard |
| Token broker | The only thing that holds the GitHub App key |
| Asana bridge | Keeps a ticket for every piece of agent work |
| Status page | Shows what’s running and how recent runs ended |
All of it is OpenTofu. Each package tests itself without network or AWS access, and CI runs all of it on every PR.
What I’d tell someone building one
Start with the boundaries, not the prompts. The prompts changed every day. The boundaries didn’t. Token scope, the guard, the diff check, and human merge are what let me trust it enough to leave it running.
And make the agent prove its work with the same check your team uses. Groundskeeper doesn’t get its own definition of done. make verify passes or the PR doesn’t exist. That one rule did more for review quality than anything I put in a prompt.