Loop X: An AI Co-worker Built on the CLI I Already Pay For
Claude co-worker and Claude Tag looked great until I did the seat math, so I built Loop X: a self-hosted AI teammate that runs entirely on my own Mac.
I had Claude’s co-worker pitch open in one tab and our team’s GitLab issue backlog open in the other. The pitch was exactly what I wanted: something that works the queue every morning, fixes what it can, flags what it can’t, so nothing assigned to anyone just sits there untouched. I’d already seen Claude Tag do something close to this in Slack, and it’s genuinely good. Then I did the seat math against how much GitLab traffic we actually generate, and it didn’t pencil out next to the Claude Code subscriptions we already had running on every machine on the team.
That’s the part that stuck with me. The capability I wanted wasn’t exotic: list my assigned issues, work them one at a time, fix what’s fixable, comment on what isn’t, send me a digest. claude -p can already do every step of that from a terminal. What I was actually being asked to pay extra for was the scheduling, the safety rails, and the dashboard around it. So I spent the next few weeks building exactly that layer myself, on top of the CLI access I was already paying for, running entirely on my own machine.
What Loop X actually is

It’s a stdlib-only Python project that runs as a couple of launchd agents on my Mac. One is a scheduler that polls every fifteen minutes and fires whichever loops are due. The other is an always-on local dashboard. Every task, issue triage, inbox labeling, CI diagnosis, meeting prep, is a loop definition in loops.json, and each one, when it runs, shells out to claude -p (or codex exec, if that’s the CLI I’ve pointed it at) with a prompt built from the loop’s own instructions file. Nothing is hosted. There’s no account with a third party that sees my issues, my mail, or my calendar; the loop’s state, its history, its drafts all live under ~/.loop-engineering on the machine I’m sitting at.
The flagship loop is GitLab issues. Every weekday at 10am it lists everything assigned to me across every project I’ve configured, and works through them one at a time, never in parallel, which took some getting used to since the instinct with agent work is always “why not fan these out.” For each issue it does exactly one of three things: fix it in an isolated git worktree and open a merge request once the project’s own lint and test commands pass, answer it with a comment if no code change is needed, or escalate with a comment asking for clarification if the ask is ambiguous or verification fails. Then it sends me a Slack digest and updates a PROGRESS.md file, so the next run knows what happened. So do I.
Then I kept adding loops
Once the GitLab loop was running clean for a couple of weeks, I noticed the same shape of problem everywhere else in my day, and started adding loops for each one. All of these are running on my own machine right now, not a pitch deck.
Gmail and Outlook, triaged but never touched. The inbox triage loop reads unread mail since the last run, puts it into one of a handful of categories (Loop/Urgent, Loop/Action, Loop/FYI, and so on), and drafts a reply for anything urgent, left sitting in the mailbox’s own Drafts folder for me to read, edit, and send myself. It cannot send mail. Not “instructed not to.” Cannot: neither the Gmail nor the Outlook integration contains a send function at all, and for Outlook the OAuth token it requests is scoped without Mail.Send, so sending is impossible at the token level even if the code tried. It also can’t archive, delete, or mark anything read. The only writes it’s allowed are creating a label and dropping a draft.
Calendar, fifteen minutes out. The meeting prep loop watches my Google Calendar and, forty-five minutes before anything with another attendee on it, pulls together a brief: the agenda, any GitLab issues or MRs linked in the invite description, recent mail with the organizer and attendees (never me, and never message bodies, just subject/sender/date), and, if it’s a recurring meeting, what the summary and follow-ups were last time. That last part is the one I didn’t expect to matter as much as it does. Recurring meetings drift, and having “here’s what we said we’d follow up on” land in Slack before I’ve even opened my laptop lid has already saved me from walking into a sync without the context from the one before it.
Topics, in plain English. The topic monitor isn’t tied to GitLab or mail at all, it runs live web search against whatever I tell it to watch. Each topic is just a name and a sentence describing what counts as notable for it, so one topic can watch something like AI model releases and funding rounds while another watches a specific project’s changelog, same loop, same code, different brief string. It dedupes against the last 7 days so I don’t get the same story twice, and a quiet day still sends a one-line “nothing new” instead of silently skipping, which matters more than it sounds like it should, because a loop that goes silent on a quiet day is indistinguishable from a loop that’s broken.
RSS, ranked instead of skimmed. The RSS watch loop does the thing I actually wanted an RSS reader to do for years: pull new entries, rank them against a list of interests I typed into a settings box, and send me the ten best with a one-line reason each, instead of a feed count climbing toward four digits I’ll never clear. This one gets an extra layer of suspicion, because feed content is the one input in the whole system that comes from the open internet rather than from an account I own. The prompt tells the model as much, it runs with no tools and no shell access at all, and anything it writes still passes through the same notification sanitizers as everything else before it reaches Slack.
The two things that were actually hard
The first was deciding which loops deserved real tool access at all. GitLab issues needs a shell, git worktrees, the ability to run a project’s own test suite: a lot of surface area, because fixing code requires touching code. But inbox triage, meeting prep, the topic monitor, and RSS watch don’t need any of that. They need one model call over data I’ve already fetched, with an opinion back. I ended up splitting the loop types in two: the full runtime with tools and MCP servers for anything that writes code, and a sealed plugin model (no tools, no shell, isolated per item so one bad entry can’t take down the whole run) for everything that’s really just “look at this and tell me what matters.” That split did more for cost than any prompt tuning did, because most of my daily loop volume turned out to be the second kind.
The second was the trust problem, and it’s the one I underestimated going in, because it shows up differently in every loop. For GitLab, it never merges its own MR, full stop; that’s always a human click, every fix happens on its own branch in its own worktree, and a verification failure escalates instead of retrying. For inbox triage, the “always mark this sender urgent” rule for VIPs isn’t a line in the prompt I’m hoping the model respects, it’s enforced in plain Python after the model responds, and senders on an exclude list are filtered out before the AI ever sees the message at all. For every loop that touches a real account, the rule I kept coming back to is the same one: never trust the prompt alone for anything you’re not willing to have go wrong once. Writing the part of each loop that does the actual work took an afternoon or two apiece. Writing the part that knows what it’s never allowed to do, no matter how well it’s been behaving, took the other three weeks.
The moment any of this stopped feeling like a toy wasn’t dramatic at all, which is exactly why it stuck. I’d open Slack some mornings to four separate digests sitting there: GitLab issues worked overnight, an urgent mail draft or two waiting for a once-over before I sent them myself, a meeting prep brief with the right follow-ups already attached, a topic briefing that actually had something in it that day. Nothing in any of those messages demanded I do anything right away. That was the point. Four queues that used to sit until I had a free hour had already been worked by the time I made coffee.
Making it something anyone can run, not just me from a terminal
I didn’t want this to live only inside my own comfort with a shell. bin/scripts/build_macos_app.sh wraps the whole dashboard in a native window with its own bundled Python environment, so the end result is a Loop X.app you can drag to Applications and open like any other program, no terminal required after the first install. It attaches to the scheduler and dashboard already running in the background, or starts them itself if they aren’t.

It’s the exact same dashboard either way, same buttons, same chat box, same loop list. The only thing that changes is whether it’s sitting in a browser tab at loop.x or in its own window with its own icon in the Dock, and that difference is the one that actually matters for handing this to someone who’d otherwise have to ask what 127.0.0.1:8420 is supposed to mean.
The connector model is what makes adding a new account to any of this a form, not a code change. A connector is just an account with a type (GitLab, Gmail, Slack webhook, RSS feed list, Jira, Notion, a calendar) and the capabilities it offers, issues, mail, notify, feed, and so on; a loop asks for a capability, and any connector of the right type can satisfy it. Adding Telegram or Discord as a notification target, or a second Gmail account, is filling out a form on the Connectors page and clicking Test, not touching a line of code.
And because all of this runs unattended, the part that actually makes it usable day to day is knowing what happened without having to go dig. Every loop posts its own Slack digest the moment it finishes, and the dashboard’s Activity view has an embedded chat assistant I can just ask things like “what’s connected right now” or “what did the GitLab loop do this morning” and get a straight answer pulled from the actual run history, not a guess.
What changed
I went into this thinking the product I was shopping for was “an AI that works your issues.” It wasn’t. The model call was never the scarce part. I already had it, included in a subscription I was paying for regardless. What Claude co-worker and Claude Tag were actually selling was the orchestration around that call: the scheduling, the guardrails, the dashboard, the policy about what an unattended agent is and isn’t allowed to touch. That’s a real product, and a fair one to charge for. It’s also a product I could build myself with the access I already had, because none of those pieces are secret. They’re decisions, not capabilities, and once I’d made them for one loop, making them again for mail, a calendar, and a pile of RSS feeds was mostly repetition.
The part I didn’t expect is how much more of the codebase is boundaries than behavior, and how consistently that held across every loop I added, not just the first one. The code that fixes a GitLab issue, drafts a reply, or writes a meeting brief is each a few hundred lines wrapping a single model call. The code that decides when any of them is allowed to act, when it has to stop and ask, and what it’s never permitted to touch no matter how well things have gone recently, that’s most of the project. If I rebuilt this tomorrow I’d start with the boundaries file first and treat the actual task logic as the easy part, because that’s the order it turned out to matter in, four loops in a row. Loop X is on GitHub at github.com/encoreshao/loop-engineering, MIT licensed.