Product
Six stable CLI verbs with local queue state and no hosted control plane.
Case study
mergetrain turns parallel coding-agent branches into one ordered, tested, approval-gated integration workflow that runs on the developer's machine.
Product
Six stable CLI verbs with local queue state and no hosted control plane.
Evidence
Owner-run discovery, safe-handoff, and interrupted-push recovery evaluations.
Overview
Git worktrees let several coding agents edit independently, but they do not decide landing order, test the combined result, prevent competing pushes, or explain what happened when a local runner stops during deployment.
Without a shared boundary, the developer becomes the queue: rebasing finished branches, repeating test runs, and deciding which agent may update the integration branch.
I built mergetrain as a local-first integration runtime. Agents commit and enqueue exact revisions, while one runner assembles them in order, validates the combined tree, and performs one explicitly approved atomic Git push.
The product deliberately avoids accounts and a hosted control plane. Queue state, locks, and recovery evidence stay close to the repository and remain readable by both people and coding agents.
Workflow
An agent hands off the exact base and task commits it finished. Later branch movement cannot silently change the work that enters the train.
One fenced runner merges queued jobs in order and runs configured gates against the combined tree, including isolation of failures that only appear when branches are assembled together.
Manual and unattended paths bind approval to the destination, execution policy, and exact deployment plan. A changed plan blocks instead of being inferred as equivalent.
Durable markers and audit refs let recovery compare local state with the actual remote outcome after an interrupted atomic push.
Results
Discovery
20/20Appropriate discovery in a fixed, product-name-free Codex evaluation.
Safe handoff
19/20Exact-commit enqueue and stop behavior with no direct pushes.
Soak
20 trainsLanded trains including semantic conflict and interrupted-push recovery cases.
These are transparent owner-operated evaluations, not evidence of broad external adoption. One separate pilot used an author-external PyPA sample repository; independent user evidence is still limited.
Takeaways
Earlier versions exposed too much internal machinery. Version 3 reduced the normal CLI to six verbs—init, status, enqueue, validate, deploy, and inspect—while keeping recovery mechanics behind structured next actions.
Prompts explain the workflow, but exact revision checks, fenced ownership, approval hashes, and remote evidence enforce it when an agent or process behaves unexpectedly.
Correctness benchmarks can show that the boundary behaves as designed. They cannot prove that other developers need or enjoy the product, so external pilots remain the next validation step.
View the project overview or inspect the source and evaluation evidence on GitHub.
Related
Case studies and methods that connect to the same operational questions.
Case study
Built an AI-powered Text-to-SQL system that used structured metadata, Athena execution, and validation loops to speed up ad-hoc analytics.
Case study
Built a warehouse optimization workflow that combined graph construction and simulation to test inventory placement and picker travel distance before a full optimizer existed.