Overview
Claude Code is good at planning and reviewing; Codex is good at implementing. This project connects them: Claude plans a change and dispatches tasks through MCP tools, Codex “workers” implement them in parallel, and a local service tracks every job, runs the repository’s checks on the result and reports a verdict before anything is merged.
The engineering challenge
AI coding agents are long-running, can fail halfway, and must not trample each other’s files or the developer’s working tree. Running several at once needs the same things as any job system: durable state, a scheduler, cancellation that actually stops work, recovery after a crash, and safe retries.
Architecture
- Thin MCP adapter → durable service. Claude Code talks to a stateless MCP adapter, which calls a detached local service over an authenticated, localhost-only HTTP API. Closing Claude Code never kills running jobs.
- One runner process per job. The service forks a runner for each running job, which drives Codex through its SDK and streams events back. Recording each runner’s PID and creation time lets the service cancel or clean up exactly its own processes.
- SQLite state. Workers, jobs, events, idempotency keys and workspace locks live in SQLite with schema migrations. Every state change and its event row are written in one transaction, so a crash never leaves a half-recorded job.
- Isolated workspaces. Each writing worker gets its own Git worktree and branch, so parallel agents can’t conflict and the developer’s checkout is never touched.
Key decisions
- A real scheduler: one active turn per worker, follow-ups queued in order, and a global concurrency cap. Queued jobs are claimed with a compare-and-set, so a cancelled job can never start.
- Cancellation that holds up: a job is marked cancelled only after its process has exited and its last messages have drained. If it doesn’t exit, the service kills the whole process tree. A job that finished just before the cancel arrived is honestly reported as completed.
- Crash recovery without surprises: on restart, jobs that were running become
interruptedand are never re-run automatically; queued jobs simply continue. - Idempotency keys on create and submit calls, so a retried request can’t duplicate work.
- Baseline-aware checks: the service runs the repository’s quality gate on each worker’s result and compares it with the gate at the worker’s starting commit, so pre-existing failures aren’t blamed on new work.
- User-controlled model policy and a live dashboard showing workers, jobs, events and gate results.
- Security: a random bearer token with user-only file permissions,
Host-header validation against DNS rebinding, and the token never appears in logs or tool output.
Verification
Backend behaviour (thread creation, resume after a fresh process, streaming, cancellation mid-command) was verified with spike scripts before the design was fixed. The test suite covers worker lifecycle, integration, the MCP surface, model policy, quotas, gate baselines and profiles.
What I’d do next
Reattaching to a running turn after a service crash (the SDK doesn’t support it yet), and a fork operation for branching a worker’s conversation.