Claim the issue before you cut the branch
Shipped
0.8.0 lets several sessions work one repo at the same time. start refuses an issue another session already holds, either through a run directory on this machine or a marker comment on the issue itself, and prints where to resume. board grew a Run column so a claim is visible before anyone picks. A fresh lane fetches its base branch before the worktree is cut. The override is its own flag, --take-over, and it archives what it displaces instead of deleting it.
The general problem is older than agents: several workers, one queue, no server to hand out tickets. What follows is the version of that I would build again, with GitHub as the shared board and the filesystem as the referee.
Two places a claim has to live
The plan was simple. Open a terminal per issue, start a session in each, let them run. Each run already had its own directory, its own worktree, its own branch and its own comment on the issue. What nothing did was ask whether that identity was taken.
A claim has to answer two different sessions:
- A session on the same machine. It shares the run root, so a file can stop it.
- A session anywhere else. It shares nothing but the issue, so the issue has to carry the claim.
Keep both. The remote half is what a second laptop sees; the local half is the only one that can be made atomic. The run root here is ~/.agentruns/<owner>__<repo>/issue-<n>, overridable with AGENT_RUN_ROOT so the walkthrough can run in a scratch directory.
Step 1: the local claim is an exclusive create
The obvious version checks whether run.json exists and then writes it. Two starts inside that window both pass. Node’s fs exposes a flag for exactly this: 'wx' is documented as “like 'w' but fails if the path exists”, and underneath it is O_CREAT | O_EXCL, which POSIX requires to check and create atomically with respect to every other open() naming the same file. There is no check-then-write; the kernel does both in one step, and the loser gets EEXIST.
// claim.mjs
import { mkdirSync, writeFileSync } from 'node:fs';
import { execFileSync } from 'node:child_process';
import { join } from 'node:path';
import { homedir } from 'node:os';
const ROOT = process.env.AGENT_RUN_ROOT ?? join(homedir(), '.agentruns');
export function runDir(owner, repo, number) {
return join(ROOT, `${owner}__${repo}`, `issue-${number}`);
}
// The local half. The filesystem is the referee: 'wx' creates the file
// only if it does not exist, and the second writer gets EEXIST.
export function claimLocally(dir, run) {
mkdirSync(dir, { recursive: true });
try {
writeFileSync(join(dir, 'run.json'), JSON.stringify(run, null, 2) + '\n', { flag: 'wx' });
} catch (err) {
if (err.code !== 'EEXIST') throw err;
throw new Error(
`a run already exists at ${dir}. ` +
`Resume it with \`next --run-dir ${dir}\`, or start over on top of it with --take-over.`,
);
}
}
The refusal names the resume command. Most of the time the person hitting it is the same person who started the other session an hour ago, and what they want is to get back into it, not to start over.
Step 2: the remote claim is a comment the next session can read
Each run posts one comment on its issue and keeps rewriting it as a checkpoint. Put a marker on its first line, <!-- runner:run owner/repo#42 -->. On the rendered page it is invisible: CommonMark treats a line beginning with <!-- as the start of an HTML block and passes it through unparsed, so GitHub shows the human text under it and nothing else. In the JSON body it is a string any session can search for.
Reading it costs nothing extra if you already list issues. gh issue list takes comments as a --json field, so one call returns every open issue with every comment body, which is what a board needs anyway. Two more rules make the marker safe to trust:
- Only the region of the body before the first
<details>counts. The checkpoint archives the run’s approved artifacts in collapsed blocks, and an archived plan can quote any string, including your finished marker. - A finished run writes a second marker,
<!-- runner:finished -->, on its own line. A laterstartreads that as a completed record, not a live claim, so a reopened issue is not refused by its own history forever.
// claim.mjs (continued)
export const markerFor = (owner, repo, number) => `<!-- runner:run ${owner}/${repo}#${number} -->`;
export const FINISHED_MARKER = '<!-- runner:finished -->';
const commentIdFromUrl = (url) => Number(String(url).split('#issuecomment-')[1]) || null;
// Only the part of the body before the first <details> counts. Everything
// after it is archived artifacts, and an archived plan may quote anything.
const runRegion = (body) => String(body ?? '').split(/\n<details>/)[0];
export function claimedIn(comments, owner, repo, number) {
const mine = markerFor(owner, repo, number);
for (const c of Array.isArray(comments) ? comments : []) {
const region = runRegion(c.body);
if (!region.startsWith(mine)) continue;
return {
url: c.url,
commentId: commentIdFromUrl(c.url),
finished: region.includes(FINISHED_MARKER),
};
}
return null;
}
// One `gh` call for the whole board; filter per issue in memory.
export function listIssues(owner, repo) {
const out = execFileSync(
'gh',
['issue', 'list', '--repo', `${owner}/${repo}`, '--state', 'open', '--limit', '200',
'--json', 'number,title,comments'],
{ encoding: 'utf8' },
);
return JSON.parse(out);
}
claimedIn returns the comment’s URL on purpose. A refusal that says “claimed” and stops is a dead end; one that links the comment tells the second person whose run it is and what state it reached before they decide anything.
Step 3: start, in the right order, with an override that archives
Check the remote claim first, because it is the only half another machine can see. Then take the local claim, because it is the only half that is atomic. The override is a separate flag rather than a second meaning for --force: taking over someone’s run is a different decision from skipping a drift check, and an unattended run should never be able to pass it.
What take-over does with the displaced run matters more than the flag. Move it aside into superseded/<timestamp>/ inside the same run directory and print the path. A plan, an evidence file and four rounds of review are hours of work; the person who wins the take-over may want them, and the person who lost certainly does.
// start.mjs
import { existsSync, mkdirSync, readdirSync, renameSync } from 'node:fs';
import { join } from 'node:path';
import { runDir, claimLocally, claimedIn, listIssues, markerFor } from './claim.mjs';
// Never delete a displaced run. Move everything it wrote aside and say where.
export function supersede(dir) {
const stamp = new Date().toISOString().replace(/[:.]/g, '-');
const archive = join(dir, 'superseded', stamp);
mkdirSync(archive, { recursive: true });
for (const name of readdirSync(dir)) {
if (name !== 'superseded') renameSync(join(dir, name), join(archive, name));
}
return archive;
}
export function start({ owner, repo, number, takeOver = false, issues = null }) {
const dir = runDir(owner, repo, number);
const issue = (issues ?? listIssues(owner, repo)).find((i) => i.number === number);
if (!issue) throw new Error(`#${number} is not an open issue on ${owner}/${repo}`);
// Remote claim first: it is the only half another machine can see.
const claim = claimedIn(issue.comments, owner, repo, number);
if (claim && !claim.finished && !takeOver) {
throw new Error(
`#${number} is claimed by another session: ${claim.url}\n` +
`Resume it there, or pass --take-over after reading that comment.`,
);
}
let archived = null;
if (takeOver && existsSync(join(dir, 'run.json'))) archived = supersede(dir);
// Local claim second: the exclusive create is the tiebreaker on one machine.
claimLocally(dir, {
owner, repo, number,
startedAt: new Date().toISOString(),
comment: claim?.finished ? null : (claim?.commentId ?? null),
});
return { dir, archived, marker: markerFor(owner, repo, number) };
}
The issues parameter is the seam that keeps this testable offline. A real run leaves it null and the gh call fills it; the walkthrough below passes fixtures. The real tool posts the marker comment right after the local claim succeeds, with gh issue comment, and the checkpoint rewrites that same comment by id from then on.
Step 4: fetch the base before the branch
Claiming the issue does not help if the lane is cut from a stale base. Every worktree in a run starts from origin/<base>, and that ref is only as fresh as the last fetch that happened to cover it. Fetch it explicitly, with a forced refspec, right before git worktree add:
// worktree.mjs
import { execFileSync } from 'node:child_process';
export function fetchBase(repoPath, base) {
const refspec = `+${base}:refs/remotes/origin/${base}`;
try {
execFileSync('git', ['-C', repoPath, 'fetch', '--quiet', 'origin', refspec], {
stdio: ['ignore', 'ignore', 'pipe'],
env: { ...process.env, LC_ALL: 'C' }, // match git's English stderr, not a translation
});
} catch (err) {
if (/couldn't find remote ref/i.test(String(err.stderr ?? ''))) return; // a base that only exists locally
throw err;
}
}
export function cutLane(repoPath, dir, branch, base) {
fetchBase(repoPath, base);
execFileSync('git', ['-C', repoPath, 'worktree', 'add', '-b', branch, dir, `origin/${base}`], { stdio: 'inherit' });
}
Two parts of that refspec are load-bearing. The explicit src:dst is what makes the fetch update refs/remotes/origin/<base> in a clone whose configured refspec does not cover that branch, and --single-branch clones are exactly that: later fetches update only the remote-tracking branch they were cloned with. The leading + overrides the fast-forward rule, so a base someone force-pushed still lands. I wrote up how a single-branch clone made two repos look stale in the press 0.5.1 entry; this is the same fix, applied before every lane instead of after the fact.
Use it: two sessions, one issue, no network
Drop the three files above into a directory and add this driver. It runs four starts against fixtures: a clean issue, the same issue again, an issue another machine has marked, and a finished claim that gets taken over.
// try.mjs — two sessions, one issue, no network
import { readdirSync } from 'node:fs';
import { start } from './start.mjs';
import { markerFor, FINISHED_MARKER } from './claim.mjs';
const owner = 'acme', repo = 'widgets', number = 42;
const noClaim = [{ number, title: 'flaky build', comments: [] }];
const say = (label, fn) => {
try { console.log(`${label}: ok`, fn()); }
catch (err) { console.log(`${label}: refused\n ${err.message.split('\n').join('\n ')}`); }
};
say('session A start', () => start({ owner, repo, number, issues: noClaim }).dir);
say('session B start', () => start({ owner, repo, number, issues: noClaim }));
const live = [{ number, title: 'flaky build', comments: [{
url: 'https://github.com/acme/widgets/issues/42#issuecomment-1001',
body: `${markerFor(owner, repo, number)}\n**Run** started 2026-09-06\n\n<details><summary>plan</summary>\n\n**Finished** is a word this plan happens to use.\n</details>`,
}] }];
say('session C, another machine', () => start({ owner, repo, number, issues: live }));
const finished = [{ number, title: 'flaky build', comments: [{
url: 'https://github.com/acme/widgets/issues/42#issuecomment-1001',
body: `${markerFor(owner, repo, number)}\n${FINISHED_MARKER} **Finished** 2026-09-07\n`,
}] }];
say('session D, claim finished, take over', () => {
const r = start({ owner, repo, number, issues: finished, takeOver: true });
return { archived: r.archived.split('/').slice(-2).join('/'), keptNow: readdirSync(r.dir) };
});
AGENT_RUN_ROOT="$PWD/runs" node try.mjs
This is the output from my run, with the scratch path shortened to .:
session A start: ok ./runs/acme__widgets/issue-42
session B start: refused
a run already exists at ./runs/acme__widgets/issue-42. Resume it with `next --run-dir ./runs/acme__widgets/issue-42`, or start over on top of it with --take-over.
session C, another machine: refused
#42 is claimed by another session: https://github.com/acme/widgets/issues/42#issuecomment-1001
Resume it there, or pass --take-over after reading that comment.
session D, claim finished, take over: ok {
archived: 'superseded/2026-09-07T19-13-00-969Z',
keptNow: [ 'run.json', 'superseded' ]
}
Session C is the one to look at. The word Finished sits inside the comment’s <details> block, and the claim still reads as live, because only the run region is searched.
Then race it. Eight starts at once, all on a clean run root:
rm -rf runs race.log
for i in 1 2 3 4 5 6 7 8; do
AGENT_RUN_ROOT="$PWD/runs" node -e "import('./start.mjs').then(m => {
try { m.start({ owner: 'acme', repo: 'widgets', number: 42, issues: [{ number: 42, comments: [] }] }); console.log('won') }
catch { console.log('refused') } })" >> race.log &
done
wait; sort race.log | uniq -c
7 refused
1 won
One winner and seven refusals is the property the whole design rests on. If you ever see two, the run root is on a filesystem where O_EXCL is not honored, which is the first gotcha.
Gotchas
The exclusive create closes one door, and the override has to be designed around it. The obvious shape is “look for run.json, then write it”, and that is a window. wx closes it for two starts that resolve to the same directory. The red team on the plan found two things around it. Round one: the plan had --force as the single override for both refusals, but an EEXIST from the exclusive create raised the same refusal, so --force was inert in exactly the case it was most needed. The answer was a separate flag, --take-over, that archives the old run before the create instead of trying to get past it. Round three: two sessions passing different --run-dir values for one issue both fetch a payload with no marker, both pass, and the loser’s checkpoint still rewrites the winner’s comment. That stayed a documented limit rather than a fix, because the default run directory is derived from the issue and nobody passes two.
O_EXCL is only atomic where the filesystem says so. The Linux open(2) page warns that on NFS the flag is honored only with NFSv3 or later on a 2.6 or later kernel, and that programs relying on it for locking elsewhere contain a race. A run root on a network mount is the one place this design quietly degrades to the check-then-write it replaced. Keep the run root on a local disk.
Search the run region, and accept the old shape of the finished line. Runs finished before the marker existed wrote a bare **Finished** line with nothing in front of it. If the new reader only accepts the marked form, every issue those runs closed refuses start after the upgrade, because their comments read as live claims. The reader accepts both forms, but only in the region before the first <details>, since the archived artifacts under it are free text.
The destructive path is where the review findings came from. The pull request went through five review rounds, and most of what each round found was in the take-over cleanup the previous fix had added: removing a worktree that was already gone, force-deleting a branch with unpushed commits, honoring a --run-dir that named a different issue. The escape was to make every step reversible or refused. Displaced artifacts move to superseded/, git branch -d runs first and -D only when the caller asked for it, and a run directory whose run.json names a different owner/repo#N is refused before anything moves. Budget extra review for any override you add to a tool that otherwise never deletes.
Gate the network call on the flag, not on the run’s recorded state. The red team’s third round flagged that the new fetch was keyed off run.offline, so an --offline invocation against a run that had been started online still hit the network. If a flag promises “no network call”, the check has to read that flag at the call site. Anything else means the promise depends on how the run began.
Sources
- Node.js fs: file system flags —
'wx'is “like'w'but fails if the path exists” - POSIX open() —
O_CREAT | O_EXCLchecks and creates atomically against otheropen()calls on the same name - open(2), Linux man-pages — the NFS caveat on
O_EXCL - git-fetch: refspecs — the
+prefix overrides the fast-forward rule;src:dstnames the local ref to update - git-clone: --single-branch — later fetches update only the remote-tracking branch the clone was made with
- gh issue list —
commentsis a--jsonfield, so one list call carries every claim - CommonMark 0.31.2: HTML blocks — a line starting with
<!--opens a raw HTML block, so a marker comment renders to nothing
Changelog
- issueflow 0.8.0 — parallel sessions on one repo: start refuses a claimed issue, board shows claims, lanes fetch their base (#252) (04b1195)