Show filesystem activity without calling it completion
Shipped
issueflow v0.5.0 is about one experience: you dispatch a stage of a gated pipeline to a background worker, and then you stare at a status board that says briefed for minutes at a time. This release adds a live elapsed clock for the stage you are waiting on, a position line at the moment of dispatch, an expected-duration range read from past runs of the same repo, and an optional per-stage progress log. None of it required the worker to phone home. This historical release used artifact mtimes as delivery timestamps. The corrected observer below treats them only as file activity. For a complete producer/consumer completion protocol, use the bounded review guide: it binds a complete result to the intended run, round and revision before accepting it.
The dead-air problem
Any orchestrator that hands work to a slow subprocess has this gap. The worker is an LLM agent in my case, but a test suite, a migration, or a render farm behaves the same way: the orchestrator marks the stage started, the worker goes quiet, and the user is left deciding between “still going” and “wedged” with no evidence either way. Jakob Nielsen’s response-time limits put the ceiling at about ten seconds before people need real feedback, and his guidance for anything longer is a progress indicator that shows the wait is moving. A multi-minute stage with a frozen status word fails that by two orders of magnitude.
The standard fix is a heartbeat: have the worker report in, the way the health check API pattern has a service expose /health for a monitor to poll. That works, but it puts the burden on the worker’s cooperation, and the pattern’s own drawbacks section admits the gap: a check is only as fresh as the last poll, and a worker can die between polls. For a subprocess you do not control, there is a cheaper signal. The worker’s job is to produce files. Files have modification times. The filesystem exposes a change signal, but a file can be stale, partial, copied or touched by another process. Its mtime proves neither worker liveness nor successful completion.
A stage’s clock is two timestamps
The state file records when each stage was handed to a worker and when its output was accepted. That pair is the whole data model this technique needs. Here is the core, in a file called liveness.mjs:
import { statSync } from 'node:fs';
// Render a millisecond span one way, everywhere: "42s", "4m07s".
export const formatSpan = (ms) => {
const total = Math.round(ms / 1000);
return total < 60 ? `${total}s` : `${Math.floor(total / 60)}m${String(total % 60).padStart(2, '0')}s`;
};
// A finished stage's duration: briefed to delivered.
export function durationOf(entry) {
const { briefed, delivered } = entry.at ?? {};
if (!briefed || !delivered) return null;
const ms = Date.parse(delivered) - Date.parse(briefed);
if (!Number.isFinite(ms) || ms < 0) return null;
return formatSpan(ms);
}
// A running stage's elapsed time: briefed to now. The "+" marks a
// lower bound, "at least this long", never a finished duration.
export function elapsedOf(entry, now) {
const { briefed } = entry.at ?? {};
if (!briefed) return null;
const ms = Date.parse(now) - Date.parse(briefed);
if (!Number.isFinite(ms) || ms < 0) return null;
return `${formatSpan(ms)}+`;
}
durationOf needs delivered to be recorded, and in the old code it was written in exactly one place: the accept step, after a human approved the output. Which means the one stage a user is actively waiting on could never show a duration. An artifact might have appeared earlier, but that does not establish when a complete result was produced. Keep acceptance time and observed file activity separate; historical durations are comparable only when they use the same explicitly defined acceptance event.
Record activity without changing the delivery state
The observer reads the artifact’s mtime into a separate display-only field. It does not mutate the run or fill in delivered. It performs a filesystem read, so it is not a pure function in the strict sense:
const mtimeOf = (path) => {
try {
return statSync(path).mtime.toISOString();
} catch {
return null;
}
};
// Observation is display data, never an acceptance or completion event.
export function observe(run, artifactPathOf) {
const observed = { ...run, stages: run.stages.map((s) => ({ ...s, at: { ...s.at } })) };
for (const stage of observed.stages) {
const artifact = artifactPathOf(stage);
stage.artifactModifiedAt = artifact ? mtimeOf(artifact) : null;
}
return observed;
}
Node’s fs.Stats exposes the last data-modification time. It is not a producer-completion timestamp. The same value can accompany incomplete or stale bytes, and timestamp resolution varies by filesystem.
Render artifactModifiedAt as “artifact modified,” with its age if useful. A stage without an accepted result keeps its elapsed clock even when an artifact exists. Only the designated acceptance path records delivered, after validating the expected complete result. A clock advancing means time passed; it does not prove work advanced.
Set the expectation from the runs beside this one
An elapsed clock answers “how long have I waited”. The question underneath it is “is seven minutes normal”, and the honest answer lives in the durations of past runs. In this layout every run of a repo is a sibling directory, so history is one readdir away. This also goes in liveness.mjs:
import { existsSync, readdirSync, readFileSync } from 'node:fs';
import { dirname, join } from 'node:path';
const median = (sorted) => {
const mid = Math.floor(sorted.length / 2);
return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2;
};
// Past stage durations from this repo's other runs: sibling dirs of
// runDir, each holding a run.json. A sibling that cannot be parsed is
// skipped, never thrown; a history scan degrading to "no history"
// must never crash the dispatch it decorates.
export function readTimings(runDir, { schema = 2 } = {}) {
const parent = dirname(runDir);
let names;
try {
names = existsSync(parent) ? readdirSync(parent) : [];
} catch {
return [];
}
const byStage = new Map();
for (const name of names) {
const dir = join(parent, name);
if (dir === runDir) continue;
const statePath = join(dir, 'run.json');
if (!existsSync(statePath)) continue;
let run;
try {
run = JSON.parse(readFileSync(statePath, 'utf8'));
} catch {
continue;
}
if (run?.schema !== schema) continue;
for (const entry of run.stages ?? []) {
const { briefed, delivered } = entry.at ?? {};
if (!briefed || !delivered) continue;
const ms = Date.parse(delivered) - Date.parse(briefed);
if (!Number.isFinite(ms) || ms < 0) continue;
if (!byStage.has(entry.id)) byStage.set(entry.id, []);
byStage.get(entry.id).push(ms);
}
}
return [...byStage.entries()].map(([stage, values]) => {
const sorted = [...values].sort((a, b) => a - b);
return {
stage,
n: sorted.length,
min: formatSpan(sorted[0]),
median: formatSpan(median(sorted)),
max: formatSpan(sorted[sorted.length - 1]),
};
});
}
Two decisions in there matter more than the code. First, the summary is a spread, min, median, max, not a mean. Google’s SRE book makes the case that averages hide the tail: “If you run a web service with an average latency of 100 ms at 1,000 requests per second, 1% of requests might easily take 5 seconds.” Stage durations are exactly that kind of skewed distribution, and a range with a median tells the user something a mean would lie about. Second, there is no pooling across repos. Stage duration is dominated by codebase size and test-suite runtime, which are properties of the repo; a number pooled from someone else’s repo is confident and wrong. Below two samples, the dispatch line says “no past timings on this repo”, and that honesty is part of the feature.
What the user sees
At dispatch, the brief now prints position and expectation before the wait begins. With a few past runs on the repo, the two lines are shaped like this:
Step 2 of 4 · 1 approved · investigate → [design] → implement → test
design on this repo: 3 past runs, 2m34s–6m01s (median 3m40s). It unblocks implement.
While a stage runs, the status board shows a table with Since (always populated, from the same clock the durations use) and the last line of an optional per-stage progress log the worker may append to, with its age. A worker that never writes the log still shows a real Since and a dash for progress. That degradation is the contract: the mechanical clock is a waiting-time display, and the log is enrichment on top of it. A worker that ignores the instructions must look quiet, never dead.
To check the boundary, create an artifact without accepting a result. Both stages must retain their elapsed clock; only the first has an observed file timestamp:
// check.mjs
import { mkdirSync, writeFileSync } from 'node:fs';
import { durationOf, elapsedOf, observe, readTimings } from './liveness.mjs';
mkdirSync('runs/repo/issue-1', { recursive: true });
writeFileSync('runs/repo/issue-1/design.md', '# the delivered artifact\n');
// A finished sibling gives readTimings one 3-minute design sample.
mkdirSync('runs/repo/issue-0', { recursive: true });
writeFileSync('runs/repo/issue-0/run.json', JSON.stringify({
schema: 2,
stages: [{ id: 'design', at: { briefed: '2026-08-13T10:00:00Z', delivered: '2026-08-13T10:03:00Z' } }],
}));
const tenMinAgo = new Date(Date.now() - 10 * 60 * 1000).toISOString();
const run = {
schema: 2,
stages: [
{ id: 'design', at: { briefed: tenMinAgo } }, // artifact on disk
{ id: 'implement', at: { briefed: tenMinAgo } }, // nothing yet
],
};
const observed = observe(run, (s) => (s.id === 'design' ? 'runs/repo/issue-1/design.md' : null));
const now = new Date().toISOString();
for (const stage of observed.stages) {
console.log(stage.id, durationOf(stage) ?? elapsedOf(stage, now), Boolean(stage.artifactModifiedAt));
}
console.log(JSON.stringify(readTimings('runs/repo/issue-1')));
An illustrative run immediately after creating the file has this shape; the elapsed value changes with execution time:
design 10m00s+ true
implement 10m00s+ false
[{"stage":"design","n":1,"min":"3m00s","median":"3m00s","max":"3m00s"}]
The file’s presence does not remove the trailing +. A pre-existing at.delivered remains the acceptance record; this observer never manufactures one from disk. A real acceptance protocol must reject partial, stale and wrong-revision output, as shown in the linked successor.
Gotchas
- The assertion that passed on both sides of the revert. The first test for the unreadable-run fix asserted that the run’s directory name appeared in the command’s output. It did, before and after the fix, because a different column already contained the path it was derived from. The test proved nothing; reverting the fix kept it green. The escape is to parse the rendered row and assert on the specific cell the fix changed. After that change, reverting the fix alone turns the test red with the old placeholder string as the actual value. If a test guards a fix, run it against the reverted code once and watch it fail.
- An mtime is not a delivery receipt. Do not copy it into
delivered, even in memory. Keep the observation read-only and make the acceptance path validate complete output bound to the intended request before recording acceptance. A quiet file or a changing file cannot independently prove whether the worker is alive. - One corrupt sibling can take down every dispatch. The history scan walks directories this code does not own, and the first truncated
run.jsonor older-schema run it hit would have thrown mid-brief. Every parse in the scan swallows its failure and skips the sibling, and a test pins the case where a legacy-schema run sits beside a good one. The floor test matters just as much: a scan that silently matches nothing renders “no history” forever and looks correct while measuring nothing. - Byte-pinned goldens decide where you can add output. Two commands’ outputs are frozen byte-for-byte by baseline tests, so the live clock could not touch them. The board function takes
nowas an explicit option, and rendering without it is exactly the old output. Time as a parameter instead of a global is what let the feature land without invalidating the goldens, and it is also what makes every timing test deterministic.
Sources
- Response Times: The 3 Important Limits — the ten-second ceiling and progress feedback for long operations
- Health Check API pattern — the heartbeat alternative and its freshness gap
- Node.js fs documentation —
fs.Statsmtime fields and their precision - Google SRE Book: Monitoring Distributed Systems — why spreads and percentiles beat averages for skewed latencies
Changelog
- issueflow: a watched run is minutes of dead air — no progress signal during a dispatched stage, timings only shown after ship, no position at dispatch, and runs renders an old run as ‘(unreadable run)’ — cmdRuns binds its caught RunError and reports the reason and remedy per… (0b49cf0)
- issueflow: a watched run is minutes of dead air — no progress signal during a dispatched stage, timings only shown after ship, no position at dispatch, and runs renders an old run as ‘(unreadable run)’ — observe() + elapsedOf() in run.mjs, board(run, { now }) with the +… (47d9f7b)
- issueflow: a watched run is minutes of dead air — no progress signal during a dispatched stage, timings only shown after ship, no position at dispatch, and runs renders an old run as ‘(unreadable run)’ — new timings.mjs reading the run’s sibling directories, positionLine() in… (7c0a9c3)
- issueflow: a watched run is minutes of dead air — no progress signal during a dispatched stage, timings only shown after ship, no position at dispatch, and runs renders an old run as ‘(unreadable run)’ — progressPath(), the ## While you work section in renderBrief, the… (fa51818)