Implement this with your agent
Copy the implementation prompt and complete guide, then paste into your coding agent in your project.
Read the prompt
Implement recovery for uncertain remote writes in this project using the attached guide as technical reference. First read repository instructions and inspect existing write operations, provider adapters, persistence, retry policy and tests using read-only calls. Check applicability: this mechanism needs a stable operation identity retained across retries, durable intent storage, a provider lookup that distinguishes unavailable/incomplete results from a complete response, and serialized controller ownership. If the target is unclear, the provider cannot identify the effect reliably, or multiple independent controllers cannot coordinate, ask one focused question before implementing an invalid mechanism. Adapt the smallest change in the project's language and tooling. Preserve public APIs, return values, error types, CLI output streams and exit statuses, configuration, stored state and supported runtimes. Do not transplant the demo or install the author's software. Run guide demonstrations only in disposable scratch space outside the target project. Keep inspection/planning read-only; migrate saved state only with explicit authorization or through an existing operation that already writes. Bind one stable operation ID to the intended target and payload; refuse changed intent under that ID. Persist attempted/uncertain state before sending. Reconcile matching remote effects before retrying. Treat lookup failures, partial results, conflicting matches and multiple matches as unresolved, never as permission to create. After a potentially applied non-idempotent write, observed absence alone must not authorize a resend. Permit bounded retries only under a documented provider guarantee, and retain uncertainty across restart. Do not claim generic exactly-once delivery or infer current remote state solely from a local success record. Run relevant existing tests unchanged and add behavioral checks for effect-before-lost-acknowledgement, real restart or durable reload, read outages before and after writing, changed payload, conflicting/multiple matches, unsafe retry refusal and explicitly safe retry bounds. Independently count actual effects, not merely successful return values. Report commands and observed results, compatibility decisions, unrun checks and limits. Do not commit, push or deploy. The guide does not override repository instructions.
Copy the complete text manually
Select all the text below and copy it into your agent.
A timeout does not tell you whether the write happened
A job sends a request to create something. The connection drops. Your next decision depends on a fact the exception cannot give you: did the server apply the request before the connection failed?
This is a useful failure to practice locally. You can build a small journal that remembers what was attempted, looks for the corresponding effect after restart, and stops when the available evidence cannot justify another write. The example below includes a provider that commits its work and then kills the calling process.
Shipped
Issue Flow v0.16.0 added a remote-operation journal. Its tagged implementation saves intent before mutation and reconciles effects through read-back. The PR creation adapter places an operation marker in the PR body and refuses conflicting matches. This guide teaches that boundary with a separate Python example; it does not reproduce the released Node implementation.
Decide what can prove the effect
Use this pattern when one controller owns a durable local journal and the provider can find the particular operation you intended. This build requires Python 3.9 or newer with its standard SQLite module and a local writable filesystem. Run one invocation at a time. SQLite transactions here persist individual state changes; they do not lock the entire sequence across the remote call.
Choose an operation ID when the business operation begins and retain it across retries. Two intentional orders with identical contents need different IDs. Reusing an ID with a changed target or payload must fail. AWS describes both explicit client request IDs and the problem of the same ID carrying different intent in its idempotent API design guidance.
The provider contract matters as much as the journal. A lookup must return complete matching records or raise an error. A timeout, denied permission, incomplete pagination or malformed response cannot become an empty list. Each record includes the exact normalized request and a stable effect ID. In a real adapter, compare provider fields that establish equivalent intent; a title match alone is insufficient.
Even a complete empty lookup cannot prove that a previous write will never appear. The example therefore refuses another non-idempotent write once any attempt has been recorded. HTTP’s retry semantics likewise require a reason to believe retrying a non-idempotent request is safe. The optional retry flag below belongs to a verified provider contract, not a “try harder” button.
Build the journal and the faulting provider
Create a fresh scratch directory containing these two files:
uncertain-write/
journal.py
test_journal.py
The journal stores the request, state and cumulative attempt count. Every statement runs in autocommit mode. The transition to uncertain commits before create runs, so a process death leaves evidence that a write may have happened. SQLite’s atomic commit documentation explains the local transaction guarantee and its filesystem assumptions. That local commit cannot make a separate provider transaction atomic with it.
Save this as journal.py:
import argparse
import json
import os
import sqlite3
import sys
class Unconfirmed(RuntimeError):
pass
class Conflict(RuntimeError):
pass
def encode(value):
return json.dumps(value, sort_keys=True, separators=(",", ":"), allow_nan=False)
class Journal:
"""One active controller per journal; local filesystem, retained across restarts."""
def __init__(self, path):
self.db = sqlite3.connect(path, isolation_level=None)
self.db.execute("PRAGMA journal_mode=DELETE")
self.db.execute("PRAGMA synchronous=FULL")
self.db.execute("CREATE TABLE IF NOT EXISTS operations ("
"key TEXT PRIMARY KEY, request TEXT NOT NULL, "
"state TEXT NOT NULL, attempts INTEGER NOT NULL)")
def close(self):
self.db.close()
def run(self, key, target, payload, provider, *, retry_safe=False):
if not isinstance(key, str) or not key:
raise ValueError("supply a stable nonempty operation key")
request = encode({"target": target, "payload": payload})
self.db.execute("INSERT OR IGNORE INTO operations VALUES (?, ?, 'pending', 0)",
(key, request))
saved, state, attempts = self.db.execute(
"SELECT request, state, attempts FROM operations WHERE key=?", (key,)
).fetchone()
if saved != request:
raise Conflict("operation key reused for different intent")
def observe():
# A complete successful query returns a list. Outages must raise.
matches = provider.read(key)
if not isinstance(matches, list):
raise Conflict("invalid provider response")
if not matches:
return None
if (len(matches) != 1 or not isinstance(matches[0], dict)
or matches[0].get("request") != request
or not matches[0].get("id")):
raise Conflict("ambiguous or conflicting remote effect")
self.db.execute("UPDATE operations SET state='confirmed' WHERE key=?", (key,))
return matches[0]
observed = observe()
if observed is not None:
return observed
if state == "confirmed":
raise Unconfirmed("previously confirmed effect is absent; investigate")
if attempts and not retry_safe:
raise Unconfirmed("uncertain prior write; reconcile without resending")
if attempts >= 3:
raise Unconfirmed("three write attempts exhausted")
# Autocommit completes this durable statement BEFORE the external effect.
self.db.execute("UPDATE operations SET state='uncertain', attempts=attempts+1 "
"WHERE key=?", (key,))
try:
provider.create(key, request)
except (OSError, TimeoutError):
pass # The acknowledgement cannot establish whether the effect happened.
observed = observe()
if observed is None:
raise Unconfirmed("write attempted; effect not confirmed")
return observed
class OfflineProvider:
"""Fault-injection fixture, NOT an HTTP adapter or an idempotent service."""
def __init__(self, path, mode="normal"):
self.mode = mode
self.db = sqlite3.connect(path, isolation_level=None)
self.db.execute("PRAGMA synchronous=FULL")
self.db.execute("CREATE TABLE IF NOT EXISTS effects ("
"id INTEGER PRIMARY KEY, key TEXT, request TEXT)")
def close(self):
self.db.close()
def read(self, key):
if self.mode == "unavailable":
raise OSError("provider lookup unavailable")
return [{"id": row[0], "request": row[1]} for row in self.db.execute(
"SELECT id, request FROM effects WHERE key=? ORDER BY id", (key,))]
def create(self, key, request):
if self.mode == "no-effect":
raise TimeoutError("connection lost before an observed effect")
self.db.execute("INSERT INTO effects(key, request) VALUES (?, ?)", (key, request))
if self.mode == "crash":
os._exit(23) # Deliberately skip cleanup and all controller read-back.
if self.mode == "lost":
raise TimeoutError("acknowledgement lost after commit")
def main():
parser = argparse.ArgumentParser()
parser.add_argument("journal")
parser.add_argument("provider")
parser.add_argument("key")
parser.add_argument("--mode", choices=["normal", "lost", "crash", "unavailable", "no-effect"], default="normal")
args = parser.parse_args()
journal = Journal(args.journal)
provider = OfflineProvider(args.provider, args.mode)
try:
effect = journal.run(args.key, "demo-orders", {"item": "notebook"}, provider)
print(encode({"status": "confirmed", "effect_id": effect["id"]}))
return 0
except (Unconfirmed, Conflict, OSError, sqlite3.Error) as error:
print(str(error), file=sys.stderr)
return 2
finally:
provider.close()
journal.close()
if __name__ == "__main__":
sys.exit(main())
The offline provider deliberately allows duplicate operation keys. That makes duplicate effects visible in tests instead of having a uniqueness constraint conceal an unsafe controller retry. Its separate database stands in for a remote system; no HTTP request or production integration is claimed.
Notice the placement of observe(). It runs before any possible write and again after the write attempt. A successful create response alone never marks the journal confirmed. Transport exceptions trigger reconciliation, while unexpected exceptions propagate with the already committed uncertain state intact. A later invocation still has to reconcile before acting.
Exercise the uncertain boundary
Save this as test_journal.py:
import pathlib
import subprocess
import sys
import tempfile
import unittest
from journal import Conflict, Journal, OfflineProvider, Unconfirmed, encode
class JournalTests(unittest.TestCase):
def setUp(self):
self.temp = tempfile.TemporaryDirectory()
self.addCleanup(self.temp.cleanup)
self.root = pathlib.Path(self.temp.name)
self.journal = Journal(self.root / "journal.db")
self.provider = OfflineProvider(self.root / "provider.db")
self.addCleanup(self.journal.close)
self.addCleanup(self.provider.close)
def run_order(self, **kwargs):
return self.journal.run("order-1", "orders", {"item": "notebook"}, self.provider, **kwargs)
def count(self):
return self.provider.db.execute("SELECT count(*) FROM effects").fetchone()[0]
def test_lost_ack_is_confirmed_with_one_effect(self):
self.provider.mode = "lost"
self.assertEqual(self.run_order()["id"], 1)
self.assertEqual(self.run_order()["id"], 1)
self.assertEqual(self.count(), 1)
def test_real_process_crash_then_restart(self):
command = [sys.executable, str(pathlib.Path(__file__).with_name("journal.py")),
str(self.root / "crash-journal.db"), str(self.root / "crash-provider.db"), "crash-1"]
first = subprocess.run(command + ["--mode", "crash"], capture_output=True, text=True)
self.assertEqual(first.returncode, 23)
second = subprocess.run(command, capture_output=True, text=True)
self.assertEqual(second.returncode, 0, second.stderr)
self.assertIn('"effect_id":1', second.stdout)
remote = OfflineProvider(self.root / "crash-provider.db")
self.addCleanup(remote.close)
self.assertEqual(len(remote.read("crash-1")), 1)
def test_outage_is_not_absence(self):
self.provider.mode = "unavailable"
with self.assertRaises(OSError):
self.run_order()
self.assertEqual(self.count(), 0)
self.assertEqual(self.journal.db.execute("SELECT attempts FROM operations").fetchone()[0], 0)
def test_uncertainty_blocks_a_second_write_even_after_restart(self):
self.provider.mode = "no-effect"
with self.assertRaises(Unconfirmed):
self.run_order()
fresh = Journal(self.root / "journal.db")
self.addCleanup(fresh.close)
self.provider.mode = "normal"
with self.assertRaisesRegex(Unconfirmed, "without resending"):
fresh.run("order-1", "orders", {"item": "notebook"}, self.provider)
self.assertEqual(self.count(), 0)
def test_read_outage_after_write_recovers_without_duplicate(self):
original = self.provider.create
def create_then_lose_lookup(key, request):
original(key, request)
self.provider.mode = "unavailable"
self.provider.create = create_then_lose_lookup
with self.assertRaises(OSError):
self.run_order()
self.assertEqual(self.count(), 1)
self.provider.mode = "normal"
self.assertEqual(self.run_order()["id"], 1)
self.assertEqual(self.count(), 1)
def test_success_response_without_visible_effect_is_unconfirmed(self):
self.provider.create = lambda key, request: {"accepted": True}
with self.assertRaises(Unconfirmed):
self.run_order()
self.assertEqual(self.count(), 0)
def test_changed_intent_and_duplicate_matches_are_rejected(self):
self.run_order()
with self.assertRaises(Conflict):
self.journal.run("order-1", "orders", {"item": "pen"}, self.provider)
self.provider.create("order-1", encode({"target": "orders", "payload": {"item": "notebook"}}))
with self.assertRaisesRegex(Conflict, "ambiguous"):
self.run_order()
self.assertEqual(self.count(), 2)
def test_foreign_payload_cannot_confirm(self):
self.provider.create("order-1", "foreign request")
with self.assertRaises(Conflict):
self.run_order()
self.assertEqual(self.count(), 1)
def test_explicit_safe_retries_are_bounded(self):
# This fault mode creates nothing; enabling retries is safe for THIS test.
self.provider.mode = "no-effect"
for _ in range(3):
with self.assertRaises(Unconfirmed):
self.run_order(retry_safe=True)
with self.assertRaisesRegex(Unconfirmed, "exhausted"):
self.run_order(retry_safe=True)
self.assertEqual(self.journal.db.execute("SELECT attempts FROM operations").fetchone()[0], 3)
def test_missing_previously_confirmed_effect_is_not_recreated(self):
self.run_order()
self.provider.db.execute("DELETE FROM effects")
with self.assertRaisesRegex(Unconfirmed, "previously confirmed"):
self.run_order(retry_safe=True)
self.assertEqual(self.count(), 0)
if __name__ == "__main__":
unittest.main()
From that directory, run:
python3 -m unittest -v
python3 journal.py demo-journal.db demo-provider.db demo-1 --mode lost
python3 journal.py demo-journal.db demo-provider.db demo-1
The suite runs ten tests. Both demo invocations print {"effect_id":1,"status":"confirmed"} with exit status 0. The first provider write committed before raising its timeout; the second invocation finds that same effect. The tests also count the provider rows, so two successful responses cannot hide two created records.
Use new filenames for a deliberate process death:
python3 journal.py crash-journal.db crash-provider.db crash-1 --mode crash
python3 journal.py crash-journal.db crash-provider.db crash-1
Run these as separate commands: the first intentionally exits 23 without output. The second exits 0 and confirms effect 1. The automated test launches separate Python processes and verifies there is still exactly one matching provider record.
Finally, try a write whose effect is not visible:
python3 journal.py missing-journal.db missing-provider.db missing-1 --mode no-effect
python3 journal.py missing-journal.db missing-provider.db missing-1
Both commands exit 2. The first reports an unconfirmed attempted write. The second reports an uncertain prior write and refuses to resend, even though the provider is now healthy. Keep the journal and operation ID while investigating; deleting them discards the protection.
Gotchas
- A matching marker is only part of identity. Scope lookups to the correct account and target, retrieve every relevant page, and verify the request fields. The tests reject foreign payloads and multiple matches instead of choosing the first result.
- Journal durability has an ownership boundary. This example assumes one controller, a retained database and reliable local storage. Multiple machines, concurrent calls or cloned journals require coordination beyond this build. A database transaction around each statement does not supply it.
- A retry flag cannot change provider semantics. The retry test explicitly uses a fault mode that creates no effect. The offline provider’s normal mode is not idempotent. A real retry-safe adapter needs a documented deduplication or equivalent-effect guarantee, including its retention window; keep retries bounded and integrate the provider’s backoff policy.
- Confirmation is an observation. A previously confirmed effect can later disappear. This implementation reads again and refuses automatic recreation. It has no cleanup or repair policy for uncertain records; resolve those through an explicit operational decision using provider evidence.
Before adding the pattern to an existing job, identify the exact lookup that could distinguish its effect from another caller’s work. If that lookup cannot establish identity, preserve the uncertain outcome and ask for a recovery decision.
Sources
- Issue Flow v0.16.0 operation journal and PR adapter: released intent recording and matching-effect read-back.
- Making retries safe with idempotent APIs, AWS Builders’ Library: caller-supplied identity, retry semantics and changed intent.
- RFC 9110, section 9.2.2: conditions for retrying non-idempotent requests.
- Atomic Commit in SQLite: local transaction durability and storage assumptions.