The durability guarantee
This is the precise statement of the core guarantee (pillar 1 in POSITIONING.md). README and marketing assert it; this document specifies it, including exactly where it stops.
What the guarantee is
Every step that completed is durably recorded before the run depends on it, and on restart the run resumes from that record instead of redoing completed work. A non-idempotent side effect is never executed a second time.
The unit of protection is the journaled step. The moment a step's outcome is written to the store, it is safe: resume replays it from the journal rather than re-running it.
The three cases when a crash hits
- Crash before a step ran → nothing recorded → resume runs it fresh. Correct.
- Crash after a step ran and its result was journaled → resume reads the result, skips re-execution. Correct.
- Crash in the gap (the side effect fired but its result was not journaled yet) → this is the dangerous window every other system re-runs into (the double-charge). Bide wrote an attempt marker before firing, so on resume it sees "this non-idempotent thing was attempted, outcome unknown" and halts (
ResumeHalt) instead of guessing. It does not silently re-run, and it does not silently assume success.
Case 3 is the whole moat. The precise phrasing is at-most-once: the side effect fires zero or one times, never twice. It is not "exactly-once": an unresumable crash in that window can leave it having fired once but unconfirmed, and the system stops for a human/policy decision rather than pretending it knows.
The boundary conditions (where the claim stops)
- The store must survive the crash. Durability is inherited from the journal's backend. If you use the in-memory store and the process dies, there is nothing to resume from: that is a dev/test store, not a durability claim. SQLite / Postgres / etc. is where the guarantee actually lives, and only as far as that storage's own durability (fsync, replication) holds.
- The write to the store must itself be atomic/durable. The guarantee reduces to "the journal did or did not record this step"; it relies on the store committing atomically. It does not defend against the storage layer lying about a commit.
- The tool must declare its safety accurately.
ReadOnlyre-runs freely,Idempotentretries, and only an unmarked non-idempotent write gets the attempt-marker/halt treatment. Mislabel a card-charge as idempotent and you have opted out of the protection. - It is at-most-once for the side effect, not "the agent always finishes." A crash can still leave a run halted and needing intervention. The promise is safety (no double-fire, no lost completed work), not liveness (guaranteed completion without help).
The one-sentence version
Not "it can't crash," and not even "it always recovers automatically." It is: when it crashes, you never lose completed work and you never double-execute a side effect, and in the one genuinely ambiguous window it stops and tells you instead of guessing.