Changelog
All notable changes to bide are documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning. While bide is pre-1.0, a minor version (0.x.0) may include breaking API changes, which are marked Breaking: below; the journal format can change between pre-releases without a version bump. Curated highlights for each release are in docs/releases.
Unreleased
Fixed
- Documentation: the formal verification overview said the TLA+ models caught 29 bugs, all before release. They found 28 and confirmed one more found in review, and four (L1, S1, S2, S4) shipped in releases up to v0.9.0 and were fixed in v0.10.0. The overview, the README and the docs index now say so.
0.11.1 - 2026-10-04
Fixed
- A failed model call's error said
(model) (model): the run wrappedErrModelaround an adapter error that already wrapped it. It wraps it only when it is missing (#160). model/openaiwith an empty API key: a server that needs a key refuses the request with 401 or 403, which read like a bad key or a provider fault. The*agent.APIErrornow says no API key was set (openai.Newwas given an empty key); a server that needs none, such as Ollama, works as before (#160).
Changed
- Getting started works without an API key: an
agenttestscripted model runs the first agent offline. The install step gets theagentpackage (go get github.com/bide-ai/bide/agent@latest), since getting the module root leftgo.sumincomplete; the core examples run withgo run github.com/bide-ai/bide/examples/<name>@latest; and the guide says what a run prints and how a crashed run continues (#161). - The examples name a
*agent.Journaljournal, notstore, soagent.New(model, journal, ...)does not read like the removedNew(model, store, tools...)(#159).
0.11.0 - 2026-10-04
Changed
Breaking: the API redesign's new names (P15, docs/design/api-v1.md section 10.2, #156). The transitional names and the old API they stood beside are gone. The list below is the upgrade path, name by name; a module that used
github.com/bide-ai/bide/mcprequiresgithub.com/bide-ai/bide/mcptoolsinstead.- Construction:
BuildisNew(model, journal, opts...) (*Agent, error); the oldNew(model, store, tools...)and the builder methods (Use,UseTool,WithMaxTurns,SetMaxConcurrency, ...) are removed: passWithTools,WithMiddleware,WithToolMiddlewareand the other options, or derive a copy withWith.Newmakes every checkBuildmade (two tools with one name, a non-object schema, an m-of-n policy's keys), so a configuration problem fails at construction rather than at a run. - Runs:
RunMessage,StreamMessage,ResumeRun,RunTypedMessageandSession.SendMessage/SendMessageOnceareRun,Stream,Resume,RunTypedandSession.Send/SendOnce, each taking aMessage(agent.UserText(s)) and run options and returning a*Result. The string entry points they replaced (RunandStreamon a string,RunSaga,RunResult,RunSagaResult,StreamSaga, theRunTypedandRunTypedNativefunctions,AgentStream.Final) are removed: a saga isRun(..., WithSaga()), native structured output isWithOutputMode(OutputNative).AgentStreamandAgentEventareRunStreamandRunEvent;audit.Recordisaudit.RecordStream, which returns(*agent.Result, error). - Journal: the
Durableinterface and the stores'Do,HistoryandJournalshims are removed. Every function that took aDurabletakes a*Journal(agent.NewJournal(store)); a test fake that wrapped aDurablewraps theStoreinstead (and implementsUnwrap() Store, whichCapabilityand the run's identity follow). The journal verbs are generic methods:j.Step,j.Parallel,j.Signal,j.Enqueue(the channelSend) andj.AnswerInterrupt(theResume[T]wrapper). - Pauses and halts: the aliases
PendingApproval,Interrupted,Awaiting,SleepingandResumeHaltareApprovalPending,InterruptPending,SignalPending,TimerPendingandOutcomeUnknown;ResolveHaltRefisResolveHalt(ctx, j, HaltRef, Outcome), and the wrappersResolveStepHaltandApproveAsare removed (SubmitDecisionrecords one approver's decision). - Context decorators:
ContextWithIdentity,ContextWithWakerandContextWithClockare removed; passWithIdentity,WithWakerorWithClockas a run option.WithIdentityrefuses an empty identity (ErrConfig), whichContextWithIdentitybound. - Tools: the
Toolinterface isSpec() ToolSpecandCall; the old method set (Name,Description,ArgsSchema,Safety) andSpecOfare removed.Functakes its safety as an option (WithSafety) and, likeCompensatedFunc,SubAgentandRetrievalTool, returns(Tool, error);MustFunc,MustCompensatedFunc,MustSubAgentandMustRetrievalToolpanic instead.SubAgent(nil)isErrConfig. A decorator that embeds aToolinherits itsSpec, and with it the gate and the timeout; it overrides in its ownSpec(s := d.Tool.Spec(), then the fields it changes).New, and a plan flow registering a tool, refuse (ErrConfig) a tool whoseName,Description,ArgsSchema(compared as JSON) orSafetymethod disagrees with itsSpec, including a pointer-receiver method on a tool registered by value, one of another signature, and one two embedded fields declare; a method of one of those names that means something else is refused too (rename it): written for the old interface, such a method no longer overrides anything (a decorator'sSafety()declaring a side effect would be ignored, and a resume would run it again). - Packages:
mcpismcptools(github.com/bide-ai/bide/mcptools); the scripted model and its turns (ScriptedModel,NewScriptedModel,ToolTurn,TextTurn,ErrorTurn) moved toagent/agenttest, besideMemJournal,MustJournal,MustNew,Must,AnswerandIdentityContext; plan'sRegister*functions are methods of*plan.Registry.
- Construction:
Breaking:
RunwithoutWithSagaon a run journaled as a saga drives it as the saga it is (its rollback included), whereRunSaga's absence wasErrConfig;WithSagaon a run that is not a saga is stillErrConfig. A session turn and a sub-agent still refuse a drive that differs.Breaking:
audit.AuditedStorewraps aStore(not aDurable): it anchors after everyInsertbut the journal header's (also one that found its entry stored, or failed: an A3 write that committed while reporting an error is anchored), the header together with the run's first record, and never a head of the header alone; a publish that failed is covered by the run's next write (a replay no longer retries it), andAuditedStore.Reanchor(ctx, runID)anchors a run whose last write's publish failed (call it onOnError).cmd/bide-auditrefuses an export whose first record is not a journal header.
Fixed
- An
audit.AttenuatingSubAgentdelegation whose sub-run held only its journal header (its first record failed to write) was taken for a sub-run with records but no journaled authority, and every later call refused it; a sub-run holding only its header is now one that has not started.
Added
audit.AuditedStore.Reanchor(ctx, runID): anchors a run's journal now if no published head covers it, the recovery for a run whose last write's publish failed, or whose write landed under a context that was done, whichInsertcannot anchor and reports toOnErrorwith the context's error (OnErroris the signal).- CI checks that no removed name appears in Go code, godoc or docs (
internal/tools/oldnames).
0.10.0 - 2026-10-03
Added
Run API, cancellation and recovery dispatch (P14)
- The Run API, under transitional names (P14):
Agent.RunMessage(ctx, runID, input Message, opts...),Agent.ResumeRun(ctx, runID, opts...),Agent.StreamMessagewithAgentStream.Result(),Agent.RunTypedMessage[T](a Go 1.27 generic method),Session.SendMessageandSession.SendMessageOnce, which the 1.0 rewrite renamesRun,Resume,Stream,RunTyped,SendandSendOnce. Each returns aResultwhenever the run ID is valid (agent.ValidateRunID), whatever the error;Result.Outputholds a typed run's answer as journaled. The input is aMessage, so a run can start from text and images (#138). - Run options (P14): the agent-and-run options (
WithMaxTurns,WithTokenBudget,WithSystemPrompt,WithSampling,WithToolChoice,WithWaker,WithIdentity,WithMaxConcurrency,WithClock) now apply per run, andWithSaga(),WithToolFilter(names...)andWithOutputMode(OutputTool | OutputNative)are new. The journaling rule: a run's first drive journals its input and options inrun:start(RunStartgainsSession,Typed,Settings,Principal,ToolsandExt, and the kindsession_turn), and every later drive, a recovery drive included, runs under them. A later drive's different turn limit or token budget is journaled as an amendmentrun:limits:<n>; any other different setting (input, saga, tool filter, system prompt, sampling, tool choice, output mode, typed schema, principal) isErrConfig, before any model call. The tool filter is enforced at dispatch: a call outside it records an error result and never runs (model 10, finding L5). The principal (OnBehalfOf,AuthorityRef) is journaled and restored; theActorstays live (#138). agent.Cancel(ctx, j, runID, reason)andagent.Status(ctx, j, runID)(RunStatus,RunState) (P14, D1 and D8).Cancelwritesrun:cancelled, or on a saga the rollback requestrun:cancel-requested, which the saga's drive answers by rolling the run back and writingrun:cancelled. A drive checks for the cancellation when it starts (and again once it has writtenrun:startor a limit amendment, or lost therun:startinsert to another drive: model 10'sDStartandDAmendreturn toDOpen), at every turn boundary after its first, and after each won side-effect claim, before the call (model 10, L2); calls in flight finish, and a retry-safe call already dispatched in the current turn may still run (noGetper retry-safe call). A sub-run (a sub-agent's, aSubRunForrun, a delegation) reads its tree root's cancellation too, at each of those checks and from the root's own store (a sub-agent outside a saga may journal to another), so aCancelof the root stops the whole tree (model 11,NoFireAfterRootCancel). Aplanflow honoursCancel: its open reads its end markers, a node readsrun:cancelledonce it will run (after a side-effect node's claim is won), and itsrun:completeis read back. The first end marker in journal order is a run's end for every reader, and each writer of one reads the markers back (L3).Cancelof a run already cancelled returns nil, of a run otherwise overErrRunEnded.Statusreads oneLoad(L6). New errors:ErrRunCancelled,ErrRunEnded,ErrNotStartedandErrNotResumable, none in a category (ErrNotStartedis often a race with a run's first drive) (#138).- Recovery dispatch (P14):
agent.Resumer(func(ctx, runID, start RunStart) error),agent.ResumeAgent(a, opts...),agent.ResumeTyped[T](a, opts...)andagent.ResumeAny(rs...). A recovery pass reads each run'srun:startunder its lease; a run with none is skipped and reported once per process (ErrNotStarted), and read again on every pass (L7); a run noResumerdrives is reported once (ErrNotResumable); the process remembers its 65,536 most recent reports. A run a pass drove to its cancelled end (a saga's cancellation rollback included) is recovered, not a failure.ResumeAgenttakes only the deployment's options (a Waker, a clock, a concurrency cap, an identity'sActor) (#138).
Sessions
- Sessions (P14, model 12's S3): a Send turn whose run was cancelled is recorded closed, with no answer and outside the transcript, by the next message's
Send;Sendof the cancelled message (the same text) returnsErrRunCancelled. A saga turn'sCancelwrites only its rollback request: the next message'sSenddrives the turn's rollback itself, under the turn lease and without holding the session handle's mutex (other callers on the handle are not blocked behind its compensators: while it is in progress, another worker or a second caller on the same handle getsErrTurnContended), and then reads the journal again and records the turn closed. Over a store with noLeaser, the handle's mutex stays held across the rollback, so two callers on one handle never drive it at once (#138). agent.ErrTurnContended: aSession.SendorSendOncewhose turn run another holder leases returns at once with an error wrapping it, having driven nothing (no category, likeErrLeaseLost); send the message again later.Agent.Session(ctx, id, opts...)takesLeaseOptionvalues (WithLeaseHolder,WithLeaseTTL) for that lease (#137).
Tools (P12)
agent.ToolSpec(Name,Title,Description,Input,Output,Safety,Approval,Timeout), everything the agent knows about a tool, andagent.SpecOf(t), which reads a tool'sSpec() ToolSpecmethod or, for a tool without one, itsName,Description,ArgsSchemaandSafety(deprecated from the start: see Deprecated). The agent reads each tool's spec once, when it is registered, and decides every call from that copy (#117).- Tool options:
agent.ToolOptionwithWithSafety,WithApproval,WithTimeout,WithTitleandWithOutputSchema, andagent.SingleApproval(), the one-decision gate.FuncandCompensatedFunctake trailing options;SubAgent(name, description, sub, opts...)takesWithApproval, so a parent can require approval before it delegates, and refusesWithSafetyandWithTimeout. An invalid option panics with an error wrappingErrConfig, asFuncdoes for an argument type it cannot describe, until the 1.0 rewrite returns errors (#117). - Tool timeouts (
WithTimeout,ToolSpec.Timeout): the call, tool middleware included, runs under the deadline. A result returned after the deadline is recorded; an error returned after it has an unknown outcome: a side effect records nothing, fails the drive withErrToolOutcomeUnknownand halts on resume, and a retry-safe tool records the error. It bounds only tools that honor their context (#117). agent.ToolCall(Use,Spec,RunID) andToolCall.ErrorText(err), what tool middleware receives (#117).agent.ErrToolNotCalled, a condition (no category) for a tool call known never to have reached its tool. A tool middleware that ends a call without callingnextreturns an error wrapping it, and only then;middleware.ToolRateLimitandmiddleware.ToolRetrydo. The agent tracks each call's state by compare-and-swap (reached, refused by the base handler, or closed when the chain returned without entering it): a call counts as not called only when refused, or closed withErrToolNotCalled; any other closed call leaves a side effect's outcome unknown and the run halts. The base handler refuses an invocation that comes after the chain returned. TheToolMiddlewarecontract: reach the tool only throughnext(#117).- Tool-result, denial and saga-failure records journal the
Safetyand approval gate the call ran under (Record.Safety,Record.Approval); a saga rollback reads the recorded safety (#117). - A tool that wraps another says so with an
Unwrap() Toolmethod; the agent follows it to find a wrappedSubAgent, so a saga rollback and the tree's token budget recurse into its sub-run (#117). mcp.WithApproval(name, policy):Toolsfails withErrConfigwhen the policy is nil or invalid, or when the server lists no tool of that name, so a misspelt gate never leaves the real tool ungated. Each MCP tool's spec:Title(its title, else its annotations' title),Output(itsoutputSchema) andTimeout(WithCallTimeout) (#117).govern.EventToolConfigandgovern.FederatedEventToolConfig, whoseOptionspassagent.ToolOptions to the tool;audit.AttenuationConfig, and trailingagent.ToolOptions onaudit.AttenuatingSubAgent, which go to itsSubAgent(#117).
Construction under Build (P13)
agent.Build(model, journal, opts...) (*Agent, error), construction from options, andAgent.With(opts...), which returns a configured copy and leaves the agent unchanged.BuildandWithreturn every configuration problem asErrConfigwhen the agent is built: a nil model, journal, option, tool, middleware or function; two tools with one name or one namedfinal_answer; a tool name the agent's model refuses, when the model declares its rule (ToolRules); a non-object input schema; an invalid approval policy, an m-of-n policy with noWithApproverVerifiersor with two approvers on one signing key; a negative limit; aWithRetrievalk below 1; and a tool choice with an unknown mode or a tool the agent lacks.Buildis the transitional name of the 1.0New.Agent.Journal()returns the agent's journal (#127).- Agent options:
WithTools,WithMiddleware,WithToolMiddleware,WithApproverVerifiers,WithToolErrorRedactor,WithSystemPromptFunc(its function gets the run'sRunInfoand may fail),WithRetrieval,WithOptions, andWithMaxTurns,WithTokenBudget,WithSystemPrompt,WithSampling,WithToolChoice,WithWaker,WithIdentity,WithMaxConcurrencyandWithClock. The last value given for a setting wins;WithSystemPromptandWithSystemPromptFuncshare one slot. The agent'sWithIdentity,WithWakerandWithClockapply to a run whose context carries none (#127). - Option scopes are interfaces with unexported methods (
Option,RunOption,ParallelOption,StepOption,ResolveOption,LeaseOption,RecoverOption,RecoverLoopOption,RetrievalOption, andToolOption), and a setting for several scopes returns one of the combination typesAgentRunOption,ConcurrencyOption,ClockOption,SafetyOptionorLeaseControl, so an option passed where it does not apply does not compile.RunOptionvalues are type-checked now and taken by the run API in the next redesign step (#127). agent.RunInfo(RunID,RootRunID,ToolUseID,Saga) andRunInfoFrom(ctx), which describe the tool call a context belongs to, andRunInfo.SubRunFor(name), the run ID of a programmatic sub-run (<scope>>step:<name>), whichRunaccepts from that call's context (#127).agent.WithRetrievalRetry(n, base, max), aRetrievalOptionforWithRetrieval(r, k, opts...): a failedRetrieveis retried up to n more times within the retrieval step, with exponential backoff and full jitter, and only the attempt that succeeded is recorded.agent.RetrieverFuncadapts a function to aRetriever, which is how a policy wraps one (#127).agent.WithSubRuns(agentFor), a tool option that declares the agent each of the tool's programmatic sub-runs (RunInfo.SubRunFor(name)) runs with. A saga's rollback walks every programmatic sub-run a call started, latest first and after the call's own compensation, as it walks a sub-agent's: a run in a saga's tree (a saga, or a plain run started from one's call) links each one in its journal (@subrun/<call>/<name>) before the sub-run records anything, and the rollback compensates the sub-run's writes with the declared agent's compensators, or, with none declared, reports them inSagaAborted.Uncompensated. A declared agent it cannot use (on another store, or aWithSubRunsfunction that panics) is listed there with the reason. A sub-run in a saga's tree must journal to the saga's store:Runrefuses another withErrConfig(#127).agent.ToolRules, an optional interface aModelimplements (itself or throughUnwrap() Model) to declare the tool setups its provider refuses:ToolNameRule() *regexp.RegexpandRequiresToolsForRequired() bool.BuildandWithrefuse a tool name the declared rule does not match, and check nothing for a model that declares no rule.model/openaiandmodel/anthropicdeclare^[a-zA-Z0-9_-]{1,64}$, andmodel/geminiits own^[a-zA-Z_][a-zA-Z0-9_.:-]{0,63}$, so a dotted MCP tool name still builds for Gemini. Tool choicerequiredon an agent with no tools of its own is left to the run, sinceRunTypedsupplies an answer tool: a run with nothing to call, under a model that declares it needs one, fails withErrConfigbefore it opens the journal or calls the model (#127).
Recovery and leases
agent.RunFilter.LeaseLapsedadmits only the runs whose lease has lapsed, by the comparisonAcquireLeasemakes;MemStore, SQLite and Postgres evaluate it over their leases table, andstoretestchecks it (Lister_LeaseLapsed) (#126).agent.WithRecoverLapsedConcurrencycaps how many lapsed runsRecoverLoop's lapsed loop drives at once (16 by default), apart fromWithRecoverConcurrency(#126).Leaser.ReapLeases(ctx, ended)deletes the lapsed leases no recovery pass takes over (a finished run's, or one on a run the store does not hold), checking the expiry in the same statement;RecoverLoop's lapsed loop calls it on each pass, so a holder that died between its run's last write and its release no longer leaves a lease every later pass reads.MemStore, SQLite and Postgres implement it, andstoretestchecks it (Leaser_ReapLeases) (#126).- Postgres:
Opencreates two indexes on the leases table when they are missing (<prefix>leases_expiryand<prefix>leases_run_c, onrun_idunder the "C" collation), so the lapsed listing reads only lapsed leases, or pages through many in order without sorting; on a store whose role does not own the tables, the owner opens it once to create them (#126).
Delegation
audit.WithRollbackGrants(ctx, signer, grants...)binds further grants a saga rollback verifies its delegations against, beside the acting grant (WithGrant): a saga whose delegations were minted under several grants (the root grant expired or was rotated between drives) rolls back with all of them bound. Each journaled child grant is verified against its own parent among the bound grants, under that grant's signer; a grant bound this way is never minted from (model 11, finding D1) (#133).
Governance (gsm)
- The gsm machine gate, a required CI check: every gsm machine the
examples/governprograms make is checked by the two checkers extracted from gsm's Coq/Rocq proof (checkeron a machine's step tables,astcheckeron its rules), built at a pinned gsm commit. gsm'sgsmgatebuilds each program.github/gsm-gate/machines.txtlists and runs it twice with no arguments in a fixed environment; the job fails when a checker rejects a machine listed as accepted or verifies oneBuildrejects, a program makes a machine the list does not hold or misses one it holds, a run fails or the two runs make different machines, the scan finds a program or package that makes machines outside the list, or the pinned commit is not on gsm'smain. It checks the examples' machines in CI, not the machines an application builds (.github/gsm-gate, #144).
Formal models
- A TLA+ model of the tool-call state machine of P12 (model 9,
spec/tla/toolcall): the call state and began word, the base handler entered by several invocations (retries, anextleft running), the tool middleware, the loop's record decision, the errgroup, retry-safe steps that change state, and the saga rollback's re-run and compensation, under store faults, crashes, cancellations and deadlines. It found six bugs in #117 before it merged (T1 to T6: a known failure or a result recorded for a call whose tool began, did not itself fail, or was still running; "already ran" for a tool that never began; an earlier attempt of a retry-safe saga step listed nowhere; a retry-safe tool begun after its chain returned), and a rollback re-run with no end; each is a regression configuration (#123). - A TLA+ model of the run lifecycle and recovery (model 10:
run:complete,run:abortedand the reservedrun:cancelled; leased and plain runs,RecoverandRecoverLooppasses with the end-marker re-check under the lease, halts, pauses andResolveHalt, under lease expiry, stalled holders, ambiguous writes and crashes), checked in CI. It shows the pickup latency the v0.9.0 docs describe (a dead holder's run waits behind the halted runs listed before it), checks a proposed fix, and states the property P14'sCancelmust satisfy (#124). Extended for P14 before any of P14's code was written:Cancelon plain runs and sagas,Status, the per-run tool filter, per-run options that survive recovery, and not-started runs in recovery, with findings L4 to L7, whose rules P14 adopted (#129). - A TLA+ model of delegation, sub-run authority and saga trees (model 11,
spec/tla/delegation):audit.AttenuatingSubAgentgrants (minting from the bound grant, reuse on resume, expiry at mint, at reuse and before every call, the subject check, a resume under the wrong authority), programmatic sub-runs (RunInfo.SubRunFor: the saga-tree link written before the child's first write, one store per saga tree, a sub-run started after its call returned refused), the rollback's recursion into sub-agents and linked sub-runs,BindRollback, and halts propagating from sub-runs, under ambiguous writes, crashes, transient read errors, lost outcomes and the clock. It found D1 to D3, fixed in #133, each now a regression configuration (#130). - A TLA+ model of sessions (model 12,
spec/tla/sessions):SendandSendOnceturns as runs of their own seeded with the shared history, several callers per handle and handles per process, redeliveries, error replies on every session-journal write, crashes, pauses, the per-turn budget, and P14'sCancelof a turn's run. It found S1 to S4: S1, S2 and S4 are fixed in #137 and S3 by rule 16 of P14's contract (#138); each is a regression configuration (#135). - P14 implements model 10's adopted rules L2 to L7 and model 12's S3 (rule 16 of the P14 contract):
findings/cancel-saga-marker,filter-request-only,status-gets,not-started-remembered(model 10) ands3-cancel-wedge(model 12) are regression configurations now, each still failing under its old rule; the drive,Cancel,Statusand session code carries theprotocol:lifecycleandprotocol:sessionsmarkers of the actions that were on the no-code lists (#138). - Code and models are kept in step (milestone M4 of the formal-models plan): the Go code models 1, 1b, 7, 8 and 10 describe is wrapped in
// protocol:<model> begin <Action> .../// protocol:<model> endregion markers, and the Lint job runsinternal/tools/modelsync, which fails a pull request that touches a marked region without changingspec/tla/<model>/(unless the description or a commit message holdsProtocol-Impact: none (<reason>), printed as a warning), and any disagreement between the markers, the model-to-code maps inspec/tla/README.mdand the specs' action names.TestProtocolVocabularychecks that the claim code's key constructors and record kinds match the record kindsClaims.tladeclares, and runs in the same job (#125). - Apalache (v0.62.2, pinned with its SHA-256 in
spec/tla/tools.lock) checks the claim model nightly, beside TLC:spec/tla/check.sh apalache [model|file.cfg]downloads and verifies it, unpacks a fresh copy, and removes its work directories on exit, failure and interrupt. A typed wrapper (ClaimsApalache.tla) instantiates the model unchanged, and an inductive invariant (ClaimsInductive.tla) proves, for two drivers over two processes,AtMostOnceandNotStartedExclusiveon one call (attempts 0..3, 8 claim ids) and all four ofAtMostOnce,NotStartedExclusive,NoLiveOverrideandAtMostOncePerIntentwith halt resolution and the caller's second call (attempts 0..3, 6 claim ids), without the approval gate, and, under the lease check, assuming no plain run holds the live attempt at the check (PlainRunIdleAtCheck), at any depth and for any number and mix of faults within the run's 8 (6) claim ids and attempts 0..3 (claim ids are never reused, so this bounds the number of claims), with every placement of the drivers, kind of call and late-commit setting left to the solver; the drivers, processes, calls, attempts and claim ids stay bounded (Apalache results). A regression check must find #90's F2 (NoLiveOverride) symbolically, TLC checks the invariant as a plain invariant of every reachable state of three configurations, and Apalache type-checks model 9's wrapper. The new Apalache (nightly) job runs them; the required Models job is unchanged. Model 9 has a typed wrapper (ToolCallApalache.tla) and no Apalache check of its properties yet. To type model 1,Claims.tlachanged three expressions without changing their meaning; TLC passes every configuration as before (#148).
Documentation, testing and tooling
- Formal verification, an overview of bide's TLA+ models: why bide model-checks, what each model guarantees and which code it covers, the bugs the models caught before release (F1 to F5, P1, P2, T1 to T6, L1 to L3) and where each was fixed, what runs on a pull request and nightly, how the models and the code stay in step, the models planned next, and what the models do not cover. The formal-models plan's status markers and the roadmap are brought up to date (models 9 and 10, M4 done) (#128); a section explains TLA+, PlusCal, TLC and model checking for readers new to them (#131).
- The gsm convergence-proof claims (README section 4, the governance guide and its translations) state what the mechanization proves: CI-verified on Coq 8.18, 8.20 and Rocq 9.3; the federated result machine-checked for the acyclic structural core and the monotone-cycle case; the cohomological layer through the cycle basis, with the full H¹ classification paper-proven (#139).
- The gsm convergence wording states the final state once. gsm v0.11.0, which bide required until this release, had a
Buildgap: its commute shortcut did not check what event guards and effects read, so a machine where one event's guard or effect reads a variable another event writes (pay/ship) could be certified convergent when it is not. #142 scoped the docs to that while it held (no "Buildproves every interleaving converges", no proof-derived re-check claimed). With gsm v0.12.0, which fixes the gap (#154), the README section 4 and its comparison-table footnote, the four i18n READMEs, the governance guide,CONCEPTS.mdand the docs-site front page say what backs convergence now:Buildreturns a machine only after the table oracle generated from gsm's Rocq proof re-checks it in-process (and, for combinator rules inside its fragment and within a cost cap, the rules oracle), a federation's own conditions are checked by gsm's Go code, and the required gsm machine gate runs the proof's checkers on every machine the governance examples build.KNOWN-LIMITATIONS.md("Governed state (gsm)") gives the exact scope, including which pairs CC covers and that a verdict recorded under v0.11.0 is not covered by the fix (#155). - The roadmap is brought up to date: P12 to P14 done, models 11 and 12, and gsm convergence (#153).
- Tests: the tool-timeout tests that need their tool to run use a
testing/synctestbubble, so a short deadline can no longer pass while the call is being dispatched (#134). - Tests:
modelsync's fixture git repositories are isolated from auto-gc, maintenance, fsmonitor and the global and system git config, so no background git process races a test's cleanup (#136).
Changed
Run API, cancellation and recovery dispatch (P14)
- Breaking:
agent.Recoverandagent.RecoverLooptake anagent.Resumer(func(ctx, runID, start RunStart) error) instead offunc(ctx, runID) error. Migration: passagent.ResumeAgent(a), or add thestart agent.RunStartparameter to your own callback; a run an earlier version started is not driven byResumeAgent(see the kind entry below) (P14, #138). - Breaking:
RunStart.Inputis aMessage(a user message of one text part is still journaled as a JSON string, so existing records read back unchanged); read its text withstart.Input.Text()(P14, #138). - Breaking:
agent.Samplingandagent.ToolChoicemarshal with snake_case JSON names (temperature,top_p,max_tokens,stop,seed;mode,name), as run:start journals them (P14, #138). - Breaking:
RunTypedandRunTypedNativejournal the typed start (output mode and the answer type's schema), so resuming a typed run started under this version throughRun, or with another type, isErrConfig(P14, #138). - Breaking: a run's first end marker in journal order is its end for every reader (P14, L3): a run whose
run:cancelledprecedes itsrun:completereturnsErrRunCancelled, not its answer, andIsCompletereports it not complete (#138). - Breaking:
ErrNotStartedwraps no category (it wrappedErrConfigearlier in this release cycle) (#138). - P14 writes the kind of every run it starts in
run:start(RecordedStart(...).Kindisagentfor an agent run). Arun:startan earlier version wrote, with no kind and no typed start, is driven by any agent entry point, as before: a session turn, aSendOnceturn or a typed run in flight across the upgrade still resumes through its session orRunTyped(its input is still held). Upgrading: recovery does not drive such a run throughResumeAgentorResumeTyped: the record does not say whether a plain run or a typed one started it, so both returnErrNotResumablefor it (Recoverreports it once per process). Finish the runs in flight before upgrading, or recover them with aResumerof your own, afterResumeAgentinResumeAny, that knows which entry point started each (RunTypedMessagefor a typed run,ResumeRunfor a plain one) (#138). WithToolChoice(ToolChoice{Mode: "none"})is enforced at dispatch: a tool call the model makes anyway is refused with an error result, as a call outside the tool filter is (a typed run's answer tool excepted) (#138).- A saga's rollback request is read before a recorded failure, and a failure's rollback of a saga whose rollback request exists ends
run:cancelled(one moreGet, on that rollback only), and so does a saga sub-run's failure rollback once its tree root was cancelled (up to twoGets more, from the root's store). A drive withoutWithSagaof an aborted saga (RunMessage,ResumeRun) reports*SagaAborted, notErrConfig; a cancellation's rollback that findsrun:abortedfirst reports the abort (#138). - Breaking: a finished agent run returns its recorded answer only to a drive with the input its
run:startrecorded;Run,RunSaga,Streamand the rest with another input areErrConfig, as for an unfinished run (#70) and a finished flow's run. Before, a finished run answered any input with its own answer, so aSendOncekey reused for another message, whose run had finished and was not yet recorded, was answered with the first message's reply and recorded as that message's turn (#137, review R137-2).
Sessions
- A session's turn run is driven under its lease when the store implements
Leaser(MemStore, SQLite, Postgres), asagent.Leasedrives a run: the run loads its journal only once it holds the lease. A worker that finds another holder leasing the run returns its recorded answer if the run has finished, andErrTurnContendedotherwise (model 12, finding S4,TurnLease). Cost: a firstSendturn overMemStoretakes about 14 µs, 25 allocations and 1.7 KiB more (BenchmarkSession_Send, the lease's renewer goroutine and timer), and over SQLite or Postgres two more single-row statements (the lease's acquisition and release) (#137). Leaser.ReapLeasesalso deletes the lapsed leases of runs whose ID contains>(a session's or a sub-agent's run), which no recovery pass takes over, so a worker that died mid-turn does not leave a lease every later lapsed pass reads.MemStore, SQLite and Postgres implement it, andstoretest'sLeaser_ReapLeaseschecks it. A customLeasermust do the same (#137).
Tools (P12)
- Breaking:
agent.Safetyis plain data,{ReadOnly, Idempotent}: comparable, and journaled. The approval gate isToolSpec.Approval, set withagent.WithApproval(agent.SingleApproval())(forSafety{RequiresApproval: true}) oragent.WithApproval(&agent.ApprovalPolicy{...})(forSafety{Approval: ...}).ApprovalPolicyencodes asneedandapprovers(#117). - Breaking:
agent.ToolHandlerisfunc(ctx, ToolCall) (json.RawMessage, error): tool middleware readscall.Useandcall.Spec, and passesnexta copy with otherUse.Argsto rewrite arguments (#117). - Breaking:
agent.Request.Toolsis[]agent.ToolSpec, sorted by name, andToolsDigesttakes[]ToolSpec(the digest is unchanged) (#117). - Breaking: a tool result's
read_onlyfield is replaced bysafetyandapproval(Record.ReadOnlyis replaced byRecord.Safety, andRecord.Approvalis new); a result journaled by an earlier pre-release carries nosafety, so a rollback over it treats every completed call as a write (the safe reading); resuming across pre-releases is not guaranteed (the journal format staysbide.journal.v1-dev) (#117). - Breaking:
govern.EventTool(gov, EventToolConfig{...})andgovern.FederatedEventTool(gov, FederatedEventToolConfig{...})replace their positional forms;audit.AttenuatingSubAgent(name, description, sub, AttenuationConfig{Store, Narrow, Rules}, opts...)replaces its positional form (#117). - Breaking:
mcp'sToolsfails withErrProtocolwhen a server lists a tool whoseoutputSchemais not an object schema (a null one is no output schema) (#117). - Breaking: new construction panics, each with an error wrapping
ErrConfig:audit.AttenuatingSubAgenton a nilAttenuationConfig.StoreorNarrow, or an optionagent.SubAgentrefuses;govern.EventToolandgovern.FederatedEventToolon an invalidagent.ToolOptionin theirOptions(andEventToolonAttestedwith an emptyPolicyDigest) (#117). - Breaking:
mcp.WithSafetysets onlyReadOnlyandIdempotent; gate an MCP tool withmcp.WithApproval.WithCallTimeoutalso sets the tool'sToolSpec.Timeout, which the agent applies (#117). - Breaking: only
agent.SingleApproval()asks for the one-decision gate; it is a distinct value, never inferred from a policy's shape, so anApprovalPolicyliteral with no approvers ({Need: 1}included) isErrConfig, inagent.WithApprovalandmcp.WithApproval.ApprovalPolicy.Clonecopies a policy, aSingleApprovalincluded.ApprovalPolicyhas an unexported field, so an unkeyed literal (ApprovalPolicy{2, approvers}) no longer compiles: name the fields (#117). - Breaking: tool middleware passes
nexttheToolCallit was given (or a copy with otherUse.Args); a call whoseUse.NameorUse.IDit changed, or aToolCallit built itself, fails withErrConfigand the tool is not called (#117). - Breaking:
SagaAborted.UnknownOutcomelists saga steps that failed with an unknown outcome (a retry-safe step that returnedErrToolOutcomeUnknown, or an error after its deadline), which may have committed; their failure records carryoutcome_unknown. The rollback reports them rather than take them for steps that changed nothing, and does not run their compensators on a result they never returned (#117). - Breaking:
Newrefuses (asErrConfig, before any model call) a tool that unwraps (Unwrap() Tool) and is aCompensator, and one that wraps a sub-agent and has aTimeoutor gives it anotherSafety(#117). Newvalidates every tool'sToolSpec.Approval: a policy a customSpec()returns that no option would build (aSingleApprovalwhose fields were changed, an m-of-n policy with no approvers) fails every run withErrConfigbefore any model call, where it used to be approved, fire, and then fail to encode its result (#117).- Breaking: a decorator that embeds a tool (
struct{ agent.Tool }, overridingCall) has noSpecmethod, so its spec, read from the old method set, would have no approval gate and no timeout, and the gated tool inside would run ungated.New(and plan's tool check) now fails closed: a tool that embeds, at any depth, a tool with anApprovalor aTimeoutits own spec lacks isErrConfig("decorator hides the approval gate; implement Spec or Unwrap"). A decorator keeps them by implementingSpec, orUnwrap() Tool:SpecOfof a tool with noSpecmethod takesTitle,Output,ApprovalandTimeoutfrom the first tool on itsUnwrapchain that has one. The check walks the wholeUnwrapchain, inNewand in plan'sBuilder.Tool(atBuild) andRegisterTool(#117). - Behaviour change for
mcp.WithCallTimeoutusers: an error the server reports after the call's deadline (a late JSON-RPC error, a known failure) now has an unknown outcome under the agent's timeout rule, so a side effect records nothing and halts on resume instead of recording the failure. The agent cannot tell a late failure from a late answer to a call that took effect, and takes the safe side (#117). - The journal encodes
SingleApprovalas{"single":true}, the bide protocol's form (ApprovalPolicyhasMarshalJSONandUnmarshalJSON); m-of-n policies encode as before (#117). - The approval policy decodes strictly:
{"single":true}alone isSingleApproval;singlewith any other value or besideneedorapprovers, a policy with noneed, and duplicate, unknown or mistyped members areErrProtocol. That isApprovalPolicy.UnmarshalJSON, for a policy being configured or received; the approval inside a storedRecord(DecodeStoredRecord,DecodeRecord, and so a proof bundle'sRecord()) is read leniently, ignoring members this version does not know, so a journal a newer version wrote stays readable (#117). - A call that reached the base handler but returned before its tool's
Callbegan (a saga-arguments write cut off by the saga's cancellation) counts as not called: a per-call word, set by compare-and-swap immediately before the tool and sealed by the loop when the chain returns, proves no call began and none can (#117). - A tool call's state is terminal once the middleware chain returns (a refused call too), so a
nextleft running cannot reach the tool after the loop recorded the call; its saga arguments are journaled only once it reaches the tool, and "already ran" (ErrToolReinvoked) is decided by the compare-and-swap that begins the tool, so it is said only of a tool that began. Once the chain has returned, no invocation ofnextbegins the tool, even a retry-safe one an earlier invocation reached: it used to begin again after the call's result was recorded, and after a saga's compensation (found by the TLA+ tool-call model, T5) (#117). - Breaking (behaviour): a side effect whose tool began and did not itself fail is never recorded as a known failure. When the chain returns an error for it (a middleware turned its success into an error, or returned a retry's refusal, or left the call running) and the context is live, the error wraps
ErrToolOutcomeUnknown, nothing is recorded, and the run halts. TheToolMiddlewarecontract says so. Found by the TLA+ tool-call model. In a saga the same holds for a retry-safe step that changes state (Idempotent, notReadOnly): its failure is recorded with an unknown outcome and listed inSagaAborted.UnknownOutcome, never skipped by the rollback as a step that made no change; outside a saga a retry-safe tool's error stays an ordinary failure (#117). - In a saga, a compensable retry-safe step that changes state (
Idempotent, notReadOnly, aCompensator) journals its accepted arguments before each call whether or not a middleware changed them: the record is the step's "may have begun" marker, since such a step writes no attempt marker. When an earlier drive's attempt journaled it and left no outcome, a later known failure of the step (the tool's own error, a middleware's refusal,ErrToolNotCalled) is recorded with an unknown outcome and listed inSagaAborted.UnknownOutcome, and a rollback reports such a step that was later denied; within one drive the same holds for any retry-safe write a middleware ran again after an invocation that did not itself fail. A store fault writing that record (or a rewritten call's arguments) fails the run and records nothing, so a re-drive calls the tool; it used to be recorded as a known failure. Found by the TLA+ tool-call model (T3) (#117). - Breaking (behaviour): a result needs positive proof: a tool middleware chain that returns a result while any invocation of the call's tool is still running in the process (a middleware that left
nextrunning and answered itself, from a cache say; a sibling invocation; one a cancelled drive left behind, counted per run and call) has an unknown outcome (ErrToolOutcomeUnknown), so a side effect halts and a retry-safe saga step is listed inSagaAborted.UnknownOutcomeand never compensated, where it used to be recorded as succeeded and could be compensated before its effect landed. A saga rollback that re-runs a retry-safe step to learn the result to compensate, and gets an unknown outcome (a result check that rejects every success, say), reports the step inSagaAborted.UnknownOutcomeand finishes, instead of stopping on every drive. A rollback re-run that a middleware answered without reaching the tool is not taken for the step's result either. The count is read whatever state the chain ends in, so a re-drive's cache answer while the earlier drive's invocation still runs is unknown too (T6). Found by the TLA+ tool-call model (T4) (#117). - Tool timeouts are judged by the deadline itself (a context's error lags its timer), for the tool's timeout and the run's own deadline; only a call whose tool was actually called can be late or unknown: a call a tool middleware ended first is a known failure, and a tool whose deadline passed in the middleware is not started. Plan flows' Tool nodes apply the wrapped tool's
ToolSpec.Timeoutwith the same rule (#117). NextOnceKeyandSafety.Idempotentdocument that once keys are scoped to one tool call: a retry the model makes is a new call with new keys, so dedup across the model's retries needs a business key from the arguments (#117).
Construction under Build (P13)
- Breaking:
agent.WithRetrieval(r, k)is an agentOption, not a model middleware: the agent retrieves for the run's user message as a journaled engine step, at a drive's first model call, and every model middleware sees the request with the documents in it. Migration:a.Use(agent.WithRetrieval(r, k))becomesagent.Build(model, j, agent.WithRetrieval(r, k))ora.With(agent.WithRetrieval(r, k)). A retrieval outside an agent run no longer exists. The retrieval runs before the model middleware chain, so model middleware neither retries it nor prevents it:middleware.Retryno longer retries a failed retrieval (useagent.WithRetrievalRetry), and a model middleware that refuses the call (a policy gate, a spend cap) runs after the query has reached theRetrieverand the documents are journaled. A policy that must keep a query from the store wraps theRetriever(seeagent.RetrieverFuncand the RAG guide) (#127). - Breaking:
agent.RetrievalTool(name, description, r, k, opts...)takes the tool's name and description and the tool options, asFuncdoes;RetrievalOption,RetrievalNameandRetrievalDescriptionare removed (#127). - Breaking:
trace.Instrument(tracer, opts...)returns anagent.Option. Migration:trace.Instrument(a, tracer)becomesagent.Build(model, j, trace.Instrument(tracer))ora.With(trace.Instrument(tracer))(#127). - Breaking: the context decorators
agent.WithIdentity(ctx, id),agent.WithWaker(ctx, w)andagent.WithClock(ctx, now)are renamedContextWithIdentity,ContextWithWakerandContextWithClock(transitional), and theWithnames are the options. A value bound to the run's context takes precedence over the agent's option (#127). - Breaking:
agent.WithNow(aResolveOption) is replaced byagent.WithClock, andagent.StepSafetybyagent.WithSafety, which a tool and aStepboth take.WithClock(nil)isErrConfig(WithNow(nil)was ignored) (#127). - Breaking:
agent.Parallel(ctx, d, runID, tasks, opts...)takes the tasks as a slice and the concurrency cap asWithMaxConcurrency(n)(a negative n isErrConfig); the positionalmaxConcurrencyis removed (#127). - Breaking:
LeasetakesLeaseOptions,RecoverRecoverOptions andRecoverLoopRecoverLoopOptions.WithLeaseHolderandWithLeaseTTLfit all three;WithRecoverInterval,WithRecoverConcurrencyandWithRecoverErrorsfit onlyRecoverLoop, so passing one toRecoverorLease, which ignored it, no longer compiles (#127). - Breaking:
agent.RunScopeandagent.InSagaare replaced byRunInfoFrom:SubRunID(info.RunID, info.ToolUseID)is a call's sub-agent run ID, andinfo.Sagawhether its run is a saga. A saga rollback's re-run of a retry-safe call now carries that call'sRunInfo(it carried whatever scope the rollback's context held) (#127). - Breaking (behaviour): whether a tool call is in a saga is its own run's flag (
RunInfo.Saga). A plain run started from a saga's tool call (child.Runwith aSubRunForID) is not a saga, and neither are its calls: they no longer inherit the saga mode of the call that started the run, so such a run's compensable calls do not journal their accepted arguments, and its sub-agents run as plain runs. Start the sub-run withRunSagato keep it a saga. A sub-agent (SubAgent) called from a saga runs as one, as before (#127). - A programmatic sub-run is refused (
ErrConfig) once the tool call whose context names it (SubRunFor) has returned, so a goroutine that outlives its call cannot start a sub-run no call owns, which no rollback and noRecoverwould reach. In a saga, a programmatic sub-run's name must be at most 96 bytes once escaped, so its link can name it (#127). Agent.WithSystemPromptandAgent.WithSystemPromptFuncshare one slot: the later call wins (the function used to win whatever the order) (#127).- The system prompt function (
WithSystemPromptFunc) is called once per drive just before the drive's first model request, and not at all by a drive that sends none: reading back a finished run, or a resume whose pending tool calls pause or halt, no longer fails when the function does. A resumed turn's pending tool calls now run before the prompt function is called (it was called first, before them) (#127).
Recovery and leases
- Breaking:
agent.RunFilter.Admitstakes a third argument,lapsed func() bool, which reports whether the run's lease has lapsed; it is called only when the filter setsLeaseLapsed(#126). - Breaking:
agent.Leaserhas a fourth method,ReapLeases(#126).
Delegation
audit.AttenuatingSubAgentreuses the grant a resumed delegation journaled instead of minting another, journals a delegation that ran without a grant, and refuses a delegation resumed under other authority, from an expired bound grant, or onto a sub-run that has records but no journaled authority, with anErrConfigthat records nothing (the run stops once siblings in flight finish; a sibling's pause is joined to it), so a re-drive under the right grant continues it. A delegation cannot run past its grant'sNotAfterUnix: every tool call in its sub-run is refused, recorded, once the child grant has expired, and a delegation resumed after its journaled grant expired fails, recorded (a saga rolls back). A child grant's subject must be the sub-agent's name. Its rollback binding verifies the journaled grant (signature under the bound signer's key, attenuation of the bound parent, its subject) and stops withErrConfigwhen no grant and signer are bound. A failure to read or write the delegation's authority in the store records nothing, so a resume retries the delegation, as for a plainSubAgent; only value records count as journaled authority (#117).- Breaking (journals): a saga journaled by an earlier pre-release that holds an ungranted
audit.AttenuatingSubAgentdelegation cannot be rolled back after the upgrade: its sub-run records no authority, and the rollback stops (ErrProtocol) rather than guess. Finish or roll back such sagas before upgrading; journals are not promised across pre-releases (#117).
Performance
- The counting-store budget (docs/design/api-v1.md, 10.3) is raised by P14's reads, a maintainer decision: every model turn after a drive's first costs one more
Get(run:cancelled), a side-effect call one more (run:cancelledonce its claim is won), a completion one more (two on a saga: the end markers read back), and a recovery pass fourGets per driven run (itsrun:start). The #138 review added oneLoadafter a first drive writesrun:start(and after a limit amendment), reading the entries' names only; a sub-run's checks read its tree root's cancellation (up to twoGets each); a flow node that runs readsrun:cancelled(oneGet), and a flow's completion and its resumed open read the end markers. Over a transitionalDurablewith noJournal, each of these reads is aHistory(P14, #138). - A recovery pass costs less than it did before P14 although it reads each driven run's
run:start(rule 14): the commonest start (a plain text agent start, with or without its kind) is read without the JSON decoder, exactly, its salt checked as the full decoding reads it (anything else takes the JSON decoding, into a pooled record whose name, kind and salt are matched without decoding strings; a record with any member but those and its result, or one of them twice, takes the full decoding, so a start either path reads is the start the full decoding reads, which a fuzz target checks), and a leased drive starts its renewal goroutine at the first renewal (half the TTL) rather than with the drive, so a short visit (a halted run, a run over since the listing) starts none.BenchmarkRecoverPass10kagainstmainin the bench workflow's A/B on the final code: -33.6% and -30.8% time, -6.6% B/op, -5.9% allocs (runs 36957790848 and 36957793253; +95% with the first P14 draft). Every leased drive (Lease, a session turn) gains the renewer change. - P14's per-run reads (the
Loadafterrun:start, the turn boundary'srun:cancelled, the end marker's read-back) cost the closed-loop fanout scenario its tail onMemStore, whose one mutex every read and write takes: a drive's own cancellation checks and a writer's read-back no longer look the run up in the journal's known-runs set on a miss (the drive opened, or the writer just wrote to, the run),MemStorecopies entry data outside its mutex (entries are never modified once appended),run:startis encoded in one pass, the input in place (the bytes unchanged), and a drive grows its goroutine's stack once, before it starts (probeDriveStack, sized against the drive's stack need, whichTestDriveStackHighWaterpins), rather than twice deep inside it. The bench workflow's A/B againstmainon the final code (runs 36957790848, AMD EPYC 9V45, and 36957793253, AMD EPYC 7763; 21 interleaved runs per binary, medians):cmd/benchoverhead (mean, p90, p99) -2.5%, -30.4%, -37.5% and -1.9%, -16.7%, -14.3%; fanout -1.1%, -9.5%, -5.2% and -5.6%, -9.4%, -13.7%; benchstat (10 runs per ref)BenchmarkRunTurns+4.0% and +2.1% time, neither significant (p=0.165, p=0.280), +0.3% B/op, +0.9% allocs (8 per run);BenchmarkToolCallSideEffect-1.7% (not significant, p=0.796) and +3.3% (p=0.035) time, +0.5% B/op, +2.6% allocs (8 per call). Before these changes fanout was +3.4% to +10.6% at p99. - Breaking:
agent.Recordshrinks from 368 to 256 bytes, from the 384-byte Go size class to the 256-byte one, so each heap copy of a record (a decode, an encode, a store's put) is smaller. Two groups of rarely set fields move behind embedded pointers:*agent.ModelTurnholdsFinish,RawFinish,Model,PromptDigestandToolsDigest(set onStepModelrecords only), and*agent.ApproverSignatureholdsApprover,ApproverAlgandSignature(set on decisionsSubmitDecisionrecords only). Direct access to those fields no longer compiles. Migration: read them with the nil-safe methods of the same names (rec.Finish(),rec.Approver(), ...), which return the zero value for a record that carries none, and build a record withModelTurn: &agent.ModelTurn{...}orApproverSignature: &agent.ApproverSignature{...}. The journal encoding is byte for byte the same: the members keep their names, order andomitemptybehaviour, and existing journals decode as before. Measured A/B against main on two 4-vCPU runners (bench.yml, 21 interleaved runs per binary, AMD EPYC 7763 and 9V74): the overhead scenario ran +0.7% and +1.6% runs/s, mean latency -0.4% and -1.6%, p90 -1.0% and -9.1%, p99 -4.4% and -7.8%; the fanout scenario ran -1.4% and +1.2% runs/s, within the runners' noise, with mean, p90 and p99 within 2% either way. In the Go benchmarks (benchstat, 10 runs per ref),RunTurnsallocates 4.4% fewer bytes (73.2 to 70.0 KiB/op) andToolCallSideEffect6.0% fewer (26.1 to 24.5 KiB/op), with 1.3% more allocations (a model record'sModelTurnis its own allocation);RecoverPass10kis unchanged, and no time per op changed significantly (#140).
Governance (gsm)
- Upgrade: bide requires gsm v0.12.0 (
govern,govern/postgreslog,govern/redislog,govern/sqlitelog,examples/governandintegration; was v0.11.0), so an application that usesgovernbuilds with gsm v0.12.0. ItsBuildchecks every pair it checks for CC (every event pair, or only the pairs declared withIndependent) exactly, with no footprint shortcut, which fixes v0.11.0's gap for guards and effects that read another event's writes, and returns a machine only after the table oracle (and, within its fragment and cap, the rules oracle), generated from gsm's proof, re-checks it in-process (scope in known limitations). Upgrade impact, from gsm's upgrade notes: rebuild every machine, sinceBuildmay now reject one v0.11.0 accepted; event names, variable names and enum labels must be unique; closures (effects, repairs, morphism maps) must return states of their machine;Machine.MergeProjectionreturns an error for out-of-domain values;BuildOrSynthesizefalls back to synthesis only when the compensation failed; re-issue gsm certificates withCertify;Report.Stringtext changed;Buildcosts more (no measurable slowdown for the governance examples).PolicyDigestis unchanged, so anchored policy digests stay valid. A convergence verdict orConfluenceCertificaterecorded under v0.11.0 is not covered by the fix: rebuild and record a new one. The gsm machine gate pins ecaf453, the v0.12.0 release commit (#154). - Behaviour change:
govern.CertifyConvergencereportsConvergesonly for a machineBuildreturned (Report.Assuranceset). Under gsm v0.12.0 a machine can pass gsm's WFC and CC checks and still be refused by the in-process oracle gate (Report.OracleDisagreement); its certificate now saysConverges: false(#154). - Three governance examples declare the gsm rules of six registries with combinators instead of closures (
examples/govern/compose's inventory,coordination's two registries,mesh's three lines), so both of gsm's extracted checkers, not only the table checker, certify them in the gsm machine gate. Their behaviour and output are unchanged (#146). - The gsm machine gate's pinned gsm commit followed gsm's
main: 219dcaa, whose rules checker carries normalization-confluence#11's fixes (#147); 150133b, whose table checker carries normalization-confluence#12's speed-up (#149); c2eecd6, which runs the proof's table oracle in-process on every successfulBuild(gsm#12), so the examples also Build under it (#150); fcf884a, which records only the synthesized machines a program receives, not the candidate synthesis certifies (gsm#16), so the gate reads 23 machine records again (#151); and 031db4e, whoseBuildalso runs the proof's rules oracle in-process (gsm#17), so every registry the governance examples build withBuildis certified by both oracles (the synthesized machine incoordinationby the table oracle), with no measurable change in run time (#152).
Formal models, CI and dependencies
- The required Models job checks four configurations at a time (one TLC worker each) and, on a pull request, only the models whose
spec/tla/<model>/directory the pull request changes; the merge queue and main still check every model. The largest passing pull-request configurations of models 1, 2, 7 and 8 (1.4 to 2.9 million states) and two of model 1's liveness checks run nightly, each safety one with a smaller pull-request counterpart of the same invariants, so every path keeps a passing configuration and its vacuity run on pull requests. The job takes about 6 minutes, down from 9 to 15.5 (#132). - CI runs its two slowest jobs in parallel shards, with no check dropped: the Models configurations in six shard jobs by configuration directory (
check.sh's newTLC_DIRS), and the Linux Go tests in four jobs (the core'sagentpackage, the rest of the core, the other modules, and thenojsonv2run), each with the same flags as before. The required check names are unchanged: a final Models and Test (ubuntu-latest) job needs every shard and fails unless each succeeded (or nothing needed checking)..github/scripts/shards.shfails CI unless every module, every package of the split core and every pull-request configuration is in exactly one shard, naming the line to change, and its self-test proves it catches a missing, doubled or unknown entry. With every model and every module checked, the Models workflow took 3.5 to 4.1 minutes on this pull request's runs, against 7.6 to 7.8 on pushes to main and 9.8 to 10.7 in the merge queue before, and the CI workflow 5.2 to 6.0 minutes, against 7.5 to 8.4 on main (each from the run's creation to its last update) (#143).
Deprecated
- The string entry points are transitional (P14):
Agent.Run(ctx, runID, input string),RunSaga,RunResult,RunSagaResult,Stream,StreamSaga,AgentStream.Final,Session.SendandSession.SendOnce(strings), and the functionsRunTypedandRunTypedNative. The 1.0 rewrite removes them forRun,Stream,RunTypedandSession.Send/SendOncewith aMessageand run options (#138). agent.SpecOfis transitional: the 1.0 rewrite givesToolaSpecmethod, andt.Spec()replacesSpecOf(t)(#117).agent.Newand the builder methods (Use,UseTool,WithSampling,WithToolChoice,WithSystemPrompt,WithSystemPromptFunc,WithMaxTurns,WithTokenBudget,SetMaxConcurrency,WithApproverVerifiers,WithToolErrorRedactor) are transitional: the 1.0 rewrite removes them and renamesBuildtoNew. They still change the agent they are called on.Newkeeps its lenient checks; onlyBuildandWithrefuse a reserved tool name and a non-object input schema (#127).agent.ContextWithIdentity,ContextWithWakerandContextWithClockare transitional: the run API takes the identity, Waker and clock as run options (#127).
Removed
- Breaking:
Safety.IdempotencyKey(the SDK never called it;Idempotentsays the same, and the tool derives its own downstream key or usesNextOnceKey),Safety.RequiresApprovalandSafety.Approval(seeWithApproval) (#117). - Breaking:
agent.ToolSafety,agent.WithToolSafetyandagent.ToolErrorText: tool middleware readsToolCall.SpecandToolCall.ErrorText, and no tool data travels in the context (#117). - Breaking:
govern.AttestedEventTool; anEventToolConfigwith aPolicyDigestis the attested form. Migration:AttestedEventTool(gov, name, desc, event, digest, safety)becomesEventTool(gov, EventToolConfig{Name: name, Description: desc, Event: event, PolicyDigest: digest, Attested: true, Safety: safety}). The old form accepted an empty digest and still recorded the state digest and acting identity;EventToolConfig{PolicyDigest: ""}is the plain tool, which records neither, so a config that setsAttestedwith an emptyPolicyDigestis refused (EventToolpanics withErrConfig) rather than silently dropping the attestation (#117).
Fixed
Run API (P14)
plan: a Tool node called a tool whose context's deadline had already passed (its timer not yet run), which the agent's base handler refuses; it now fails withErrToolNotCalledwithout calling the tool (P14, #138).
Sessions
ResolveHaltRef's live-driver check leased everything before the first>of the halted run's ID. It did not see a live session turn's driver, which leases the turn's run (<id>>@turn/<n>), so it could resolve a call the turn was running, and it took a root run named like the session for a live driver (*HaltInFlight). It now leases the halted run's tree root: the turn's run for a session turn and every sub-run inside it, the root run otherwise, the runRunInfo.RootRunIDnames (#137, review R137-1).- A stale
Sessionhandle refused every new message for a turn another handle had finished: a handle whoseSendof that turn's message failed or paused kept the turn open in its own view, andSendof another message returnedErrConfigwithout reading the journal. It now reads the journal once before it refuses (model 12, finding S1) (#137). - Two callers on one
Sessionhandle sending one message (a redelivery, or a retry while the first still ran) recorded its turn twice, forSendand for oneSendOncekey: the second caller's append started past the slot the first had recorded and reloaded. A handle now skips the append of any run among the turns it has loaded (model 12, finding S2) (#137). - Two workers given one session message both drove its turn's run, and each counted only the spend it had seen, so a turn under
WithTokenBudgetspent up to its budget once per worker (twice its budget in the test). Over a store with leases one worker drives a turn at a time (model 12, finding S4) (#137).
Tools (P12)
agent.Newread a tool's spec twice, once for its name and once for the rest, so a tool whose answers differed had its calls decided by the second (#117).
Delegation
- A saga rollback did not recurse through a tool that wraps a sub-agent (
audit.AttenuatingSubAgent): the sub-run's compensable writes were left in place and the delegation reported uncompensated. A resumed tree's token budget missed such a sub-run's spend the same way (#117). - A saga rollback into an
audit.AttenuatingSubAgent's sub-run now compensates under the child grant and identity the sub-run journaled (none, if the delegation ran without a grant), never the parent's broader authority (#117). - A saga rollback into a resumed
audit.AttenuatingSubAgentdelegation whoseAttenuateFuncgave each grant its own ID stopped: the sub-run held two grants. The delegation now keeps its first grant (#117). - A saga whose
audit.AttenuatingSubAgentdelegations were minted under two grants (the root grant expired or was rotated between drives) could never finish its rollback: each journaled grant was verified against the one grant bound, and the walk stopped at the first minted from the other, whichever was bound. Bind every such grant (WithGrantandWithRollbackGrants); a journaled grant whose parent is none of them is still refused (ErrNotVerified). Minting from an expired bound grant now says to keep it bound for the rollback beside the live one (model 11, finding D1) (#133). - A saga rollback into an
audit.AttenuatingSubAgent's sub-run ran a retry-safe write of the sub-run again after the delegation's grant had expired: the rollback rebound the child grant without the mark the expiry guard checks. The mark is kept, and a re-run the guard refuses is listed inSagaAborted.UnknownOutcomeand the rollback goes on (model 11, finding D2) (#133). - An
audit.AttenuatingSubAgentdelegation under a grant bound with a nil or typed-nil signer (a nil pointer in theSignerinterface) panicked: a typed nil when it minted or resumed, a nil one when it resumed, and a nil one was recorded as a failure when it minted.WithGrantnow binds the grant with no signer, and the delegation is refused withErrConfigand records nothing;WithRollbackGrantsbinds nothing without a signer (#133). - A saga rollback skipped a sub-agent call with an error result (in a plain run inside the saga's tree, a sub-agent whose run failed after a compensable write), so the sub-run's writes were neither compensated nor listed. The rollback now walks a sub-agent call's sub-run whatever its result (model 11, finding D3) (#133).
Governance (gsm)
examples/govern/meshwas not convergent (each signal event set the level, so order mattered); gsm v0.11.0 wrongly certified it; signals now only raise the level (max), which commutes.TestSignalsCommutechecks every pair of signals in both orders, and the example no longer shows a line standing down, which no order-independent event can do (#145).
Recovery and leases
RecoverLooptook over a dead holder's run late when halted runs were listed before it: each pass visited every unfinished run in order, halted ones included, so the takeover waited about one visit per halted run, and a lease that lapsed just after the pass tried the run waited for the whole next pass (finding L1 of the run-lifecycle model, #124). It now runs a second loop on the same interval, with slots of its own, that drives only the runs whose lease lapsed, so a dead holder's run is taken over within about one interval of its lease lapsing however many halted runs the store holds. The full pass is unchanged and still re-drives halted runs and runs that held no lease (#126).
0.9.0 - 2026-10-01
Added
Formal models
- A TLA+ model of the claim protocol (attempt claims, not-started records, numbered retries, remembered claims, the resume gate, the Step flight and halt resolution), checked with TLC in CI by the new Models workflow, a required check on every pull request; see spec/tla (#100).
- The claim model also covers
ClaimAttempton a single key (plan flows), thependingClaimseviction, and a halt resolution running in a driver's process (#106); it follows the Store/Journal core as merged, including the claim on the lease path (#107). - The claim model covers the approval gate (1-of-1
Approveand m-of-nSubmitDecisiontallies) together with halt resolution (#108). - A TLA+ model of flow semantics (switch and loop replay, per-iteration keys,
run:complete,Flow.ResolveHalt), checked in CI (#110). - A TLA+ model of spend accounting (
@llm,@spend,@spend-late, hedged losers, two drivers), checked in CI (#111). - A TLA+ model of the bide protocol's claim rules for remote tool calls (claim at assignment, begin records, abandons), checked in CI (#112).
- The approval model covers approvers' key sets (refused on any overlap) and a verifier resolver that changes between the gate's check and its count (#115).
- The design and plan for formal models of the coordination protocols, docs/design/formal-models.md (#99).
Store and Journal
agent.Store, the storage port (Insert,Get,Loadof anEntrywith an opaque, commit-orderedSeq), with its requirements A1 to A8 documented on the type, andagent.Journalover it (NewJournal,Get,History,Records,Format): the journal owns memoization, the record encoding and salt, the journal format header, attempt claims, not-started records and recording an outcome after the caller's context is cancelled (#92).- The journal format header: every run's journal starts with an
@journalrecord (agent.StepHeader) namingagent.JournalFormat(bide.journal.v1-dev, the one tag every pre-release writes until 1.0, which switches tobide.journal.v1and refuses every pre-release journal); a run in another format, or with no header, is refused with*agent.JournalVersionError(wrappingagent.ErrJournalVersion) before anything is read or written (#92). agent.RunFilter(After,Prefix,ExcludeHolding), which SQL stores evaluate in their query (#92).agent/storetest: every store requirement (64 goroutines racing one name through three handles, prefix-closed reads, byte fidelity, context, iterator hygiene), the header checks, shared in-flight steps, claim reuse, record fidelity, andCheckWrapper, which checks, given at least two contexts that differ in what the wrapper reads from a context, that a store wrapper's mapping of run IDs and names does not depend on the context (A1), and its use ofUnwrap(#92).agent/agenttest.CountingStore, and a test that holds the engine to an exact budget of store round trips per operation (#92).Record.Raw(the stored bytes),Record.Salt(),Record.ClaimID(),Record.Format,Record.Redacted(the reserved redaction tombstone), andRecord.MarshalJSON, which writes the journal encoding (#92).store/sqliteimplementsagent.Leaser, on a lease connection of its own with a short busy timeout and expiry from the database's clock;Openopens a writer, readers and a lease pool (#92).store/sqlite.Newandstore/postgres.Newover a*sql.DB,WithTablePrefix, and a schema version row: a database with a newer schema is refused (#92).- Journals in one process over one store share in-flight steps. A claim whose insert failed (it may have committed) records that its attempt did not start, keyed by its claim (
attempt:not-started:<claim>:<marker>), so a re-drive in any process re-attempts the effect instead of halting over one that never ran. If that record cannot be written either (or is written and reported failed), the process remembers the claim, and the next claim of the marker in the process, or a resume that meets it, writes the record again; every claim takes a fresh id, so an effect never runs under a marker that is, or can become, recorded as not started.plan's conformance check ignores not-started records (#92). - Benchmarks
BenchmarkRunTurns,BenchmarkToolCallSideEffect,BenchmarkStep,BenchmarkRecoverPass10k,BenchmarkAnchoredInsert,BenchmarkSQLiteInsertandBenchmarkPostgresInsert(#92).
Pauses and halt resolution
agent.Pause, the sealed interface every pause satisfies, withagent.RunRef,IsPauseandAsPause. Its five kinds areApprovalPending,InterruptPending,SignalPending,TimerPendingandOutcomeUnknown; code outside the package cannot add one (#90).agent.ResolveHaltRef(ctx, store, HaltRef, Outcome, ...), one resolution for tool andStephalts, withOutcomeUnknown.Ref,HaltRef,OpRef,OpKind(OpTool,OpStep),HaltCause(HaltCrashed,HaltContended) andOutcome(whoseEvidencemarks the resolution reconciled). The 1.0 rewrite renames itResolveHalt(#90).- Verbs named for the pause they answer:
SubmitDecisionwithDecision,AnswerInterruptandEnqueue;agent.Wake(#90). agent.HaltInFlight,agent.HaltAlreadyResolved,agent.ErrAlreadyResolved(wrapsErrConfig) andagent.WithoutLiveDriverCheck(#90).
Flows
agent.RunStart.Kind(agent.RunKind:RunKindAgent,RunKindFlow) andagent.RunStart.Flow(agent.FlowRef): aplanflow's run recordsrun:startwith kindflow, its flow's name and its input as JSON, andRecordedStartreads them back. An agent run'srun:startrecords no kind, which reads asRunKindAgent(#103).plan.Flow.ResolveHalt:agent.ResolveHaltReffor a node of the flow, which first checks that the halt names a node of this flow, that the run is a run of this flow (its recorded start names the flow and its recorded topology digest is the flow's), and that a successful outcome decodes as the node's output type, since a resolution is final and one the flow could not read would leave the run unable to continue (#103).agent.ErrNoLiveAttempt(wrapsErrConfig) (#103).plan's fault-schedule exploration of flow lowering (crashes, store errors that did or did not commit, and cancellations at every write, across processes) as a permanent test, bounded by default (BIDE_EXPLORE=1explores every process plan) (#103).
Model calls and spend
agent.ModelCall(Request,Model,RunID,Turn),ModelCall.AddHookandModelCall.Attempt,agent.ModelResponse(Message,Usage,Finish,RawFinish) andagent.ModelAttempt(what a hook'sAftersees, includingDiscarded, the discarded spend a replayed turn reports) (#104).agent.CallModel(ctx, model, req, mw...)sends one model call outside an agent through the same model handler an agent uses: hooks, clipped request slices and the response checks (#104).- Model records journal
Record.Finish,Record.RawFinish,Record.Model(theModelInfoof the model that answered) and per-turnRecord.PromptDigestandRecord.ToolsDigest(agent.PromptDigest,agent.ToolsDigest: SHA-256 of the system prompt and of the tool set the turn was sent, after middleware).ModelInfohas JSON tags (#104). middleware.CostMeter.Snapshot()returns aCostSnapshot(Answer,Spend,AnswerUSD,SpendUSD) read under one lock (#104).trace.Modelrecordsgen_ai.response.finish_reasons; an empty reason is recorded asstop, as the journal records it (#104).agent.ModelCall.OnAnswer(key, fn):fnruns once per recorded turn with the turn's answer, after the journal holds it, however many targets or attempts a middleware sees; it reports whether the call belongs to a turn (#104).- A run that ends (completes, pauses or fails) waits, for at most two seconds in all and not past its context, for model requests still in flight (a hedge loser, a request a middleware left running) and journals their usage in a late spend record,
@spend-late/<id>, whichResult.Spend, the budget andReplaycount. A driver whose model record another driver of the run recorded first journals its request's spend there too (#104).
Proofs and signatures
audit.Signerandaudit.Verifiername their scheme with a typedaudit.Alg(an alias of the newagent.Alg) and exposePublicKey(), and aVerifierreports itsKeyIDs(), so everyaudit.Verifieris anagent.ApproverVerifier;audit.NewVerifierandParsePublicKey(andverify.NewVerifier) refuse a weak Ed25519 key (audit.ErrWeakKey);audit.NewVerifier(alg, pub),audit.VerifierOf(signer), and the<alg>:<hex>key text form (audit.FormatPublicKey,audit.ParsePublicKey). The standaloneaudit/verifypackage gainsverify.Head,verify.Verifierandverify.NewVerifier, and checks tree heads under ed25519, ML-DSA-65 and the hybrid (#105, #109).audit.ErrNotVerifiedandaudit.ErrMalformed(wrapsagent.ErrProtocol;audit.ErrFormatnow wraps it), the sentinels every verifier's error wraps (#105).audit.JournalExport(formatbide.audit.journal-export.v1) andaudit.ExportJournal: a run's journal as each record's stored bytes, the input ofbide-audit proveandprove-absent.audit.JournalLeafHash, the leaf hash a redaction tombstone records; a redacted record keeps its place in every tree (#105).- Format constants
audit.STHFormat,audit.AnchorEntryFormat,audit.GrantFormat,audit.JournalExportFormat;audit.EvidenceKind(a typed string) withKindRunCertificateandKindConsistency;agent.ReasonAlg(#105). agent.Decision.Algandagent.Record.ApproverAlg: an m-of-n decision is journaled with the scheme it was signed under (#105).
Approval key identity
agent.ApprovalPolicy.ValidateKeys,agent.ApprovalTally.Excluded,agent.ReasonSharedKey,agent.ReasonNoKeyID,audit.KeyID,audit.CheckEd25519PublicKeyandaudit.ErrWeakKey;KeyIDsonaudit.Ed25519Verifier,MLDSAVerifierandHybridVerifier(#109).
Documentation, testing and tooling
- The
RecoverLoopdocumentation (godoc, the recovery guide, known limitations) states that a dead holder's run is picked up within about one interval of its lease expiring only while a pass is short: a pass costs about five store round trips per unfinished run it lists, halted runs included, so a long pass delays pickup. Bounding a pass's cost follows v0.9.0. - The bide protocol (
bide.protocol.v1), an accepted design for SDKs in other languages, with its claim rules checked by model 2; not implemented yet (docs/design/protocol.md, #95). - The pre-1.0 API redesign proposal, docs/design/api-v1.md (#64); the roadmap adds bide underneath other agent frameworks (#102).
scripts/release.shpushes at most three tags per push,release-modules.ymlchecks a module tag on manual dispatch, andbench.ymlhas an A/B mode that runs two refs interleaved in one job (#101).- Tooling: a nightly Explore workflow (
.github/workflows/explore.yml, also on demand) runs the claim and flow-lowering fault-schedule explorations at their full bound (BIDE_EXPLORE=1) and keeps the schedule signatures (BIDE_EXPLORE_SIGS) as an artifact kept 30 days; it is not a required check (#119, #121). - The concurrent claim exploration is deterministic: its scheduler finds quiescence with
testing/synctestinstead of timers, runs one driver at a time, and branches on the races between drivers that share a step call in flight, so every run explores the same schedules (more than before, with each one the timing-based scheduler reached) and no longer slows or times out under CPU load (#118). bench.ymlandcmd/bench/README.mdreport latency as the mean (concurrency / throughput), p90 and p99; the p50 stays in the raw output only, since the closed-loop harness's p50 is bimodal (#116).
Changed
Store and Journal
- Breaking: every journal starts with the
@journalheader, soHistoryreturns it first and record indices (audit leaves,ProveRecord) shift by one; journals written by v0.8.0 and earlier have no header and are refused, not resumed (#92). - Breaking:
agent.Lister.Runstakes aRunFilterand returns an iterator of run IDs in ascending order;RecoverandRecoverLoopask the store for the runs holding no terminal marker and read none of the finished ones (#92). - Breaking:
agent.Capabilitytakes aStoreand followsUnwrap() Store; the engine still finds capabilities behind aDurable(#92). - Breaking:
Record.SaltandRecord.Claimare read-only: useSalt()andClaimID(); the journal sets them (#92). - Breaking: a
Stepthat is not retry-safe and returns a pause (Interrupt,Sleep,Await, a pending approval) isErrConfig; its marker stays, so its next attempt halts. Inside a tool call the error is not recorded as a tool failure, so the call halts on resume rather than being retried under a new id. Put the pause in a retry-safe step of its own (#92). - Breaking:
store/sqlitenames its tablesbide_steps,bide_leasesandbide_schema_version, andstore/postgresits leases tablebide_leases.sqlite.Openrefuses a file whose v0.8.0-or-earlier journal table (steps) holds rows, withErrJournalVersion, rather than open it as empty; Postgres refuses each such run on its first read, without writing to it. Stop every v0.8.0 (or earlier) node before starting this version: the two do not share leases, and v0.8.0 cannot read the new journals (#92). - Breaking: a store wrapper whose
DoandHistorycome fromMemStore, a SQL store or an embeddedagent.Durable(directly or through another wrapper) while itsInsert,GetorLoadcomes from elsewhere isErrConfigwhere aDurablegoes, since thoseDoandHistorywould write past its methods; passagent.NewJournal(wrapper)instead (#92). agent.Durableand theDoandHistorymethods ofMemStoreand the SQL stores are transitional: they go through a Journal over the store, so existing code keeps working.*JournalimplementsDurable.agent/durabletestisagent/storetestunder its former name (#92).MemStorehonors its context, and lists runs in order (#92).- A live tool result, a claim and a completion are written with one Insert, without a read first (#92).
- A
Stepthat loses its claim to a call of the step in flight in this process, and finds that call failed, halts withHaltContended(its claimant was live); aStepthat loses its claim with no call in flight still halts withHaltCrashed.ResolveHaltReffinds aLeaserthrough aJournaland through store wrappers, so its live-driver check uses the lease onstore/sqliteas onMemStoreandstore/postgres(#92). ResolveHaltRef(and its wrappers) claims the attempt after the live one before it records the outcome, on a store that leases runs as on one that does not, and returns*HaltInFlight(whose newAttemptfield names that attempt) if a driver holds it already, so a resolution cannot override a driver that revived a remembered claim and ran the effect: afterWithMinHaltAge, or under the lease, which a plainRundoes not hold. If recording the outcome then fails, the resolution's claim stays live, and the operation halts until it is resolved again (#92).- A
SagaAbortederror lists its uncompensated writes without saying each lacked a compensator: the list also holds a call whose outcome is unknown and a call whose tool is gone (#92). - The DST and reference-model crash suites inject their crashes at the storage port, under a Journal, so they cover the journal header, claims and not-started records;
agentalso carries an exhaustive fault-schedule exploration of the claim protocol, bounded by default (BIDE_EXPLORE=1runs the full exploration) (#92).
Pauses and halt resolution
- Breaking: the pause types are renamed:
PendingApprovaltoApprovalPending,InterruptedtoInterruptPending,AwaitingtoSignalPending,SleepingtoTimerPendingandResumeHalttoOutcomeUnknown. The old names stay as deprecated aliases until the 1.0 rewrite, so code that names the types or matches them witherrors.Askeeps compiling, but composite literals and the removed fields break: each type embedsRunRef(RunID,RootRunID), so a composite literal namesRunRef;Interrupted.KeyisInterruptPending.Name,ResumeHalt.ToolUseIDandToolNameareOutcomeUnknown.Op.IDandOp.ToolName, and the never-setAwaiting.Promptis gone (#90). - Breaking:
Waker.Schedule(ctx, Wake) error. A failed schedule fails the run with an error wrappingErrStorageand records nothing for the sleeping call (it used to pause with no wake registered), andRecoverLoopschedules it again on its next pass.MemWakerkeys a wake by itsRunIDandNameand resumes itsRootRunID(#90). OutcomeUnknown.Causesays why a run halted:HaltContendedwhen another driver won the claim during this drive, otherwiseHaltCrashed(no live claimant known to the halting driver, which is not proof). Breaking:ResolveHaltRef, and theResolveHaltandResolveStepHaltwrappers, refuse to resolve an effect a driver may still be running: on a store that leases runs they hold the root run's lease while resolving and return*HaltInFlightwhile a driver holds it; on a store that cannot (a custom store with noLeaser) they requireWithMinHaltAge(measured from the live attempt), unlessWithoutLiveDriverCheckis given; aHaltContendedhalt always requiresWithMinHaltAge. A resolution that conflicts with an outcome already recorded returns*HaltAlreadyResolvedinstead of nil (#90).ResolveHalt,ResolveStepHalt,Resumeand the channelSendare deprecated wrappers ofResolveHaltRef,AnswerInterruptandEnqueue(ApproveAs, the wrapper ofSubmitDecision, is removed below);ResolveHaltandResolveStepHalttake the cause asHaltCrashed(#90).- The loop,
Recover,RecoverLoopandMemWakerdetect pauses withIsPause(#90).
Flows
- Breaking:
planlowers every node ontoagent.Stepthrough the engine's step hook: a node runs as the Step named by its node key,node:<name>(node:iter:<n>:<name>in a loop body), so it has the Step's claim protocol (a fresh claim id, the not-started record, numbered re-attempts) and pause guard. A node that halts returns*agent.OutcomeUnknownwithOp: agent.OpRef{Kind: agent.OpStep, ID: "node:<name>"}, a pause thatRecoverLoopdoes not report as a failure, andagent.ResolveHaltRefclears it with the node's output. A node cancelled after its claim and before its body is re-attempted on the next drive instead of halting. A node's body that returns a pause from a node that is not retry-safe isErrConfig. A node reads the journal by point reads, never by loading the run.plan.HaltAmbiguousis removed (#103). - Breaking: the keys a flow writes change: a node's result is
node:<name>(was<name>), its attempt marker the Step'sattempt:step:node:<name>(wasattempt:<name>, aStepValuerecording{"retry_safe":...}, which is no longer written or read: the Step marker is itself the attempt's recorded safety), and a loop iteration's keys arenode:iter:<n>:<name>(wasiter:<n>:<name>). A retry-safe node writes no marker, so a node attempted as retry-safe and relabelled a side effect since runs again under a claim, as aStepdoes; a node attempted as a side effect still halts if it is relabelled retry-safe (#103). - Breaking: a flow's
Runrecordsrun:startbeforeflow:digest; resuming a flow run with an input whose JSON differs, under another flow's name, or driving it with anAgent(or a flow over an agent's run) isErrConfig, and records nothing.Conformrequiresrun:startto name the flow and recognizes the Step markers; it ignores the journal header and not-started records (#103). - Breaking: a flow's run records
run:completewith the flow's name and its output when its terminal node finishes, soRecoverandRecoverLoopskip it (they re-drove every finished flow run on every pass) andagent.IsCompletereports it; a later drive with the run's input returns the recorded output with one point read, whatever the flow's topology is now. AnAgentrefuses a finished flow run, and a flow a finished agent run (#103). - Breaking: a flow's input is compared with the recorded one as canonical JSON (keys sorted, each number as the exact decimal value it denotes, with no rounding through a float, so
1,1.0and1e0match and 2^53 and 2^53+1 do not), so decodingRecordedStart'sInputwithout loss and passing it toRunresumes the run (#103). - Breaking: a flow input whose JSON has a repeated object key, an escaped lone surrogate or invalid UTF-8 is refused with
ErrConfig, since two different texts would compare as one input; numbers compare exactly whatever the size of their exponent (#103). - Breaking: a
Step's value, and every other value the engine journals (an interrupt answer, a signal, a channel message, a timer's wake time, anAwaitFordeadline and outcome, a session's turn records, a saga's failure text, a resolution's evidence), is journaled without HTML escapes (<, not\u003c), as the journal encoding andResolveHaltRefwrite values, andResolveHaltRefcompares an outcome with the recorded one as JSON values (whitespace, key order, number spelling and string escaping aside), so resolving the recorded outcome again is not a conflict whichever way it was escaped (#103). - Breaking:
agent.Step, aParalleltask and a Step's resolution refuse an empty step name withErrConfig, inside a flow node or not (#103). - Breaking:
ResolveHaltRef(andResolveHalt,ResolveStepHalt) refuse an operation with no live attempt marker, one that never halted, withErrNoLiveAttempt, so a resolution cannot record an outcome for an effect that never ran (#103). - Breaking: an
agent.Steporagent.Paralleltask a flow node's body runs for the flow's run is recorded under the node's key (node:<node>:step:<name>, and per iteration in a loop body): a loop body's Step used to replay its first iteration's result in every later iteration, so its effect ran once (#103). - Breaking:
Conformreplays the run's routing from its recorded choices: a record of a node on an untaken arm, of a loop iteration the run did not reach, or of a node under the wrong kind of key (an iteration outside a loop, or a loop body node outside one) is a divergence, and so is a journal that records nodes withoutrun:startandflow:digest, a completed journal that lacks a node, choice or loop iteration its routing reaches, and arun:completewhose output is not the recorded output of the terminal node the run reached (compared as JSON values, so HTML escaping does not matter); a Step a node's body ran is recognized as the node's (#103). - Breaking:
Flow.Runrefuses (ErrConfig) the run IDsagent.Runrefuses: an empty one, or one in the form reserved for sub-agents and session turns (#103). - A loop whose body is one node (the head is the switched node) feeds the head its forward predecessor's value; it used to get a nil input (#103).
- Breaking:
node:,switch:andflow:are reserved step-name prefixes, so aStepinside a flow body cannot collide with a flow's records.ResolveHaltRef(andResolveStepHalt) accept a flow node's key as a step name (#103).
Model calls and spend
- Breaking:
agent.ModelHandlerisfunc(ctx, ModelCall) (ModelResponse, error), andMiddlewarewraps it. A middleware passes on the call it received or a copy with fields changed; the model handler refuses aModelCallbuilt from scratch withErrConfig, since it would drop the hooks outer middleware added.Request.MessagesandRequest.Toolsarrive clipped at every handler, so an append in one hedged branch never writes into another's backing array (#104). - Breaking: hooks are added with
ModelCall.AddHook(append-only) and have the signaturesBefore(ctx, ModelCall) errorandAfter(ctx, ModelCall, ModelAttempt); a hook whoseBeforereturned nil gets exactly oneAfter, also when a later hook'sBeforerefuses the request. The run's spend meter is not a hook, so no middleware can hide a request fromWithTokenBudgetorResult.Spend. Requests are numbered per turn across every retried attempt and hedged target (ModelCall.Attempt); a request is numbered when it reaches the model handler, before its Before hooks, so a request a Before hook refused keeps its number (#104). - Breaking: the stream sink is claimed by one request per turn at a time. A failed claimer releases it and the next claimer emits
TurnRestarted; when the turn's response is not the one that streamed (a hedge target that was not the claimer, a fallback, a cache), the agent emitsTurnRestartedand replays the response; nothing a request sends after its turn returned reaches the caller. WithHedge, the first target to start streams live, where before a hedged turn was delivered only once the winner was chosen (#104). - Breaking: an empty finish reason from a custom
Model, or in a response a middleware built, is recorded asFinishStop; a response a middleware built withFinishLength,FinishFiltered, an unknown reason, orFinishToolUsewithout a call is the error the model's own would be. The loop still decides whether to run tools from the message's calls, never from the reason (#104). agent.Replayends each turn with the finish reason and raw reason its record journaled, and reports the turn's discarded spend inFinish.Discarded, where it used to reach the run through the context;middleware.Coston a replaying agent counts it too (#104).middleware.Hedgeisc := call; c.Model = backupand has no streaming code;Retryloopsnext(ctx, call);RateLimitandCostadd hooks, and count requests outside an agent only throughagent.CallModel;trace.Modelnames the provider and model fromagent.ModelInfoOf(call.Model);WithRetrievaljournals through the call's run, not the context (#104).middleware.Costcounts each call's answer once, the response the agent records, wherever it sits: inside aHedgea losing target's response is no longer counted as an answer (#104).agent.Replayjournals the replayed turn with the model its original record named (Record.Model), not the replaying Model, also through a Model that wraps the replay model and forwards its events; late spend journaled before any turn is reported with the first (#104).- A model turn whose record write reports an error is settled from the journal: a record that landed is the turn's (its answer functions run, and the rest of the spend is late), one that did not is a failed call's
@spend/<id>, and when the journal cannot be read the spend and the answer functions are kept by the process for the run's next drive in it, which decides from the journal; a spend record whose write failed is kept and written by that drive too. Nothing is counted twice (#104). - The engine tells its own model record from another driver's by the salt the journal stamps (it draws that salt itself;
JournalEntrykeeps it), not by comparing content, so a tool call whose arguments are not compact JSON, text with invalid UTF-8, or two drivers' records equal in every field no longer make it take its own record for another's (#104). - Spend kept for a run's next drive is keyed by the store's identity, so any Journal over the store in the process sees it, and bounded at 4096 entries, oldest runs dropped first. It follows a Durable wrapper's
Unwrap() Durable(asaudit.AuditedStoreoffers) to the store beneath it, so the run's next drive through another wrapper or Journal over the same store sees it; the list of runs holding kept spend never outgrows the entries (#104). - A claim through a Durable the engine drives by its
Do(a wrapper such asaudit.AuditedStore) now behaves as through a Journal: a claim whose marker write fails records that its attempt did not start, and if that record cannot be written either the process remembers it, under the identity of the store beneath the wrapper, so the next claim or resume in the process writes it and re-attempts the effect instead of halting over one that never ran (#104). - A
Stepdriven through a Durable wrapper (such asaudit.AuditedStore) over a Journal follows the Journal's loser rule: a driver that lost the step's claim only joins a call of the step in flight and otherwise reads the result, and never starts a call, so the claim's winner in the same process cannot take the loser's halt as its own and halt (the claim model'sWinnerNeverHalts); a joined call that fails halts withHaltContended. (#104). Unwrap() Durableis a contract, asUnwrap() Storeis: only a Durable wrapper that passes run IDs and step names through unchanged may implement it, since the engine follows it for capabilities and for the per-run state the process keeps; a key-rewriting wrapper must not.storetest.CheckDurableWrapperchecks a Durable wrapper against it (#104).- The
DurableandStorecontracts state what the engine's accounting relies on: a record is written at most once under a name,fnis called at most once per record, and every record's bytes, its salt among them, are kept exactly (#104). - Breaking:
agent.Finishhas an unexported field (the replayed turn's marker), so an unkeyed composite literal of it outside the package no longer compiles; use field names (#104). - Breaking: spend records are keyed by a fresh id,
@spend/<id>(a replayed run keeps the original's), so two drivers of one run never write their spend under one key (#104). - Breaking: a request of a model turn that is over (its call already returned, so nothing would record it) is refused with
ErrConfig(#104).
Proofs and signatures
- Breaking: proofs commit to raw stored bytes. A journal leaf is its tag followed by the bytes the journal stores for the record (
Record.Raw), never a re-encoding, so a record written by a later release with fields this one does not know verifies.ProofBundle.Recordis replaced byRecordBytes(record_bytes, base64) and aRecord()method that decodes them leniently for display and role checks only;EvidencePackageandCurrentGrantProof.Leafcarry record bytes through their bundles.VerifyInclusiontakes the record bytes. A record built in memory has no leaf (#105). - Breaking: new artifact formats, and before 1.0 a verifier reads only the current one (an older artifact is
ErrFormat; re-create it from the journal):bide.audit.proof.v3,bide.audit.absence.v3,bide.audit.runcert.v3,bide.audit.current-grant.v3,bide.audit.event-inclusion.v3,bide.audit.evidence.v5,bide.audit.sth.v5,bide.audit.anchor-entry.v1. Grant canonical bytes arebide.audit.grant.v2(tag included in the signed and digested bytes), and event leavesbide.audit.event-leaf.v3.audit.UnmarshalStrictchecks the format of every artifact it meets, including one carried inside another, and wraps any other decoding error inErrMalformed(#105). - Breaking: STH v5: a
SignedTreeHeadcarriesformat(bide.audit.sth.v5) and always itsalg, and the scheme is part of the signed bytes.TreeHead.TimestampisTimestampNanos(timestamp_nanos). Each component of a hybrid signature signs its own label (bide.hybrid.ed25519.v1,bide.hybrid.mldsa65.v1) before the message, so a stripped ed25519 half verifies neither as hybrid nor as plain ed25519 (#105). - Breaking: every ed25519-only entry point takes a
SignerorVerifier:SignTreeHead(th, Signer) (SignedTreeHead, error)(SignTreeHeadWithand everyVerifyWithare gone),Sign(head, Signer) ([]byte, error),VerifySignature(head, sig, Verifier) error,SignAbsenceRoot(..., Signer, ts),Evidence(ctx, store, runID, Signer, ts, ...),EvidencePackage.Seal(Signer)andVerify(Verifier, ...),CertifyRun(ctx, store, runID, sth, RunCertSpec)withRunCertSpec.SignerandTimestampNanos(the signer must hold the key that signedsth),VerifyRun(cert, approved, Verifier),VerifyApprovals(..., Verifier),VerifyCurrentGrant(..., Verifier), andNewAuditedStore(inner, Signer, anchor) (*AuditedStore, error)(it no longer panics on a bad key). AnEvidencePackagenames its key as{alg, public_key}instead ofpublic_key_hex, andVerifyrefuses a verifier of another scheme or key (#105). - Breaking: verifiers return
error(nil means verified):SignedTreeHead.Verify,ProofBundle.Verify,AbsenceBundle.Verify,VerifyInclusion,VerifyConsistency,VerifyAbsence,VerifyEventInclusion,VerifyAnchorInclusion,SignedGrant.Verify,VerifyDelegationChain,VerifyCurrentGrant. The report verifiers (EvidencePackage.Verify,VerifyRun,VerifyApprovals) return their report and an error wrappingErrNotVerifiedwhen it is not OK.CheckTimestampandCheckTimestampOrdererrors wrapErrNotVerified(#105). - Breaking:
agent.ApproverVerifierhasAlg(),SubmitDecisionrequiresDecision.Alg, and the m-of-n gate (andTallyApprovals, soVerifyApprovals) counts a decision only when its journaled scheme is its approver key's. The deprecatedApproveAsis removed; useSubmitDecision(#105). - Breaking:
Grant.NotAfterisNotAfterUnix(not_after_unix), andGrant.Expiredtakes Unix seconds (#105). - Breaking: the stream event types (
TurnStarted,ModelEvent,AssistantTurn,ToolStarted,ToolCompleted,ApprovalRequired,Finished,TurnRestarted) and the model events (TextDelta,ReasoningDelta,ToolCallDelta,Finish) have snake_case JSON tags, and event leaves name their kind in snake_case (turn_started,model_event/text_delta); an event of an unknown type is refused (#105). - Breaking:
audit.PoliciesUsed(records) ([]string, error)andaudit.AbsenceRoot(records, set) (root []byte, size int, err error)return an error (audit.ErrRedacted,audit.ErrMalformed) for a journal they cannot project, instead of a set that silently omits a redacted record, so an auditor's recomputation cannot confirm a head that omits what the journal commits.audit.KeyFuncreturns the record's keys ([]string, nil for none), so one model turn adds a key per call it requested;ToolUseKeyandPolicyUsedKeyfollow (#105). - Producers and verifiers of proofs read every record by its stored bytes (#105): producers that project a journal (the absence key sets,
CertifyRun's used-policy set,EventLogFromJournalandPersistJournal) refuse one holding a redacted record with the newaudit.ErrRedacted, instead of omitting it (an absence proof of a redacted call, a certificate missing a redacted action's policy); a proof'srecord_bytesmust read one way to every JSON reader (no duplicate or case-variant names, invalid UTF-8 or escaped lone surrogates; unknown fields tolerated) or they areaudit.ErrMalformed; the tool-use key set also holds every call a model turn requested, attempt markers and saga failures, so a call that was started cannot be proven absent, a retry-safe call whose result was lost included; every producer and recomputation of a key set (NewAbsenceTreeHead,SignAbsenceRoot,ProveAbsent,ProveAbsentBundle,AbsenceRoot,PoliciesUsed,CertifyRun) checks each record's stored bytes as a proof does, and so doEventLogFromJournalandPersistJournal, so a record that reads two ways isaudit.ErrMalformedrather than projected one way; each projects what a record's stored bytes (Record.Raw, which the journal tree binds) decode to, and refuses a record whose fields do not say what its stored bytes say (audit.ErrMalformed) (event leavesPersistJournalwrote before a redaction keep the redacted content: redact the event trail too);VerifyAnchorInclusion,VerifyRunandEvidencePackage.Verifycheck the format of every artifact they carry, at any depth, on a Go value asUnmarshalStrictdoes on JSON (a run certificate's heads included, when the certificate is for another run andVerifyRunis never reached);VerifyApprovalsandApprovalEvidenceread the journaled tally withUnmarshalStrict, so a tally that reads two ways isaudit.ErrMalformed, and the gate reads a recorded tally by the same rule (a tally it cannot read strictly isErrStorage, and the tool does not run);bide-audit verify-approvalsrefuses an approver-key file that gives two ids one key, compared by key identity (KeyIDs), whether or not both ids are in the policy (exit 4). bide-audit verify-approvalstakes-approved/-approved-file, the allowlistverify-evidencetakes, so a package that carries a run certificate can pass (#105).- Breaking:
bide-auditreads keys as<alg>:<hex>(or bare ed25519 hex) for-pubkeyand the approver keys, reads aJournalExportfor-journal, and maps the audit sentinels to its exit codes:ErrNotVerifiedis 1,ErrFormatandErrMalformedare 4 (#105).
Approval key identity
- Breaking:
agent.ApproverVerifierhas a second method,KeyIDs() []string: the identities of the signing keys behindVerify, derived from the public key's bytes (one per key; a hybrid reports each component). A custom verifier must implement it. An approver whose verifier reports no key identity, including anaudit.Ed25519Verifierover a wrong-length key, used to count as unable to sign and is nowErrConfigat the gate (#109).
Removed
- Breaking:
agent.DetachModelSink,agent.EmitMessage,agent.WithModel(ctx, m)andagent.WithModelCallHook(ctx, h): the model call path carries no engine data in the context. UseModelCall.Modelto retarget a call andModelCall.AddHookto add a hook (#104). - Breaking:
middleware.CostMeter.Total,Usage,SpentandSpentTotal; useSnapshot(#104). - Breaking:
trace.WithSystemandtrace.WithModel; the chat span reads the provider and model from the call'sModel(agent.Describer) (#104).
Fixed
store/postgressends each lease call (AcquireLease,RenewLease,ReleaseLease) and eachInsertas one statement that Postgres commits before it replies, instead of a transaction ofBEGIN, the statement andCOMMITin separate round trips. A holder stalled between its statement and itsCOMMIT(a SIGSTOP, a suspended VM, a long GC pause) kept the lease row, or the run's insert lock, locked, so another node'sAcquireLeaseof the expired lease, or its next insert into the run, blocked for as long as the stall lasted instead of taking the run over one TTL after the last committed renewal. An insert now takes its position from<prefix>next_seq_v1(run_id), a VOLATILE plpgsql function thatOpencreates when it is missing: called inside the insert's one statement, it takes the run's transaction-level advisory lock (the key earlier versions used, so both queue on one lock during an upgrade) and then readsMAX(seq)+1with a snapshot taken after the lock, so inserts into one run queue instead of racing for a position.Opennever replaces the function and refuses one with another definition. A statement that a repeatable read or serializable deployment fails with a serialization failure (40001) is run again, reads included, so the default isolation still does not change the outcome. Retries wait a capped, jittered exponential backoff (1ms up to 100ms) and continue until the context ends, so a caller that needs a bound on a call sets a deadline.Openrefuses, withErrConfig, existing tables that lack a uniqueness the statements depend on ((run_id, seq)and(run_id, name)on the journal,run_idon the leases), since it never alters an existing table. The schema migration, the one remaining multi-statement transaction, setsidle_in_transaction_session_timeoutso a node stalled inside it cannot hold the migration lock (#113).govern/postgreslog.Appendis one statement for the same reason: a process stalled inside an append held the entity's advisory lock, and every other append to the entity waited for as long as the stall lasted. It takes its position fromgoverned_events_next_seq_v1(entity)the same way, its retries back off the same way, andOpenrefuses agoverned_eventstable without unique(entity, seq)and(entity, append_id)indexes, or a next_seq function with another definition (#113).store/postgresandgovern/postgreslogopen in a schema whose name has an upper-case letter: the next_seq lookup castcurrent_schema()toregnamespace, which folded the name to lower case. The lookups now readpg_catalog's tables by name with the statement's snapshot, never throughto_regclassorto_regprocedure(#120).Openin both packages uses one schema for every statement: the oneWithSchema(new instore/postgresand, as an option ofgovern/postgreslog.Open, in the log) pins, or else the first schema on the search path holding the store's steps table (orgoverned_events), else the first schema on the path, with a warning logged throughlog/slog, since discovery runs again at everyOpenand a role that can create a schema earlier on the path can redirect a restarting node.WithSchemarefusesinformation_schemaand everypg_name, and a pinned schema that does not exist failsOpenwithErrConfig;Opennever creates a schema. Every table reference, the next_seq call and the migration's DDL name that schema. Agoverned_eventstable in a later schema than an empty one is now migrated in place: the migration's unqualifiedCREATE TABLE IF NOT EXISTSused to create a second, empty table in the first schema and move the log to it.Openrefuses, withErrConfig, a first relation of that name that is not an ordinary or partitioned table: a view namedgoverned_eventsbefore the table had appends go through it to another table (#120).Openin both packages refuses a next_seq function owned by another role than the table's owner, or one withoutSET search_path = pg_catalog, pg_temp, withErrConfig. The statement check now tokenizes every statement and holds it to allowlists: calls only to listedpg_catalogbuilt-ins (a function named like an unreserved keyword,conflict, counts as a call) or to next_seq under its exact schema-qualified name, once, in a one-rowINSERTinto its own table; operators only asOPERATOR(pg_catalog.<op>)from a list, and no keyword operator orARRAYconstructor; relations only the store's own tables orpg_catalog's; casts only to listedpg_catalogtypes (a domain'sCHECKcan call anything); ASCII only outside literals; no comments, dollar quotes, backslashes, typed literals or semicolons (#120).RecoverandRecoverLoopcallresumeonly for a run that is still unfinished: holding the run's lease, the pass checks the terminal markers (run:complete,run:aborted,run:cancelled) again before it callsresume, so a run another driver finished after the pass listed it is left alone and not counted as re-driven (nor is a run whose check failed, which is reported). The check covers every driver that holds the run's lease; a finish by a driver with no lease, or during a stall past the lease TTL between the check andresume, can still reachresume, which replays the finished run. Budget: a recovery pass makes three point reads (Store.Get) for each run it drives; the finished runs the listing excludes still cost nothing (#114).- Tests: the multi-process HA harness in
store/postgresgives each cluster a schema of its own. Its workers' recovery passes walked every incomplete run the package's other tests left in the shared schema, so taking over a stalled worker's run slowed with that count rather than the lease TTL, andTestHA_MultiProcessStallPastTTLtimed out in the Integration job's second pass. Its wait for the take-over is now derived from the lease TTL, the recover interval and one drive (#122).
Security
- An m-of-n approval gate counts one seat per signing key. Two approvers whose verifiers resolve to one key let that key's holder meet the quorum alone; the gate now refuses such a policy with
ErrConfigon every evaluation,TallyApprovalsnever counts either approver,audit.VerifyApprovalsreturns an error, andbide-audit verify-approvalsexits 4. Found by the TLA+ approvals model (finding F5) (#109). Upgrading: the check covers tallies this version counts. A terminal tally already in a journal is reused, not recounted, so a passed tally recorded by an earlier version stands even if two of its approvers shared a key, and its tool runs when the run resumes. Journals are not promised to resume across pre-releases; before upgrading, finish the runs paused on an m-of-n gate, or audit each one that holds a recorded tally (audit.VerifyApprovalsunder the new rules refuses a shared key) and resolve it by hand if two approvers shared a key. - Weak Ed25519 public keys are refused.
crypto/ed25519accepts keys that are not canonically encoded, small order, or mixed order: under a small-order key (such as any of the identity point's four accepted encodings) anyone can forge a signature for any message, and a mixed-order keyA + Tis a second public key forA's secret, which gave one secret two approval seats.audit.Ed25519Verifier(and soHybridVerifier),audit.VerifySignatureandaudit/verify.TreeHeadnow verify nothing under such a key,Ed25519Verifier.KeyIDsreports no identity for it (so the approval gate refuses it withErrConfig), andbide-auditrefuses it as-pubkeyor in-approver-keys(exit 4). The check costs 1 to 4 ms of CPU per new key; results are cached (a 1,024-key LRU, single-flight), so an application that resolves Ed25519 keys from untrusted input should bound or rate-limit those lookups. Found in the adversarial review of #109. store/postgresandgovern/postgreslog, in v0.8.0 and earlier as well: three kinds of role could run their code with the store's privileges, and some could choose the run's lock key. A role withCREATEon the database could create a schema named like the store's role, which the default search path ("$user", public) puts first, at any time afterOpen, with anext_seqfunction (or, for v0.8.0, tables) that the store's unqualified names then resolved to. A role that can create functions in any schema on the search path, even one after the store's own (publicon Postgres 14 and earlier lets every role create there), could definehashtextextended(text, integer), which matched the lock callhashtextextended(r, 0)better thanpg_catalog's. And a role that can create operators in any schema on the path could define=fortext[]or foroidandregtype, whichOpen's catalog checks compared and whichpg_cataloghas no exact match for, so the operator ran at everyOpen,pg_catalogsearched first or not. Now every name in every statement is qualified: the tables and next_seq with the store's schema, and every function, type, collation and operator withpg_catalog(operators asOPERATOR(pg_catalog.<op>)); the statement check refuses anything else, keyword operators (LIKE,IS DISTINCT FROMand the like) andARRAYconstructors included; the next_seq functions qualify every name and operator in their body and also run withSET search_path = pg_catalog, pg_temp; andOpenchecks the function's owner and definition. WithWithSchema, the search path plays no part in which schema the store uses; without it, discovery still trusts every schema on the path, so pin the schema. The lock key's value is unchanged, so v0.8.0 nodes still queue on it during an upgrade. The stores trust the owner of their schema and every role that can create objects in it, as they trust the tables' owner; any role that can connect can still stall a run's inserts by holding its advisory key, as before (#120).
0.8.0 - 2026-09-30
Added
agent.RecoverLoopwithWithRecoverInterval,WithRecoverConcurrencyandWithRecoverErrors: runs recovery passes until its context ends, so a dead holder's runs are taken over without another call (#58).agent.ErrLeaseLost: a lost lease cancels the drive with a cause that wraps it; failed renewals retry every ttl/20 and give up at 3/4 of the TTL (#58).agent.ErrToolOutcomeUnknown: a call lost mid-flight on a tool that is not retry-safe records no result and halts on resume (#57).mcp.WithSafety,mcp.WithCallTimeout,mcp.WithMaxResultBytes,mcp.WithMaxDescriptionBytes, with defaultsDefaultMaxResultBytes(1 MiB) andDefaultMaxDescriptionBytes(8 KiB) and the errormcp.ErrResultTooLarge(#57).modeltest.ToolNamesconformance check for provider tool-name rules (#57).agent.StepNotStarted: a tool call orStepcancelled after winning its claim but before its effect ran is recorded as not started and re-attempted, instead of halting (#65).agent.Result.Spend,agent.Record.DiscardedUsage,middleware.CostMeter.SpentandSpentTotal: usage of retried, hedged and failed requests is journaled and reported (#55).agent.TurnRestartedstream event, sent before a retried attempt's output when the previous attempt streamed any (#55).agent.ModelCallHook,agent.WithModelCallHookandagent.WithModelfor hooks around every request and per-call model routing (#55).agent.ErrNegativeUsageandagent.Usage.Validate(#55).agent.ToolErrorTextandagent.RedactURLs, so middleware can record exactly the error text the journal holds (#54).agent.Capability[T]finds an optional store capability through wrappers;audit.AuditedStoreimplementsUnwrap(#66).agent.ResolveStepHalt,agent.IsReservedStepName,agent.SubRunIDandagent.ToolResultStep(#60).agent.FinishStop,FinishToolUse,FinishLength,FinishFiltered,ErrOutputTruncatedandErrOutputFiltered(#68).provider.DefaultMaxResponseBytes(32 MiB),provider.LimitResponseandWithMaxResponseByteson each model adapter (#68, #76).agent.Describer,agent.ModelInfoandagent.ModelInfoOf, implemented by the Anthropic, OpenAI and Gemini adapters (#76).Finish.Rawcarries the provider's own finish reason, andFinish.Discardedthe usage of discarded responses (#76).modeltest.CheckFinish(#76).agent.DecodeStoredRecordforDurableimplementations to read rows back (#68).audit.CheckTimestamp,CheckTimestampOrder,DefaultClockSkew,WithVerifyTime,WithClockSkew, andbide-audit -max-clock-skewand-max-input-bytes(#68).audit.PolicyLeafName,ConvergenceLeafName,GovernedPolicyDigestandaudit.ErrFormat(#66, #68).plan.BlockName,Arm.Named,Flow.DigestV1and the config safety valueside_effect(#68).RetrievalTooloptionsRetrievalNameandRetrievalDescription, so one agent can search several stores (#62).eval.ReportFormat,eval.ErrFormat,eval.MetricDirectionandRunOutput.TraceErr(#74).- A multi-process HA test harness on Postgres that kills, stalls and restarts workers (#58).
bide-audit -jsonand-version(#79).eval.CompareoptionsWithUnscoredToleranceandWithCaseSetMismatchAllowed,Comparison.Gate,ErrRegression,ErrInconclusiveandDirectionInconclusive(#78).- CI enforces doc comments on every exported identifier with
internal/tools/doccheck(#73). agent.RunStartandagent.RecordedStart: a run's input and entry point are journaled at its first drive and can be read back, for example by aRecovercallback (#70).agent.Record.ReadOnly(read_only): a tool result records whether its call ran ReadOnly (#70).- The library modules (
govern,store/sqlite,store/postgres,mcp,trace,codec/gcf,govern/sqlitelog,govern/redislog,govern/postgreslog) are tagged<dir>/vX.Y.Zwith each release byscripts/release.sh, so they install withgo get(#85). - CI checks that every Go block in the README and
docs/compiles against the current code, and that API listings match the packages, withinternal/tools/docsnip(#88).
Changed
- Breaking: journal keys for tool results, attempt markers, approvals, compensations and sub-agent runs encode the tool-use id (
tool:<id>, sub-run<parent>><id>); runs journaled by v0.7.0 do not resume (#60). - Breaking:
ResolveHaltresolves only tool calls; useResolveStepHaltfor aStep(#60). - Breaking: step names with a reserved engine prefix are
ErrConfig, run and session ids may not contain>,planstep names may not contain:, andRecoverskips sub-runs (#60). - Breaking:
Finish.Reasonuses the neutral valuesstop,tool_use,lengthandfiltered; a cut-off or filtered turn fails withErrOutputTruncatedorErrOutputFiltered, and any other reason withErrStreamProtocol(#68). - Breaking: a reply over the response cap fails with
ErrResponseTooLarge, and atool_useturn with no call fails (#68). - Breaking: tool arguments for
Func,SubAgentandfinal_answerdecode strictly: a missing required field, an unknown or case-variant name, a duplicate or trailing data isErrToolArgs(#61). - Breaking:
schema.Fordescribes whatencoding/jsondecodes (Unmarshaler,TextUnmarshaler, integer bounds) and returnsErrUnsupportedTypefor chan, func, complex and similar kinds (#61). - Breaking:
RunTyped[T]returnsErrConfigwhenTis not an object, andRunTypedNativeon Anthropic returnsErrConfig(#61). - Breaking:
WithRetrievalsends context as a user message just before the latest user turn, one JSON object per document, journaled once per run as@retrieval/<layer>;RetrievalToolandWithRetrievalpanic for k < 1 (#62). - Breaking:
WithTokenBudgeton a parent covers its sub-agents, andResult.UsageandResult.Spendreport the whole run, including earlier invocations and sub-agents (#69). - Breaking:
agent.EmitMessagetakes the turn'sUsage(#52). - Breaking:
planstep and join functions (Builder.Step,Join2,Join3,RegisterStep,RegisterJoin2,RegisterJoin3) take acontext.Context(#66). - Breaking: audit proof JSON uses snake_case names and a
formatfield (bide.audit.proof.v2and siblings,bide.audit.evidence.v4); verifiers reject other formats withaudit.ErrFormat(#66). - Breaking:
plandigests move tobide.plan.topology.v2, so flows started under v1 do not resume; config JSON decodes strictly, and a configsafetycan lower retry safety but not raise it or clear approval (#68). - Breaking:
planconfigs must state"version": 1(plan.ConfigVersion) and use snake_case keys (loop_max);TopologyJSON keys are snake_case andTopologyNode.KindisTopologyNodeKind; config parse errors wrapErrConfig(#75). - Breaking: signed tree heads need a positive nanosecond timestamp within the clock skew, empty evidence packages fail, and
bide-auditinput is capped at 256 MiB (#68). - Breaking:
ApprovalPolicy.Validaterefuses approver ids that are equal under Unicode normalization and case folding, and invalid UTF-8 (#68). - Breaking:
middleware.RetryandToolRetrywith n < 0 returnErrConfig;Quorumwith k above the voter count isErrConfig;EventLog.EventsreturnsErrProtocolon a gap in positions (#68). - Breaking:
eval.Judgepasses only on an exactPASS, and a failed run never passes (#68). - Breaking:
eval.RequiredRunsreturns(int, error), andeval.Runreturns an error for duplicate metric names (#53). - Breaking:
eval.Matchestakes a*regexp.Regexp,eval.AgentRunnerreturns(RunFunc, error),eval.Comparereturns(Comparison, error), andGovernanceHeldpredicates take a context (#74). - Breaking: the provider HTTP kit moves from
agentto the new packagemodel/provider(ClassifyHTTPError,ClassifyStreamError,NewSSEScanner,ParseRetryAfter,ToolResultCodec,JSONToolResultCodec,EncodeToolResultOrand others);WithToolResultCodectakes aprovider.ToolResultCodec(#76). - Breaking:
Finish.Reasonis the typedagent.FinishReason(#76). - Breaking:
evalMetric.Fnreturns(bool, error); a metric that cannot score a run leaves it unscored,MetricStatreportsScoredandUnscored, and reports arebide.eval.report.v2(#78). - Breaking:
bide-auditexit statuses follow one scheme: 0 verified, 1 not verified, 2 usage, 3 no verdict, 4 input unreadable or unusable (#79). - Breaking: a session's journal and turn runs move to
"<id>>@session","<id>>@turn/<n>"and"<id>>@event/<encoded key>"(from"<id>","<id>/t<n>"and"<id>/e/<key>"), so no run ID passed toRuncan name one:Run("chat/t0")and session"chat"'s first turn shared a journal, and whichever finished first handed the other its answer. A session journaled before this change opens empty. Session ids may contain/(#56 had refused it) and anySendOncekey is allowed;agent.IsSessionRunreports a session's run IDs, andRecoverskips them, since the session resumes a turn when its message is sent again (#86). - Breaking: every run journals
run:start(its input and whether it is a saga); resuming an unfinished run with another input, or throughRunfor a saga (orRunSagafor a run), isErrConfig, and so isSendOncewith a different input on a key whose turn is still open (#70). - Breaking:
planattempt markers record whether the node was retry-safe; a node re-runs on resume only if it was retry-safe when attempted and is now, and a marker written before this halts (#70). - Breaking: a recorded approval denial is final even if the tool's gate is later removed, loosened or made m-of-n (#70).
- Breaking: saga rollback treats a completed call as a write unless its result records that it ran ReadOnly, and reports calls whose tool is no longer registered as uncompensated (or halts on one attempted with no result) (#70).
WithTokenBudgetcounts discarded and failed requests;RateLimitandCostcount every request sent;Hedgeruns each backup through the inner middleware chain;Retryableno longer retriesErrTruncatedToolArgs(#55).- A resume halts on any attempt marker without a result, whatever the tool's current safety; a tool no longer registered halts instead of failing with
ErrUnknownTool(#57). mcp.ToolsreturnsErrProtocolfor malformed or duplicate tool names, non-object schemas and oversized descriptions; an agent with two tools of the same name returnsErrConfig(#57).Leaseclaims, renews and releases under a per-call token (<holder>#<token>), so drivers sharing a holder name do not share a lease (#58).ToolStartedfires immediately before the tool is called (#65).- The OpenAI adapter merges consecutive user messages into one turn (#62).
bide-auditexits 2 for stray arguments, unknown flags,-h, or both-tooland-index, and 3 when a-checkergives no verdict (#54).middleware.LogErrorTextlogs the journaled, redacted error text (#54).docs/KNOWN-LIMITATIONS.mdis rewritten for users, grouped by area with impact and workaround (#49).- The approval guide is now Human approval (human-in-the-loop) (
docs/guides/hitl-approval.md); the olddocs/guides/approval.mdand its site URL point to it.
Removed
- Breaking:
plan.Retryableand the config safety spelling"retryable"; useplan.Idempotentand"idempotent"(#75).
Fixed
store/postgresandgovern/postgreslogrun every write transaction at read committed, whatever the deployment'sdefault_transaction_isolation. At repeatable read or serializable, concurrent steps of one run failed with a duplicate(run_id, seq)(23505) because each insert computed its position from a snapshot taken before it held the run's lock, so a tool result could go unrecorded and the run halt on resume; racing nodes and lease calls failed with serialization errors (40001), and concurrent governed appends failed with 23505 or 40001 (#93).- A timer (
Sleep,WaitUntil,AwaitFor) inside a sub-agent registers its wake under a name qualified by the sub-run, so two sub-agents of one root waiting on timers of the same name no longer share one wake that the later replaced, which left the first sleeping past its wake time (#91). examples/signals: the Await scene's resumed run no longer fails with a reused tool-use id; the example's scripted model now decides from the conversation (#91).- Recording a step costs about half what it did:
MemStore.Dodecodes a record it writes once, and a message part is tagged with its type without a decode and re-encode. The journal bytes do not change. This restores thecmd/benchthroughput lost since v0.7.0 (#89). - Resuming a run that crashed after its final answer returns the recorded answer without another model call (#59).
Replaykeeps redacted reasoning, empty thinking blocks and Gemini thought signatures (#51).Replaycarries each turn's recorded usage, and a finish reason derived from the turn, so a replayed run stops at the same token budget (#52).- A hedged stream's
Finishcarries the winner's usage (#52). - A response substituted by middleware or a hedge backup is checked for missing or reused tool-use ids (#55).
NewRateLimiterwith a non-positive interval does not block, and a cancelled call keeps its token (#55).- One
*Sessionshared by concurrent callers is safe, and a resumed session turn keeps the transcript it started with (#56). - An MCP result carrying only
structuredContentreaches the model (#57). - A retry-safe
Stephalts when an earlier side-effect attempt left a marker (#57). - SQLite and Postgres record a step's outcome even when the driver's context is cancelled while the step runs (#58).
LeaseandRecoverreturnErrConfigfor a non-positive TTL instead of panicking (#58).- The lease renewer stops before the lease is released (#58).
- A tool-use id containing
/,:or an engine key name no longer collides with other journal records, and aStepcannot mark a run complete by its name (#60). Recoverno longer hands sub-runs to the resume callback (#60).Lease,RecoverandRecoverLoopfind theLeaserandListerof a store wrapped inAuditedStore(#66).- Saga compensation undoes the arguments the tool accepted after middleware, including in rollback re-runs (#67).
RunTypedreturns the argumentsfinal_answeraccepted, and its text fallback decodes the final turn (#61).schema.Forterminates on recursive named maps and slices and on self-referential pointer types (#61).- Retrieved context stays in every model call of a run, including after resume, and a retrieved document cannot forge another entry (#62).
- A non-finite retrieval score is reported as 0 (#62).
- A panicking model or tool call marks its span failed without exporting the panic value (#54).
- A turn cut off at its token limit or stopped by a filter is an error, not the run's answer (#68).
- Stream parsers reject a second call at the same index, Anthropic deltas for unknown or stopped blocks, and OpenAI choices other than index 0 (#68).
- A negative or overflowing
Retry-Afteris clamped (#68). - Model text quoted in errors is bounded (#68).
- In-process step keys are length-prefixed in
MemStore,MemWakerand the SQL stores, so run ids and step names cannot collide (#68). - Event-log reads detect a missing position, and a negative
Appendposition isErrProtocol(#68). Quorumchecks that each recorded vote names its slot's voter and recomputes the tally (#68).planrefuses empty node and join names, and edges into a join other than its declared inputs (#68).- Mermaid labels are entity-encoded (#68).
WithMinHaltAgetreats a non-positive attempt timestamp as missing (#68).- Signing with an ed25519 key of the wrong length fails before anything is written (#68).
eval: a repeated tag no longer double-counts, Wilson bounds are exact at 0/n and n/n, andHashCasesreports encoding errors (#53).eval.AgentRunnerreports a failed journal read, and trajectory metrics fail such a run (#74).chaos.Verifyalso requires a completed run to fire exactly once (Report.Missed), andBide().Writes()counts the completion marker (#53).cmd/benchrejects non-positive-runsand-concurrency(#53).- The example store in the extension-points reference records through
agent.JournalEntry, so it salts every record and passesagent/durabletest(#87).
Security
bide-auditchecks that each bundle is a record of the role it is read as (#68).bide-audit -pubkeyalways reads a 32-byte hex value as the key, never as a file name (#68).audit.VerifyAnchorInclusionbinds an entry's sequence and run to its proof, andverify.TreeHeadrefuses the head shapes theauditpackage refuses (#68).EventLog.AddandVerifyEventInclusionrefuse invalid UTF-8 (#68).- One strict reader (
audit.GovernedPolicyDigest) decodes the governed-action digest for both the agent andbide-audit(#68). bide-audit verify-quorumrejects a tally that lists one voter twice (#54).- With content capture on, trace spans and
ToolLogrecord the redacted error text the journal holds (#54).
0.7.0 - 2026-09-29
Added
agent/durabletestconformance suite forDurablestores, and one journal encoding (agent.EncodeRecord,DecodeRecord) used by every store (#36).- Salted journal records (
Record.Salt,agent.JournalEntry) and salted event-log leaves (bide.audit.event-leaf.v2,agent.ProjectEvents), so a proof discloses nothing about neighbouring records (#42, #46). Agent.WithToolErrorRedactor; URL credentials are redacted from tool error text before it is journaled (#42).middleware.ErrorSummaryand theToolLogoptionLogErrorText(#42).audit.VerifyCurrentGrantandProveCurrentGrantfor a current-grant ledger (#35, #40).governApplyOnceon every governor, andagent.NextOnceKeyfor exactly-once operations inside a call (#34, #40).ErrQuotaExhausted,ErrResponseTooLarge,ErrStreamProtocol,ErrToolUseIDReusedandagent.ClassifyStreamError(#41, #44).schema.Gemini,schema.ErrStrictUnsupported,openai.WithMaxCompletionTokensandmodeltest.ToolConfig(#36, #41).- Fuzz targets for the SSE scanner, adapter stream parsing, journal encoding, strict JSON, every audit verifier, the
planconfig loader and the GCF codec (#44). govulncheckin CI (#38) and a manual Benchmark workflow on a standard runner (#45, #47).- How bide is verified (#32).
Changed
- Breaking: audit heads sign
bide.audit.sth.v4with versioned, salted leaves and commit to their tree kind and run; re-anchor heads made with earlier versions (#35, #40, #42). - Breaking: new signatures for
VerifyRun,VerifyDelegationChain(withScopeRules),SignAbsenceRootand the absence proofs,EvidencePackage(Seal,WithConsistencyFrom),ProveCurrentGrantandVerifyCurrentGrant(#35, #40). - Breaking:
EventLog.ProvereturnsEventInclusion, andVerifyEventInclusiontakes it (#46). - Breaking:
govern.EventLog.Appendtakes an append id,Governor.Applycan return an error,Quorumtakes a name, andbide-audit verify-quorumrequires-name(#34). - Breaking: Redis event-log keys are
govern:{<len>:<entity>}, and channel and quorum step names change; drain in-flight runs before upgrading (#34, #40). - Breaking: Postgres keeps its journal in a
bide_stepstable stored asbytea, with one position per step (#25, #36, #40). - Breaking:
agent.MemWaker.Startreturns a channel that closes when the loop has stopped (#24). - Breaking: a live model turn with a missing or reused tool-call id is an error, and Gemini tool-call ids are the provider's own or random (#41).
- Breaking:
schema.ForandOpenAIStrictcan return errors, strict schemas make optional fields nullable, andagent.Funcpanics on an argument typeForcannot describe (#36, #40). - Breaking:
RunTypedends at the first acceptedfinal_answer(#36). - Breaking:
plan.Buildrejects shapesRuncannot execute as declared, and approval-gated nodes (#30). - Breaking:
Stream.Eventsfails on any event after aFinish, and an Anthropic stream withoutmessage_stopis incomplete (#44). - Quota and billing exhaustion is
ErrQuotaExhaustedand not retried, andRetryableno longer retriesErrConfig(#41). - OpenAI sends
max_completion_tokenson api.openai.com and for o-series and gpt-5 models (#41). - With content capture off, trace spans and
ToolLogrecord an error's category, not its text (#42). - Signers,
AuditedStoreand the model adapters print a redacted string for every format verb (#42). audit.UnmarshalStrictuses onlyencoding/json, so bide builds withGOEXPERIMENT=nojsonv2(#37).- The SQLite store and event log wait up to 30s for another writer's lock (#29).
golang.org/x/textis v0.39.0 instore/postgresandgovern/postgreslog(#38).- Published benchmark numbers come from the Benchmark workflow and a labelled Mac run (#47).
Fixed
- A losing driver is never told it won a step when its reload fails (#25).
- Postgres keeps records byte for byte, accepts NUL and invalid UTF-8, and gives each step its own position (#25, #36).
Leasereleases on shutdown instead of holding until the TTL (#25).- A resumed run sends the model the same bytes the live run did, on every store (#36).
- A cancelled
Stepdoes not start (#24). - A pause in one parallel tool call lets its siblings finish and record their outcomes (#26).
- Saga rollback covers calls its abort cut off, including sub-agents, and reports idempotent writes without a compensator (#27).
- A pause or halt inside a sub-agent carries
RootRunID, and aSleeporAwaitForin a sub-agent wakes the root (#28, #34). Recoverskips a saga whose rollback finished (#31).planresumes only under the flow a run started with (#30).- A tie for the most votes is never quorum agreement (#33).
- An unknown governed event name is rejected before it is appended (#34).
- Governed appends and
EventToolcalls land once when retried (#34). AwaitForrecords its outcome once per call (#34).- Channel names containing
:no longer read each other's messages (#34). - Each
Quorumin a run keeps its own tally (#34). - Two applies inside one tool call are both recorded (#40).
- SQLite and Postgres event-log migrations are safe under concurrent opens, and Postgres opens take no table lock when the schema is current (#40).
- The Redis event-log script works on Redis Cluster (#40).
- A negative Merkle leaf index is rejected (#35).
- Wrong-length keys verify nothing instead of panicking (#35).
codec/gcfcarries integers above 2^53 exactly and rejects trailing data (#36, #44).schema.Foragrees withencoding/jsonon promoted fields,,string, arrays and embedded pointers (#36).RunTypedends atfinal_answerand does not fall back to unrelated prose (#36).- Gemini requests group parallel tool results, skip empty parts and carry thought signatures (#41).
- Anthropic requests keep
redacted_thinkingblocks and tools undertool_choice"none" (#41). - Empty OpenAI tool arguments are sent as
{}(#41). - An OpenAI or Gemini turn ends only on the provider's finish signal, not on a chunk carrying usage (#44).
bide-auditdecodes committed leaf content strictly (#44).
Security
- Journal heads and absence heads are signed under distinct domains, so a head for one tree cannot prove absence from another (#35).
VerifyRunrequires the used-policy head to belong to the certificate's run, with an auditor-supplied allowlist (#35).EvidencePackage.Verifychecks every field and authenticates the run id (#35).- Delegation chains enforce issuer continuity, expiry and full scope attenuation; demoted earned-authority grants stop verifying (#35, #40).
- Records and grants with invalid UTF-8 are refused, and audit bundles are decoded strictly (#35, #40).
golang.org/x/textv0.39.0 fixes GO-2026-5970, reachable fromstore/postgresandgovern/postgreslog(#38).
0.6.0 - 2026-09-28
Added
agent.StepSafetyandTask.Safetyto declare a durable step retry-safe (#18).Session.SendOnceanswers each inbound message once per key (#20).Agent.WithTokenBudget, rebuilt from the journal so it holds across resume; each model call's usage is recorded with its turn (Record.Usage) (#15).Usage.TotalInputTokensandUsage.TotalTokens(#14).agent.ToolSafety,agent.WithToolSafety,Safety.RetrySafeandErrToolReinvoked(#16).agent.NewStreamFuncandStream.Close(#12).agent.ErrIncompleteResponse(#13).model/modeltestconformance checks for adapter streaming (#12, #13).mcp.TrustAnnotations(#21).
Changed
- Breaking:
middleware.TokenBudgetis removed; useAgent.WithTokenBudget(#15). - Breaking:
agent.Stepis a side effect unlessStepSafetysays otherwise: it claims an attempt marker and halts after a crash (#18). - Breaking:
Parallelrequires unique task names (#18). - Breaking:
Usage.InputTokensis the uncached input on every provider (#14). - Breaking:
eval.Runreturns(Report, error)(#19). - Breaking:
mcp.Toolsapplies server annotations only withTrustAnnotations()(#21). ToolRetryretries only retry-safe tools,ToolCachecaches only read-only tools, and a tool that is not retry-safe runs once per call whatever the middleware does (#16).- Breaking out of
Stream.Eventscloses the stream (#12). Session.Sendrecords which message started each turn; a different message while a turn is open isErrConfig(#22).
Fixed
MemWakerretries a wake whose resume failed (#10).- A cancelled tool call records no outcome, and a cancelled run stops before its next turn (#11).
- A stream the consumer abandons releases its HTTP response (#12).
- A response that ends without a
Finishis an error, not the answer (#13). - Cached tokens are counted once on OpenAI and Gemini, so cost matches across providers (#14).
trace.Modelreports total input tokens, including cached input (#14).- A token budget applies per run and has no data race (#15).
- A cancelled parent waits for its sub-agent to finish (#17).
- A crash after a
Step's effect does not repeat it (#18). - Each eval samples the model afresh, and a cancelled eval returns an error (#19).
- A session turn answers its own message, and concurrent handles lose no turn (#22).
- A run's anchored tree heads only grow, and a failed anchor publish is retried (01a9afd).
0.5.0 - 2026-09-28
Added
agent.ClaimAttemptandRecord.Claim: an exclusive attempt claim keeps side effects at most once when two drivers overlap (f71efd6).govern/eventlogtestconformance suite forEventLogimplementations (f3fd665).- Log-backed governors gain
Sync(ctx)(f3fd665). agent.DetachModelSinkandagent.EmitMessagefor middleware that delivers a chosen response to a streaming caller (317c98b).- CI runs the Postgres and Redis integration suites (afc0fe6).
Changed
- Breaking:
EventLog.Appendreturns the event's position, andEventstakes a starting position (f3fd665). - Breaking:
Applier.ApplyreturnsApplied{State, Position}, andFederatedApplier.ApplyreturnsFedApplied(f3fd665). - Breaking: the Postgres event log uses a
governed_eventstable; write Redis event-log streams only through the adapter (f3fd665). MemStorestores records as JSON and returns independent copies, like the SQL stores (4912fbf).- With backup models,
Hedgestreams only the winning response (317c98b).
Fixed
- Governors sharing an event log fold in each other's events as they act, and each governed action records its log position (f3fd665).
- The Postgres event log keeps each entity's events in a fixed order under concurrent writers (7efece5).
- A hedged stream cannot panic the process after the run ends (317c98b).
0.4.0 - 2026-09-28
Added
- m-of-n signed approval:
Safety.ApprovalwithApprovalPolicy,ApproveAs,ApprovalDecisionBytes,Agent.WithApproverVerifiersandPendingApproval.Quorum(78f8db6, 3262cd1). WithDecisionCheckonApproveAs, withErrInvalidApprovalandErrAlreadyDecided(3262cd1).audit.ProveApproval,audit.ApprovalEvidenceandaudit.VerifyApprovals, andbide-audit verify-approvals(78f8db6, 994721b, 3262cd1).- A
planconfigapprovalblock (78f8db6). examples/approval, a cross-process walk-through (994721b).SECURITY.mdwith a private reporting contact.
Changed
- Breaking: calling
Runon a finished run id returns its recorded answer; useSessionto continue a conversation (c6deb76).
Fixed
- Re-invoking a finished run returns its recorded answer without calling the model, so a repeated tool request cannot run again (c6deb76).
0.3.0 - 2026-09-28
Added
ResolveHaltoptionsWithMinHaltAgeandWithNow, the error*HaltTooYoung, andResumeHalt.AttemptedAt(#2).ResolveHaltoptionWithEvidence, recorded asRecord.ReconciledandRecord.Evidence(#2).- README translations in Simplified Chinese, Russian, Hindi and Arabic under
docs/i18n.
Changed
- The
bide-auditHomebrew formula inbide-ai/homebrew-tapis maintained by hand instead of by the release pipeline.
0.2.0 - 2026-09-28
Added
agent.ToolResultCodec,JSONToolResultCodecandEncodeToolResultOr, andWithToolResultCodecon the OpenAI, Anthropic and Gemini adapters (#1).codec/gcfmodule: opt-in GCF encoding of model-facing tool results (#1).bide-auditHomebrew formula (brew install bide-ai/tap/bide-audit).
0.1.0 - 2026-09-28
First public release.
Added
agent: durable agent loop on an append-only journal, with at-most-once side effects keyed by tool-use id (Agent.Run,RunResult,Stream,RunSaga,StreamSaga,Session).agent:Safetyclassification,ResumeHaltandResolveHaltfor unknown outcomes.agent: durableStep,Parallel,SubAgent, and sagas withCompensatedFunc.agent: typed output withRunTyped[T]andRunTypedNative.agent: human-in-the-loop (Interrupt,Resume,Approve), durable timers (Sleep,WaitUntil,MemWaker), and durable signals (Signal,Await,AwaitFor, ordered channels withSend,Receive,Ack).agent: crash recovery and leasing (Recover,Lease,WithLeaseTTL,WithLeaseHolder).agent:WithSystemPrompt,WithSystemPromptFunc,WithMaxTurns,WithSampling,WithToolChoice, prompt caching, and multimodal image input.agent: retrieval seams (WithRetrieval,RetrievalTool),Replay,RenderMermaid, and sentinel error classification.model/anthropic,model/openai(any OpenAI-compatible endpoint) andmodel/geminiadapters.store/sqliteandstore/postgresdurable stores, withListerandLeaseron Postgres.middleware:Retry,Hedge,RateLimit,Cost,TokenBudget, per-attempt timeouts, and tool middleware (UseTool,ToolLog,ToolCache,ToolRetry,ToolRateLimit).trace: OpenTelemetrygen_aispans andtrace.Instrument.mcp: Model Context Protocol tools, with pagination, elicitation and tool-list changes.schema: reflection-based JSON Schema andOpenAIStrict.plan: typed flow builder (Step,Tool,Model,Edge,Switch,Join2,Join3,LoopBack), declarative config (Load), conformance checks and a topology digest.govern: convergent governed state on gsm, durable event logs (govern/sqlitelog,govern/redislog,govern/postgreslog), federation, identity binding andQuorum.audit: tamper-evident journal, RFC 6962 inclusion and consistency proofs, signed tree heads, proof bundles, absence proofs, ML-DSA signatures, anchoring (AuditedStore), run certificates (VerifyRun),EvidencePackage, signed grants and delegation (AttenuatingSubAgent, earned authority).bide-auditstandalone verifier, prebuilt for Linux, macOS and Windows on amd64 and arm64.evalstatistical evaluation harness,chaoscrash-injection harness, andcmd/bench.