scratch_* tools. These entries expire; use durable memory
for information that must survive across work.
Why it exists
Context is the only data bus between tool calls in a text turn. Every byte a tool returns is serialized back into the prompt for the next model call. A code-execution turn hides its intermediate state in the sandbox — files, variables, a 40k-line DOM parsed down to the one number the agent wanted — and returns only a summary. A browser, query, or workspace-read turn has nowhere to put a large result but the prompt. So the blob floods context, the model re-summarizes and truncates it, loses the thread it was following, and the turn fails. The scratch store is the missing out-of-context scratch space. It backs two surfaces over the same rows, distinguished by akind column:
- Auto output-handles. When a tool result exceeds a byte
threshold (
spillThresholdBytes, ~8 KiB), the full content is spilled to the store under a generated handle id and only a compact stub —{handle, preview, size, note}— enters context. Aread_handleverb slices the stored blob by byte range or returns the windows that match a substring, so a large result is read back in pieces instead of re-entering context whole. - Explicit scratchpad.
scratch_write/scratch_read/scratch_list/scratch_delete: a deliberate per-run key/value the agent writes interim notes to, over the same store.
What it is NOT
It is not durable memory. Versioned-state is what an agent chose to remember: curated, durable, retrieved across conversations, auditable forever. This is what an agent is holding right now: cheap, TTL-bounded, garbage after the run. The two never share a table, a retention policy, or a reader.internal/scratch
touches scratch_entries and nothing else; it never reads or writes the
versioned-state tables.
Scoping is multiplayer-first
A row belongs to one(run_id, agent_id) pair. The run has one owning agent.
Every read and write is scoped by both IDs and enforced in SQL. Deleting a run
or agent cascades to its scratch entries. Completing a run does not itself
delete its row; live entries otherwise remain until expiry or explicit deletion.
Retention
Entries carry anexpires_at (default one hour, scratch.DefaultTTL). Reads
filter on live expiry, so an expired row is invisible before it is collected.
runScratchSweeps deletes expired rows every few minutes as a backstop; the
primary GC is the ON DELETE CASCADE from the run row.
The spill decision
shouldSpill is deliberately conservative — when in doubt, do not spill, because
a wrongly-spilled result some downstream contract expected verbatim is a bug,
and a result left in context is merely large. It never spills when the
observation carries structured out-of-band meaning the stub would strip:
ArtifactID, MutationComplete, Metadata, or a memory-write receipt
(MemoryEntryID/Persisted). It skips the always-on internal verbs by name
(report_progress, the memory and scratch verbs, tool_search). It spills only
the result classes that actually flood context — query, workspace,
browser — and leaves effect-class results alone. A spill failure is never a
turn failure: the full result is returned to context as if the layer were not
there, with a loud log.
Relationship to compaction
This is complementary tointernal/turn/compaction.go, not a replacement.
Compaction trims the whole conversation window once it is already too big;
output-handles keeps one oversized result out of the window to begin with.
Compaction still runs over everything that does enter.
Inspection limits
There is no scratch HTTP endpoint or dedicated dashboard inspector. The store’sListScratchEntriesForRun query returns live-entry metadata (keys, sizes,
previews and source tool). The previously proposed
GET /api/agents/{slug}/runs/{run}/scratch route is not implemented.