Specification
The design axioms and execution model the implementation answers to.
0. This Document
This specification is normative for behaviour: what an agent, a person or a test can observe the system do, and what must always hold. Where the code disagrees, the code is presumed wrong, but the disagreement is a finding for a person to resolve, never something to align silently in either direction.
A sentence belongs here if it states a guarantee — something whose violation a person could observe — or the reason for one that no axiom already supplies. How the code achieves a guarantee, which module holds it, a configuration key or its default, and the alternatives that were rejected do not belong; the code, its comments and the configuration reference hold those. A new guarantee is written here as well as in the code that provides it; a new mechanism is not.
Every section is amendable, the axioms included. The weight of the case scales with what rests on the section: a proposal quotes the sentences it changes, names what elsewhere it invalidates, and says why the reason recorded for the current rule no longer holds. Nothing here changes until a person has approved the proposal.
1. What This Is
A single Node/TypeScript process running several LLM-backed agents. Each agent is a distinct identity that appears in our Mattermost workspace as a bot user, holds conversations with staff, and executes a fixed set of hand-written tools — subject to human approval on every consequential action.
The agents perform non-engineering business work: monitoring inboxes, drafting replies, researching on the open web including JavaScript-rendered pages. Each has its own system prompt, model, tool set, skill set, and private memory.
1.1 The Problem This Solves
We previously ran a third-party agent framework (Hermes). It failed in the following ways, and this design is largely a response to them.
- Arbitrary destructive commands against SQLite databases Unrestricted filesystem access; difficulty inspecting entire chained commands; reliance on shell as the primary work tool; agent not respecting the system prompt.
- Scattered files across the server Unrestricted filesystem access.
- Overwrote documentation with incorrect content Unsupervised writes; difficult to inspect, and rapid iteration meant nobody read the payload.
- Spawned sub-agents whose work was uninspectable Runtime agent creation; no durable identity per worker.
- Ignored system prompts, pursuing solutions at any cost Instructions were advisory; no hard constraints.
- Applied one-off patches to its own codebase Buggy, vibe-coded software with no boundary between the agent and the framework running it.
- Multi-agent gateway broken Third-party coordination layer we did not control.
- Mattermost integration missing critical features No approval mechanism existed at all.
These group into four classes, each with a structural response:
- Capability — reached things it should not have. Response: OS-level confinement and hand-written tools (A2, §6).
- Supervision — the action looked fine, the content was wrong. Response: approval gate with full payload disclosure (§6.3).
- Constraint — instructions were suggestions. Response: framework-enforced limits, never prompt-enforced (§5, §7.4).
- Substrate — the coordination layer was someone else’s. Response: Mattermost as the only substrate (A1).
1.2 What It Is Not
Not an autonomous agent platform. Agents do not run continuously, do not decide when to act, do not write their own instructions, and do not create other agents. Every consequential action stops and waits for a human.
The design trades throughput and autonomy for supervisability and predictable failure. That trade is deliberate.
2. Design Axioms
Most of this specification is a consequence of these five constraints rather than an independent choice.
A1 — Mattermost Is The Substrate
Mattermost is the only substrate for identity, addressing, and work delivery. There is no internal message bus and no scheduler-to-agent RPC path.
Triggers may originate anywhere — cron, HTTP, a mailbox poll — but an agent is only ever activated by a Mattermost post. SQLite holds a control state that points at posts; it never holds work that exists nowhere else. Every pending item in the system corresponds to a post you can scroll to.
Why: the alternative is a coordination layer living only in process memory — invisible in the client, lost on restart, and in the Hermes case owned by someone else. Under this rule, if work is happening there is a post you can point at. Debugging is scrollback.
A2 — No Ambient Code Execution
Every capability is a hand-written TypeScript function with a schema, reviewed and immutable at runtime. No dynamic tool creation, no runtime agent spawning.
Shell access is a capability that may be explicitly granted to an individual agent. It is not the default and not the normal way work gets done. Where granted:
- Every command requires approval.
- The agent runs as its own dedicated OS user, one per shell-holding agent.
- That user owns its home directory and nothing else, mode
700. - That user reads the agent’s own workspace and writes nothing in it, so a file the agent wrote reaches a command by its path rather than retyped into one. No other agent’s workspace is visible to it, and the framework’s own state is not.
- The framework’s own tree is unreadable by agent users.
- Agent home directories are unreadable by other agent users.
- The shell’s network reach is not policed. The address policy of §3.4 binds the web toolset, which is ungated, and not
shell::run, whose control is the approval on every command: a person reads the command before it runs, and that reading is the whole of the network policy. The preamble says so (§3.8).
The capability surface of an agent is therefore fully enumerable by reading config.
A3 — Deterministic Activation
Non-LLM code decides when an agent acts. The agent decides what to do once activated. An agent can never schedule itself, wake itself, or extend its own operating window.
Why: an agent with autonomous initiative and idle cycles will find work nobody asked for.
A4 — Clean Stop Over Graceful Degradation
When anything goes wrong — denial, malformed output, budget exhaustion, restart, rate ceiling — the agent stops and says so. It does not retry, route around, or degrade quietly.
Why: a system that degrades gracefully hides its failures. We would rather absorb interruptions than debug an agent that has been quietly wrong for six hours.
A5 — Visibility Is Not Consent
Every action an agent takes is recorded and surfaced. Blocking approval is reserved for actions with consequences, because approval frequency is inversely related to approval quality: a gate that fires constantly is answered reflexively and stops being a control.
Ungated does not mean unsupervised. This is why reads are ungated, why the status post exists, and why the full trace is retrievable on demand rather than streamed into the channel.
3. Concepts
3.1 Agent
A persistent identity with a name, persona, system prompt, assigned model, assigned tool set, assigned skill set, and private memory.
The prompt is written in configuration, either inline or as a markdown file the configuration names beneath the resources root, because prose escaped into a JSON string is prose nobody reviews. Which form is used changes nothing downstream.
Each agent is a Mattermost bot account with its own access token, addressable as @{username}, shown with a BOT tag in the member list. Prose names it by its display name — declared in configuration, else the username with its first letter capitalised — which is also the name its bot account shows. A name is no mention, so an agent written in passing is woken by nothing (§4.5); the handle is for addressing it.
Agents are persistent colleagues, not task-scoped job runners. There is one instance of each agent, continuously; conversations are episodes in an ongoing relationship rather than independent invocations.
Provisioning is a deployment act, not a runtime one: before the app starts, a separate process reconciles Mattermost against the configuration — the team, the channels, one bot account per declared agent, and the token each is addressed with. It converges; boot only refuses. Channel membership is not converged: every bot joins the main channel, an agent joins the channel its mail is announced to, and the system bot joins every channel the configuration names, so its notices (§3.2) have somewhere to land. Who else belongs in a channel is set in Mattermost by the people who run it, and §3.10 is checked against that membership rather than against a second declaration that would have to agree with it. Against a Mattermost it did not start, provisioning refuses before creating anything if the server’s settings forbid bot accounts, personal access tokens, or the callback address, naming each one. Nothing the running framework does creates an account or a channel, and there is no runtime agent spawning: every worker has a durable identity and its output is in a channel.
3.2 System Bot
A separate Mattermost bot account used for all mechanical output: triggers, boot announcements, interruption notices, stall notices (§7.6), and refusals.
Invariant: a post from the system bot is never an agent thinking. All of its output is fixed strings or templated facts, never LLM-generated. The same rule applies when the framework posts under an agent’s account — budget notices and stop notices are deterministic code speaking as the agent, not the agent speaking.
Nothing the system bot posts enters an agent’s queue (§5.2). Trigger delivery is governed by §4.2 instead.
A notice names an agent by its display name (§3.1), not its @, since a mention from the system bot activates the agent it names (§3.10); the handle is written with its @ only where the post means to start that agent’s turn, as a trigger’s delivery does (§4.2). A command a notice quotes keeps the handle it takes.
3.3 Turn
The unit of execution. One activation of one agent in one channel: assemble context, call the model, execute tools, produce output, terminate.
A turn is created by exactly one of:
- a human posting in a DM with the agent
- a human posting in a respond-to-all channel
- a human mentioning the agent in a channel
- another agent mentioning the agent
- the system bot posting a trigger that mentions the agent
Posting is performed by the framework, not by a tool. Text emitted alongside a tool call updates the status post and is transient. Text emitted with no tool call is the turn’s final output and terminates the turn, and may be empty only in a turn that owes no reply, after a unit post of the turn addressed a colleague (§3.15). Such a turn may also end at that unit post, where the call that made it asked to (§3.15); nothing is posted after it. Posting does not consume the action budget. A framework tool may return a post beside its result, which the framework publishes under the agent’s account from the tool’s own template — the arrangement §3.7 uses for an approval prompt and §3.15 for a work unit’s state change. A plugin tool cannot return one.
3.4 Tool
A hand-written TypeScript function exposed to the model with a schema. One tool per action — there are no tools whose behaviour branches on an action argument; a tool that would is two tools.
A tool’s identity is two segments, [namespace, tool], rendered per audience: operators, approvers, config, traces, and errors see mail::send; the model sees mail__send. A call arriving in the display form is resolved all the same, since the framework’s own approval posts put mail::send in the agent’s window (§8.1). A call naming the tool segment alone — send — is resolved when exactly one granted tool carries it, because a model drops a namespace far more often than it invents a tool. No spelling resolves to a tool the agent was not granted; a name claiming anything else is answered as §7.2 says.
A text names another tool as the way on only where the agent is granted it. A text rendered for a call or a turn — a result, a refusal, a reminder, a section of the prompt — offers a tool only where the acting agent holds it by its full namespace::tool identity, and names the next true step or none where it does not. A description, which every agent granted the tool reads alike, may name the tools it works with.
A toolset is a namespace and everything that belongs to it: its tools, the services they may reach, the settings they are configured by, the storage collections they own, the skills they ship.
Each tool declares, alongside its description and parameters:
approval— present means the tool always gates (§3.7): the function renders the payload the approver reads and cannot decline. What it renders is the model’s own arguments, the operator’s own configuration, and the records in the toolset’s own storage that the call acts on — never text a call the tool itself makes returns, which is unverified once anything external supplied it. An approver asked to delete a record reads the record as it is stored, not the model’s description of it. Whether a granted tool can act without a human is therefore answerable from config alone, and/collegium inspectmarks it per tool; a tool whose gate would depend on its arguments splits into two tools. A plugin tool must state the choice,nullfor “this does not gate” (§3.14), because an omitted field cannot be told from a forgotten one.retryable— whether a call is safe to run again with the same arguments: a timed-out one may be reported to the model as a plain failure (§7.2), and a restart may run its turn again (§7.3). A mutation declares it only where repeating it changes nothing; false, the default, ends the turn on a timeout as an unconfirmable side effect.budgetExempt— never billed against the action budget (§5.3). Framework toolsets only; a plugin cannot declare it.concurrent— may run alongside the other concurrent calls of the same completion (§5.1): a read that neither depends on nor disturbs what another call in the batch touches. A browser action is not one.supersedable— a page whose identical re-read replaces it (§3.8): a later result of the turn with the same content, at whatever address served it, costs a line naming this one while it is shown, and is shown again, naming it, once it has been replaced. It plays no part in what is collapsed.
A tool reaches only what its toolset declared. A tool that creates a durable record returns its disclosure — body, description, reference, anything superseded — and the turn writes it into the trace (§3.6).
Reads are generally ungated: search, fetch, read mail. Writes, shell commands, and anything externally visible carry approval.
Authority parameters are never model-supplied. Any argument determining whose authority an action carries is fixed in tool settings — the from address on outbound mail, credentials for external services. The model may request that mail be sent; it cannot choose who it appears to be from, and boot refuses a mail-granted agent without a mailbox (§3.13).
Filesystem scope. The workspace tools and shell::run are confined to the agent’s own directories — the former by path confinement, the latter by OS permissions (§A2). They are not the same directory: the shell user reads the workspace and cannot write it, and the workspace tools do not reach the shell’s home (A2). Both paths are stated in the preamble (§3.8) rather than left for the agent to discover by failing, as are the commands the shell offers, probed once at boot. The shell’s network reach is bounded by the approval on every command and by nothing else (§A2). Purpose-built tools may write to real systems by their own internal logic; those are individually reviewed and their write targets are fixed in code, never chosen by the model.
Reading the workspace is not a shell command. The workspace toolset’s reads — workspace::list, workspace::read, workspace::find, workspace::grep, workspace::stat — take typed arguments under the same path confinement as workspace::write, so there is no command string in which a second command could hide. Each is ungated, because the confinement that bounds a write bounds a read of the same directory. The shell’s own directory is not reached by these tools; a read there is still a shell::run under the gate.
Shell output a result cannot carry is kept where the agent can read it. A command’s result carries a bounded prefix of each stream. Where the output ran past it and the agent holds workspace::read, the whole capture is written to a file in the agent’s workspace and named in the result, so the rest is a ranged read rather than a second command under a second approval. The capture is itself bounded, says how much it dropped, and only the most recent few are kept. An agent without workspace::read gets the prefix alone.
Browsing. The web toolset drives a real rendered browser in a turn-scoped session: web::navigate, web::click, web::fill, web::hover, and web::select act against element refs from the latest snapshot, each action returning the page as markdown: its form controls and any tab it opened first, then the page, shown up to a fixed width with the rest read on by reference (§3.8). The controls are bounded too, the rest listed after the page, since an application page can carry hundreds. Where a page grows with each action, so that the view shows only what the snapshot before it showed, the result says where in it the page first differs. In the page an image is its alt text alone: its address is a read the model cannot make. An action is captured once the page has redrawn from the script requests it started, as far as a bounded wait allows. The status a result carries is that of the document the page settled on, not the first one it was served: a bot check that clears itself, or a page its own script moves on, has loaded another after the one asked for. A ref the page hides — a submenu that opens on hover — is marked hidden, and an action attempted on it is a typed refusal naming web::hover; among the controls, the hidden follow every visible one under a line saying they must be revealed first, so the controls’ bound falls on them first. A drop-down’s options are listed once, bounded: beside its ref on a snapshot, and inline on a fetched page. web::select chooses one by its label or value; an option it does not offer is a typed refusal that names the options whose labels contain what was asked. A ref no longer on the page is refused with the page as it now is. An action that fails on an element that is present and visible is reported as that element’s failure, not as the page failing to load, since the page is still there to act on; a fill aimed at a drop-down is refused at once, naming web::select. A tab the page opens for itself is closed and reported with its address, so the session never moves onto a page nothing asked for. It may submit forms and sign in where the task calls for it — a great deal of the open web is unreachable otherwise.
It is ungated all the same. The per-agent grant decides who may browse at all, the status post traces every action live, and a per-action approval would stall autonomous turns while approving only an entry URL — never the later actions that could transmit. This is the widest ungated surface in the system and is named as such rather than hidden: an agent that browses can transmit on its own authority.
What a fill types is shown like any other argument — in the status post, in the trace, and in the snapshot that follows — so nothing in this toolset is built to carry a secret. The one compensating control is scope, which is enforced, not requested: only http(s) URLs are opened, and only on the public internet — a file:// URL, or a host naming this machine or its own network, is a typed refusal. The check runs on every request the session makes, not only the address the model asked for — a redirect, a same-session link click, or a page’s own sub-resource fetch is judged the same way, with the name resolved and every answer judged, so a public URL that leads somewhere private does not reach it. A load it stops is the typed refusal naming the address it stopped, not a page that failed to load; a sub-resource it stops leaves the page without it. A deployment may declare its own network browsable, so that pages it serves itself are reachable by its agents; the declaration lifts the private-address refusal for every agent and every request, leaves the scheme rule standing, and is logged at boot, so a deployment that has opened its network has said so in its configuration and in its log. A deployment may also close hosts to its agents by name, each with every subdomain beneath it; nothing is closed by default. The refusal is the same typed one, on either route, whichever way the private-address rule is set. It is the operator’s control alone, never an agent’s, and it too is logged at boot.
Sessions are fresh anonymous contexts disposed at turn end — no cookies persist and no login outlives the turn that made it. A deployment declares how many may be live at once, since each holds a rendered page. A turn that would open one past that number is refused, not queued, because a wait inside a turn is a stall the model cannot see; the refusal says that a session frees only when the turn holding it ends, and that web::fetch needs none. robots.txt is not consulted: these are agent-driven reads at human pace, not crawling. A resource that is not HTML is a typed refusal, and one web::fetch reads — a PDF, a text file — is refused naming it.
Beside the browser sits web::fetch: one plain HTTP GET, no session and no script, converted to markdown by the same rules and bound by the same URL policy on every redirect hop, with the connection pinned to the address that passed. A 429 or a 503 is asked once more when the site names a wait, or after a second for a 429 that names none, provided the wait leaves the retry time to answer within the request’s own timeout; the result says it was retried and after what. Nothing paces requests to a host before it refuses one: a fixed limit would slow the many bursts that are served to spare the few that are not. An email address Cloudflare cloaks for its script to decode is decoded in the conversion as that script would, so a fetch reads the address a browser shows rather than a cipher; one that does not decode is left as served. It is the cheap first read for static pages, and it never starts a browser, since whether one is worth its slot is the agent’s call: a page that yields nothing without JavaScript is a typed refusal naming web::navigate, and so is an HTML page the site will not serve it — a 401, 403 or 429 whatever it says, a 503 on a CDN’s refusal page, or a bot check in the page’s place at any status — whose body describes the refusal, not the page. A success with an empty body is reported as empty, naming no browser. An error status over nothing readable is reported as that status; only a 404 or 410 says there is no page at the address, and a result with either says too that an address the agent built, rather than read off a page, proves nothing about the page it was after. A textual non-HTML body is returned as it is, at any status, since the browser does not open one: an API’s JSON error is its answer. A fetch reads a page’s main content: the one main landmark where the page marks exactly one, else the page less its navigation and its page-level banner and footer. The result says how many characters that left out, and wholePage reads them too, since a page may keep what is wanted in its footer; a browser snapshot keeps its navigation, because its refs include it. A result holds a bounded window of what was read, its width one constant, and says where to read on; find returns instead where each of a few phrases occurs, with a bounded span of text around each and the offset a window reads from, so a field in a long page costs a search and a small window rather than the page. A find searches the whole page, and says so when window arguments were given. A find shows text shared by several places once, and returns the page whole, saying so, where its places would take more text than the page. A PDF is read as its text layer, each page under a marker giving its number, and is windowed and searched the same way; a PDF with no text layer — a scan — is a typed failure saying so, since nothing here reads text from an image. The parser never runs in the app’s process: a PDF’s compressed streams are untrusted, a file of a megabyte can inflate to gigabytes, and a parse cannot be stopped in the middle of a page without stopping the process running it. Each PDF is read in a child process of its own, killed once its resident memory passes a fixed cap or the read’s own deadline passes, and offered first to the kernel’s out-of-memory killer should that act before the cap does. Pages reach the app as they are read, so a read stops once its text passes the ceiling a page’s markdown has, at its deadline or at the cap, keeps every page read whole by then, and says how far it got and why. A read stopped before its first page is a typed failure naming the limit, and one stopped short is never taken for a scan, since the pages it did not reach may carry the text. A fixed number of PDFs are read at once, so the cap bounds their sum as well as each; a read beyond them waits within its own deadline, and one still waiting when it passes is a typed failure saying that asking again may read it. A PDF past the body cap is not read at all, since a PDF cut short does not parse.
A certificate that fails to verify, on either route, is a typed failure in the framework’s words, never the runtime’s, whose advice about its own flags reads to a model as a fault in this deployment. The site is named as at fault only where the error establishes it — a chain sent without its intermediates, an expired or self-signed certificate, one issued for another name; an issuer nothing here trusts is as likely a gap in this deployment’s trust store, and is said to be either.
web::search queries a search provider and returns ranked summaries — title, URL, snippet — never the pages themselves. The provider and its key are fixed in the agent’s web settings, never model-supplied, and an agent whose settings name no provider is not offered the tool. Throttling is a result the model reads; refused credentials or an unanswering provider end the turn.
Grants and settings are one config mechanism. An agent is granted namespaces or single ns::tool refs; a namespace grant covers tools added to that namespace later, so a plugin update can widen an existing grant with no config change — accepted, since installing a plugin is already a trust decision. Effective settings per toolset are the deployment defaults merged shallowly with the agent’s own, then parsed against that toolset’s own schema. Settings for an ungranted toolset are an error, and a granted toolset whose merged settings fail its schema is an error — “mail requires a mailbox” is what the mail settings schema says, not a rule anyone maintains. A tool may declare that it works only under certain settings: a namespace grant leaves it out until the agent’s settings enable it, and naming it by ref without them is a boot refusal.
Core tools are framework machinery, not grantable capability. builtins::now, results::read, skills::load, and triggers::resolve are in every agent’s tool set — reading the clock, reading on in a result of the turn that is shown in part or as its line, loading an assigned skill or one of its references, clearing a trigger the framework itself raised — and naming one in config is an error. A plugin cannot contribute one.
3.5 Skill
A procedure stored in the repository as a directory holding a SKILL.md with a description, an optional list of the tools it calls, and supporting reference documents beside it. Skills are validated at boot, and a malformed or undeclared one refuses the boot.
Each agent’s system prompt carries a skill manifest: the names of its assigned skills plus one line of description each. The full body is pulled into context on demand via skills::load, and one of that skill’s references by the same call naming it. The body returned carries the skill’s reference index, so an agent that has loaded a skill can never be unaware which references it has.
A skill shipped by a toolset is namespaced like a tool — bookmark::saving-bookmarks — and granted the same way; any toolset may ship skills, framework or plugin. Framework skills belonging to no capability keep bare names, which cannot collide, because every other skill carries a namespace. Granting a namespace’s tools does not grant its skills, or the reverse. Core skills — handing-work-to-a-peer and understanding-collegium — are in every agent’s manifest, are not grantable, and naming one in config is an error.
A skill may name the tools its procedure calls. Where it does, boot refuses an agent granted that skill without them: a procedure whose third step is a tool the agent does not hold is a turn that goes wrong in the middle, where the pairing was knowable before the process started. The list is the author’s claim about the procedure, not a grant and not a filter — it reaches neither the model nor the tool registry.
Agents cannot write skills. Skill authorship is an administrative act. The manifest is what makes on-demand loading safe: an agent may fail to judge when a procedure applies (an ordinary error) but can never be unaware that it exists (a structural blind spot).
Where a procedure is load-bearing, encode it as a single hand-written tool rather than describing it in a skill: an ordering that must hold, or a payload that must carry a particular part, such as the criteria a hand-off will be judged by (§3.15). A procedure in a prompt is a suggestion; the model will eventually skip the step that mattered. What a skill is left with is judgement.
3.6 Memory
A per-agent store, reachable by agents only through tools. Each entry has a description (the trigger) and a body (the content).
- Descriptions are in context on every turn, after the window (§3.8), under a count of the entries held against their cap. Bodies are loaded on demand.
- Writes, revisions and deletes are ungated — the single exception to A5.
- Entry count, description size, and body size are all capped by the memory toolset’s settings (§3.4), which the preamble states for the agent (§3.8). An over-length description or body is refused, never truncated. A write or a revision reports the body’s size against its cap, and a revision refused at the cap states what the entry holds and what the change would make it, since the remedy is room made in that entry, not a second entry beside it. A write at the entry cap evicts the entry whose body was read longest ago, since the store already knows which entries an agent keeps needing; the write’s result, the trace and the record’s line in later windows name what it evicted.
- Every entry carries provenance: written-at timestamp and originating post ID. Entries are shown to the agent, in the trace, and to
/collegium memoryby a short reference, which the store resolves back, refusing rather than guessing if it ever matched two. - A body is read back with its age and its written-at date above the text, and with the age of its last revision once it has one, so a model reasons about staleness from the ages and about provenance from the date. None of them is in the listing, so the prompt does not change as entries grow older.
- An entry keeps its reference for life. A revision —
memory::appendadds to a body,memory::replacesubstitutes passages of it, all or none,memory::rewritereplaces the whole of it — edits the entry in place and counts it. The description carries over unless the revision names a new one;memory::replacenaming only a new description changes the description alone and counts as a revision. The written-at date and the entry’s place in the listing carry over; the provenance is the revising turn’s. A correction still reads as a correction to an operator:/collegium memorylists how often and how lately each entry was revised, and the trace names each revision’s count and what it replaced — passages, a whole body, a description. - A revision is one step, not a delete and a write. Two steps leave the only copy of the body in the model’s context between them, and two turns revising one entry from two channels would each rewrite a body the other has not seen. An append or a replace sends only the change and applies it to the body as stored; a passage
memory::replacedoes not find exactly once is refused as a result (§7.2), and so is every other edit of the same call; where the turn has not read the revision it quoted from, the refusal says so first. A rewrite sends a whole body, and is guarded as a delete is. Concurrent turns of one agent cannot lose each other’s writes, evictions or revisions. - A turn deletes or rewrites an entry only as it has seen it. A turn has seen an entry’s revision when it read the body, wrote or rewrote the entry, or revised it from the revision it had seen. A delete or a rewrite of an entry the turn has not seen, or of one revised since, is refused as a result that names the read to make (§7.2), since it would discard a body the turn never read or what another turn added since. An append or a replace needs no such check: it applies to the body as stored. What a turn has seen ends with the turn. A prune through
/collegium memoryis an operator’s, and is not checked. - Memory is per-agent and never shared between agents. A reference that resolves to nothing — one since deleted, or another agent’s — is refused, and the refusal says it may have been deleted and, as the preamble does, that memories are private.
Why ungated: gating a memory write would block an entire turn on a triviality. Memory formation cannot sit behind human latency or it will not happen. Deletion inherits the exemption: an agent that cannot retract a fact it now knows to be wrong carries that fact into every later turn.
Compensating control: a write or a revision returns a disclosure — description, body, reference, a revision’s count, anything superseded — which the trace records, and the status post carries the call like any other tool call (§8.1). This is detection, not prevention — the write has already happened.
Memory is one of two paths by which information crosses channels — the other is search (§3.8). Both are intentional, both leave a trace, and the channel window itself is strictly channel-scoped.
3.7 Approval
A blocking request for human consent, rendered as a post under the agent’s own account with interactive buttons — the payload is model-proposed content, which §3.2 forbids the system bot to carry.
- Approve — the tool executes; the turn continues.
- Deny — the turn terminates; the agent posts asking how to proceed.
- Deny with reason — opens a dialog; the reason is fed back as the tool result and the turn continues under the same budget.
The prompt must show the full payload, not just the intent. Where a payload exceeds what a single post can carry, the prompt shows a bounded prefix inline and the complete payload as an attachment (§6.2). Once resolved, the prompt post is rewritten into a terminal state and its buttons removed. The prompt is recorded as the turn’s own, so the agent’s window leaves it out as it leaves out every post the turn authored (§3.8): the decision reaches the model through the call’s result, never as words of its own it then continues.
The prompt also carries what the framework knows and the payload cannot say. Above the payload sit facts an approver would otherwise reconstruct by scrolling: which action of the turn’s budget this is, and who asked for the work — by name, with a bounded excerpt of their own words, and only where a human asked. Where a colleague asked, the line names it by its display name (§3.1) and, beside it, the person whose request that colleague’s own chain descends from (§7.4), with the same excerpt, so an approver on a delegated call is not deciding on a colleague’s word alone. Where the turn was started by a colleague’s assignment post, the line also names that work unit (§3.15), and tags the person without the excerpt: the unit carries the instruction the call serves, and words from before it would read as that person asking for this call. A trigger-initiated turn names the trigger’s source and item instead, and says that no person asked, except where the operator’s schedule raised it (§4.2). Where the call follows a reasoned denial of the same tool in the same turn, the line names that denial, so the payload reads as the amendment it is. Every word of that line is framework-authored from the turn record; nothing a tool returned and nothing a page, file or message contained ever reaches it. Nothing here is a gate.
There is no timeout. The agent waits, for days if necessary, until a human answers. Work queued behind it accumulates without bound (§5.2). Since nothing expires and nothing chases, /collegium approvals (§8.4) is how a human finds what is still waiting.
Who may approve: any human present in the channel. There is no separate approver list, because a config-based roster would be a second access-control system that silently drifts from channel membership. A decision from outside the channel is refused.
There is no standing approval, and no session or per-run grant. A gated tool gates on every call; nothing a human clicks makes the next call cheaper. A standing grant can only bind a call’s effect — a tool and a target — never its content, which is different every time; and every tool that gates today gates because of its content, so a grant bound to mail::send → alice@example.com would authorise every future body sent to Alice, which is precisely the failure §1.1 records as nobody reading the payload.
3.7a Ask
A tool may declare ask instead of approval: a blocking request for a fact only a human has, rendered under the agent’s own account like an approval, with the same no-timeout and channel-presence rules as §3.7 and listed beside the approvals by /collegium approvals (§8.4), but resolved by a text answer rather than approve or deny. It carries the same framework-authored line above the question as §3.7 puts above a payload, and a person that line names is tagged: a parked question notifies the person its work descends from by construction, whatever the model wrote. The answer is fed back as the tool’s result and the turn continues under the same budget — there is no denial, because there is no action to refuse. A tool that would gate on consent uses approval; a tool that would solicit information uses ask; no tool declares both. Where the turn wrote text alongside the call, that text leads the prompt, so the person deciding reads the agent’s own reason for asking.
ask::human is the one framework tool that uses it: a question, and optionally a few short answers offered as buttons alongside free text. It is not a substitute for approval — a tool that needs permission still gates, and the framework preamble tells the model so.
3.8 Context Window
Each turn assembles context fresh from the store:
- System prompt
- Skill manifest (§3.5)
- Tool definitions
- Channel window — recent posts from the current channel, interleaved with the trace of this agent’s own turns there (§8.2), walked backwards until a token budget is exhausted. The trace is never a peer’s: an agent reads its colleagues through their posts alone — never their status posts, which are their trace rendered.
- Posts this turn answers — the posts §5.2 defines as what the turn answers, oldest first: a person’s or a colleague’s with who wrote it and an excerpt, a trigger’s by its id and source and never quoted; each where it sits in the window, or, older than the window, with the tool the agent holds that reads it whole
- Today’s date, in the operator’s timezone, which it names
- Channel — its name and handle, or who a direct or group message is with
- Memory descriptions (§3.6)
- Pinned posts — every post pinned in the channel, oldest first, whoever wrote it and however long ago (below)
- Earlier actions — a fixed number of lines naming what this agent itself did in this channel before the window reaches, newest first, in the replay form defined below. Its own turns only, never back past the channel’s episode boundary. Mechanical and derived: nothing is stored, nothing is summarised, and no model decides what was worth keeping.
- Peer roster, and the people who posted here (§3.11)
- Open work — the units this agent created or was assigned in this channel and has not closed (§3.15), oldest first, one line each with where the other party stands, or a line saying none is open, so that absence reads as a fact and not as an omission; then what starts the agent’s next turn here. Read on demand for the rest, as a memory’s body is.
Nothing whose text can differ between two turns of the same agent in the same channel precedes the window. The first three items change only with the deployment. The rest follow the window as one message, rendered afresh each turn and opened by a line saying it is the framework’s and not a post; the turn’s own calls and results follow that message. It is sent in the user’s role, because a system message part-way through a conversation is one a provider may move back ahead of the window. The date, an age, the memory listing, a pin and a channel’s membership all change between turns, so a section that carries any of them goes after the window, however seldom it changes.
The window names each post’s author and what they are — a person, an agent, or the system bot — from what the store recorded when the post was observed, and without an @: a name that read as a mention was copied back into replies. An agent is named by its display name (§3.1), so it is copied as prose; a person and the system bot by username, which is what a reply that tags the person needs. A post’s attached files are named in the window, one line each — name, type and size — whether or not anything can read them. A file an agent cannot read is a fact it states, not a fact it is spared: an agent that answers the caption as though it were the whole message is wrong in a way nobody can see.
A pinned post is a standing instruction from the channel’s people. Memory is per-agent (§3.6) and the window is recency, so a ruling a person gives a channel has no other home that every agent there sees and that outlasts the window. Every post pinned in the channel is rendered after the window for every agent present, whoever wrote it: author and what they are, as the window names them, the day it was written and its id, then its text delimited as a search hit is. Pins are set by people in the client; no agent holds a pin operation, which is the guarantee, since Mattermost does not say who pinned a post. The section is held to a fixed token cap. The newest pin is always shown whole, since one post is bounded by the substrate’s post limit; older pins that fit under the cap follow it, and the rest are counted and named by id rather than dropped in silence (A4). A pinned post is not bounded by the episode boundary — /collegium reset leaves it standing — and a forgotten one is not listed (§8.4); /collegium clear deletes pins with the posts (§8.5). The preamble says a pinned post outranks a colleague’s paraphrase of it.
The system prompt contains the agent’s own prompt, the shared behavioral baseline, its optional personality, and the framework preamble.
The preamble states the one runtime fact about the filesystem an agent cannot otherwise see without failing: the directories its file tools point at. workspace::write and the workspace reads share one directory; shell::run runs in a different one, and reads the workspace without writing it (A2). Naming them grants nothing: confinement is path rooting and OS permissions (§6.1), neither of which depends on the model not knowing where it is.
What is deliberately not there: the time of day, and the host’s operating system, git state and processes. The date is in the message after the window, where its change at midnight costs the cache nothing; the time is a tool call, builtins::now, because a time stated as the turn starts is stale within minutes and would still be read as the time now. The rest is ambient state the agent was not granted, and a snapshot of it in the prompt is a read nobody approved. An agent that needs git state holds shell::run and asks for it under the gate.
The behavioral baseline is rendered for every agent on every turn: task intent, routine autonomy, scope changes, recovery from failure, proportionate verification, untrusted content wherever it arrives (tool results, mail, webhook and quoted text in the channel), collaboration and a colleague’s authority, how a person and a colleague are named, reply discipline, what is worth keeping in memory and what to believe when a memory disagrees with what the agent can see now, and disagreement. The memory paragraphs render only for an agent that holds memory. The baseline is held to a budget rather than a list: an instruction added is paid for by one removed, since compliance with all of them at once falls with their count. These are advisory instructions about how the model should work, separate from runtime facts and optional tone. A personality adds a stance to that baseline, selected per agent or by default from a fixed set the framework ships.
The preamble describes runtime behavior in Simplified Technical English, using the deployment’s actual budgets and exemptions: the shape of the window and who authored what, the framework’s message after the window, pinned posts and who sets them, context retention within a turn, interim text and the status post’s audience, posting and what starts the next turn (which Open work states instead for an agent that holds tasks, §3.15; to an agent that assigns or reports units, that a unit post may ask to end the turn, which ends there only where it owes no reply and no unit a move), steering, approvals and the shape of the prompt, the budget, memory and its caps, the file directories and the shell environment, peers and the one-colleague rule, the colleagues an agent may hand a unit to where they are declared (§3.15), what an @ does for a person, triggers — a schedule’s instruction, the quoted bodies mail and webhooks carry, and the id that resolves one — and the prohibition on self-modification. Every sentence states a runtime fact that holds whether or not the model complies. Additional deployment guidance belongs in a personality or an agent’s own prompt.
An agent’s context budget is the one it declares, else the deployment default, else a fixed share of its turn ceiling (below); a model cannot be offered without a recorded window.
There is no threading. All posts are channel-level, so context is pure recency: no structural marker indicates where one piece of work ended and the next began. /collegium reset {agent} provides a manual episode boundary; context never reaches back past the most recent one, pinned posts aside (above). The boundary hides the agent’s own earlier actions as well as the posts: its turns’ calls and results are cut at the reset itself, so an action taken after the channel’s newest post but before the reset is hidden too. /collegium clear is the boundary for every agent at once, with the record removed rather than hidden (§8.5). DM context follows the same mechanism as any other channel.
The window’s oldest entry holds still. Once a window has been built, the next one for the same agent and channel starts from the same oldest entry for as long as everything since it fits the budget; when it no longer fits, the window is trimmed from the old end, below the budget, and holds there again. A window built afresh, after a restart or a clear, is walked to the same share a trim leaves, so it too has room to grow. A window whose start moved every turn would never be cached past the system prompt. A restart costs one cache miss.
A tool result replays as what it was, not what it said. A tool may return, beside its text, the line later turns see in place of it: a skill body replays as [loaded skill x], a page as its address and size, a long shell output or mail message as its size. A result longer than 2,000 characters whose tool returned no line replays as the tool’s name and the result’s size, whatever the tool; a shorter one is kept as it was, being cheaper to keep than to describe. The turn that made the call reads the text, whole or as its view with the rest by reference (below); a later turn reads the line and can make the call again if it needs the text. Within a turn, a result stays as it arrived until relief under pressure replaces it with its line and its reference (below). A result byte-identical to one the turn still shows costs one line naming it; one identical to a result since replaced is shown again, naming the earlier one. A supersedable result is identical when its content is: a tool may name the part of its result that is its content, and a page read successfully is its body, whatever address served it. The line that answers such a read names the earlier one and says only that the content is identical, never why: an ignored parameter, a soft 404 and a login wall all look alike. The trace keeps every result in full.
The budget is charged what the window renders. A post costs its text as the model reads it, attachment lines included. A response of the agent’s own costs its text and each call’s name and arguments — never its reasoning (§3.12) — and is costed as one unit with the results and decisions answering its calls, since a call is sent beside its result or not at all; a result costs its replay line where it has one, and a written memory costs its one line. Each unit is costed before it is admitted, and the walk stops at the first that does not fit rather than skip it for something older, which would leave a gap the model cannot see.
A turn’s own context is bounded too, and a result is a record. The budget above bounds what a turn starts with; the agent’s turn ceiling bounds what it may accumulate. The ceiling is the one the agent declares, else the deployment default, and never more than a fixed share of the model’s window, the rest being the completion still to be written and the slack an estimate owes a tokeniser. It is declared rather than taken from the window, because a window measures the most a model accepts, not what one turn should carry; at a million tokens, limits drawn from the window never fire. A budget that is not below the ceiling refuses boot, since a window that filled it would leave the turn no room to start. Every tool result is kept whole in the turn’s record; what the turn’s context holds is a view of it. A result no longer than a fixed share of the ceiling (at most a fixed size) enters whole, and a longer one enters as its first part — shorter still where its tool declares a narrower view — followed by a line giving its whole size, when it was recorded and the reference by which results::read reads on from any offset, or finds phrases in it, for the rest of the turn; each read is billed like any call. A reference lasts as long as the turn that made the call. Under pressure, results the model has already read are replaced by one-line stand-ins carrying their reference and when they were recorded, largest first, until the context is well under the ceiling, so one rewrite buys room for many completions; the views of the results just received may be shortened together, never without their reference, and one that cannot keep even a floor shows only its size and its reference. Nothing is cut silently and nothing becomes unreadable, which is what A4 requires. The ceiling is checked before each completion, so every call a completion made runs; a call that would park a person is not run while the context is over the ceiling. A turn whose context still exceeds its ceiling once every result it has read is a line ends, saying so, and the queue drains as §7.1 says.
Prompt caching. What precedes the window changes only with the deployment, and the window’s oldest entry holds still, so a provider that caches prompt prefixes can reuse, at each turn start, everything up to where the previous turn’s window ended; the turn pays afresh only for what was recorded since and for the message after the window. Caching never changes what the window contains or preserves stale context. A provider that serves a model from several upstreams caches each one separately, so a model reference may name the upstreams to prefer, in order, and the framework passes them on; the model itself never changes (§9). /collegium usage reports cache reads where the provider supplies them, and each turn records them beside its prompt tokens.
Search reaches what recency cannot. conversations::search is a read over the post store — a case-insensitive substring over post text, optionally bounded by author and date — returning each match as post ID, channel, author, time, and text: the text delimited rather than requoted, and bounded per hit around the match, with the whole post readable by its ID. Reading a post by its ID is the same read under the same bounds, so a post out of reach reads exactly as one that was never written, and the post that triggered the turn is never among the matches, being already in context verbatim. It returns posts alone: never a trace, which is per-agent and need-to-know (§8.3), never stored reasoning (§3.12), and never a status post or a notice — and the tool’s own description and its empty result both say so, since a bound the model cannot see is one it searches around. An agent’s own replies are posts like any other, and are found, and so are the posts that assign, report or close a work unit (§3.15), which are a kind of their own and the durable record of delegation. It is ungated, as reads are, and billed against the action budget (§5.3): a search that can run forty times is the case the ceiling exists for.
Search is bounded exactly as the window is, applied per channel. It reaches only channels the agent is a member of now, and back no further than each channel’s own most recent episode boundary. A forgotten post (§8.4) is as absent from a result as it is from the window, or /collegium forget would be a lie.
It never widens an audience. A match surfaces in the current channel only if everyone who can read the current channel could already read its source. From a public channel, search reaches public channels alone; from a private channel or DM it reaches public channels and any private channel or DM whose members include everyone present. A result names its source channel, so the model can see it is quoting from elsewhere.
Why a structural rule rather than an instruction: a colleague knows not to repeat a private conversation in a public room; a model quoting a raw search result will do so by accident. What was said in private and is wanted elsewhere crosses by memory (§3.6), where the human saw the disclosure at the moment it was kept.
Search reaches what was said; the earlier-action lines reach what this agent did. A trace is per-agent and need-to-know (§8.3), so search cannot return one, which would leave an agent able to find a colleague’s sentence from last week and unable to recall the file it wrote itself two turns ago. The lines are carried mechanically instead — no second model call, no stored interpretation of the channel to disagree with the channel. They say what the agent did, never what a result said.
3.9 Work Channel
Each agent generally has a dedicated Mattermost channel containing that agent and its authorized humans.
All channel configuration is declared in configuration. Agents have no input into topology. Membership is not part of that declaration: it is held in Mattermost and set by the people who run it (§3.1), so what configuration states is which channels exist and how each one triggers.
3.10 Triggering Mode
A per-channel flag set at provisioning:
- mention-required (default) — the agent acts only when explicitly
@-mentioned - respond-to-all (opt-in) — every human post in the channel starts a turn
- DMs are respond-to-all inherently, as a property of the channel type
In respond-to-all channels, agent-authored and system-bot posts must not trigger, or an agent will reply to its own output and loop immediately. The rule is: human-authored post, or a mention from an agent or the system bot.
A respond-to-all channel contains at most one agent. Every human post starts a turn for every agent present, so two agents in one such channel produce two concurrent turns on the same task — precisely the harm §4.5 exists to prevent, arrived at without a mention. This is checked against Mattermost membership at boot and on every membership event; a violation discovered at runtime trips the global halt (§7.4).
3.11 Peer Roster
The set of other agents present in the current channel, rendered into every turn’s context after the window (§3.8), excluding the agent itself and the system bot. This is how an agent knows which peers it can reach.
Each peer is listed by name and handle, with what it can do: its expertise, and the toolsets it was granted, by namespace — never its individual tools and never their settings. The name is how the agent writes the peer; the handle, written with its @, is how it addresses one (§4.5). An agent that sees a colleague’s name and nothing of its reach plans against capability it cannot see. The namespaces are the grants configuration states (§3.4), which A2 already makes enumerable, so listing them tells the agent nothing an operator could not. The core namespaces every agent holds are left out.
Membership is what Mattermost holds now, never a copy that could drift from it.
Beside the roster, the people who posted in the channel, a handful, the latest poster first, each by @username: a person is always named with the @, which notifies them and starts no turn, where a colleague is @-tagged only when addressed. They are read from the author kind the store recorded (§3.8), not from membership, which does not say who is a person, and no further back than the window reaches: a reset or a forgotten post hides its author as it hides the post.
3.12 Thinking
Private reasoning is any intermediate model computation that is not emitted as user-visible output or as a tool invocation. Such reasoning:
- is never requested as part of an agent’s response;
- is never posted to Mattermost;
- is never included in an approval prompt;
- is never returned by
/collegium traceor/collegium inspect; - is never written to a log line.
It is stored beside the completion that produced it for one purpose: a thinking-mode provider refuses to continue from an assistant message whose reasoning it is not handed back, so it is handed back to the provider and to nothing else, unaltered — and only for the rounds of the turn in progress, since those are the only ones a provider continues from (§5.2). An earlier turn’s reasoning is kept but never replayed, and costs the window nothing (§3.8). No cached prefix is given up for this: a turn’s own rounds follow the message after the window (§3.8), so the next turn’s request departs from that turn’s at that message, before any round whose reasoning it leaves out. Where a provider reports reasoning-token usage, the framework may store the count.
How hard a model reasons is configuration. A model ref may state a reasoning effort in its provider’s own vocabulary, and the framework states it to the provider in that form. Where config states nothing, the provider’s default stands.
3.13 Mail
An agent may act as at most one email address, fixed in its mail settings (§3.4). Which agents can send mail is therefore answerable by reading config alone. The mailbox boundary is enforced by the provider — credentials scoped to that one mailbox — never by the framework asking politely.
Inbound is deterministic code, not an agent noticing. A poll reads the mailbox on a configured interval and records one trigger row per arrival (§4.2); the system bot announces it when the channel is idle. On first connection the mailbox is read to its head and nothing is announced — existing mail is not news. After that nothing is missed, across restarts and extended downtime, and nothing is announced twice. Backlog drains at a bounded rate rather than all at once.
An announcement shows the thread, not a wall. The body is split at the boundaries the sender’s client wrote and rendered as one blockquote: the newest message first under its real headers, then each quoted message under the sender, date, and subject its boundary carried, verbatim. The split is deterministic code, never an agent’s reading. A boundary no rule recognises leaves the text where it is, still quoted, so an unknown client degrades to one segment rather than to dropped text.
Handling an announcement marks the message read, so the same item is not worked twice by a human and an agent. This rides the ordinary triggers::resolve path.
Reading is ungated; sending is gated. Listing, searching, and gathering a conversation return sender, subject, receipt time, whether attachments ride the message, and a short preview — only opening one message returns a body, and a body is never truncated by the mail tool; the one thing that may shorten what the model is first shown of it is the turn’s view of a long result (§3.8), which says so and reads on by reference. A conversation says how many messages it holds. Search takes one phrase, which the provider matches against sender, subject and body; it exposes no query language, so an adapter is not free to offer one. Bodies are presented as readable text whatever the sender’s formatting. Attachments are described — name, type, size — and never opened; a request for their content is a typed refusal. Their absence is stated too, so a message with none is not read as one whose attachments were withheld.
Every send discloses what will leave: the full recipient list, the subject, and the entire body, in the approval prompt (§3.7). Recipients are explicit arguments and there is no bcc field anywhere, so a recipient the approver cannot see is unrepresentable rather than merely disallowed. A reply stays in its conversation because the provider threads it, never because the model managed it. Drafting is not sending: proposing wording in conversation touches no mailbox and needs no approval. A mailbox may name an HTML template in its settings; the body is then rendered from markdown into it, raw HTML escaped, so the approved text is still all the model contributed to what leaves.
A send is never retried. A refusal the server answered — nothing left — is an ordinary result the agent may act on. An outcome that cannot be established is reported as unresolved and ends the turn (§7.1), because feeding “unresolved” back to a model invites exactly the duplicate the rule exists to prevent.
Boot proves every mailbox, and distinguishes what waiting cannot fix. An announcement channel that is a direct message or that the agent does not belong to, and credentials the provider refuses, are boot refusals naming what is wrong. A mailbox that is merely unreachable is not: the system bot says so in the announcement channel and the framework runs, polling until it recovers and saying so again when it does. A provider’s failure mid-session fails the call that met it and never stops the process; a mailbox that stays unreachable is announced as above.
Mail is optional: a deployment with no mailbox configured runs exactly as it did before mail existed, and a partially configured one is refused before the system runs, naming what is missing.
3.14 Plugin
A unit of operator-supplied capability living outside the framework: a directory of TypeScript sources mounted into the deployment, named in config.json, compiled and loaded once when the process boots. The framework performs orchestration and carries no business domain — a deployment’s own concerns belong in plugins, so the framework upgrades without ever occupying the same files as what a deployment added.
A plugin is a toolset (§3.4), and its layout is the declaration. The directory’s name is the plugin’s namespace, its storage scope, and its skills’ qualifier: one identity, stated once. Its config module declares the settings schema agents are configured by and the storage collections it owns; each tool file default-exports one tool named by its basename; each skill directory is a skill in the §3.5 layout. There is no manifest of contributions. A plugin does not alter how the framework activates agents, orders turns, budgets actions, or resolves approvals. Contributions appear under the namespace and agents opt in per grant exactly as with framework capability: installing a plugin grants it to no one. Framework namespaces are reserved.
Refused, never skipped; failures are startup failures. A file the conventions cover either loads or stops the process from starting, naming the plugin and the file. Nothing about a plugin is discovered mid-turn, and nothing malformed is silently left out.
A plugin tool no agent is granted is warned of, not refused. The state is legitimate — a tool may be held back on purpose — but it is also exactly the state a forgotten grant leaves, and nothing else surfaces it.
Plugins are fully trusted; agents are not. A plugin is operator-written code running with framework privilege — installing one is not different in kind from editing the framework, and the safety model constrains what an agent may reach, not what an operator may install. What a plugin decides is whether its tools gate (§3.7) and what the prompt shows; what it does not decide is whether its actions are seen — every plugin tool call is disclosed in the status post and recorded in the trace exactly as a framework tool’s is (A5) — or how the framework budgets actions. A plugin’s tool declares its gate explicitly — a render function, or null — and a tool that declares neither is refused: trust in a plugin is total (§9), which is exactly why what it claims must be legible.
The boundary is structural. A plugin imports @collegium/sdk, zod, and node: builtins — nothing else; every other bare specifier is a boot refusal. Both packages resolve to the framework’s own copies, and each declared range is checked against the deployment’s version at boot. A tool body returns plain text — or text beside a disclosure — and raises the only two failures it controls: invalidArguments, returned to the model, and unresolved, ending the turn as an unconfirmed side effect. A plain throw is a semantic failure like any other (§7.1). What the SDK does not hand over, a plugin cannot touch.
Storage without schema ownership. A plugin persists durable records in the framework’s own store, scoped to its namespace, validated against its declared collection schemas on write and parsed on read — the one qualified read perimeter, because rows may outlive the schema that wrote them. It owns no tables, no migrations, and no database client, and cannot reach the framework’s tables or another toolset’s rows. A record carries the declared shape plus an id, createdAt, and updatedAt the store stamps; a taken id is refused, never overwritten; a patch is merged and the whole parsed again, so no patch leaves an invalid record. Beyond lookup by id, a collection answers one query grammar: an AND of conditions over top-level scalar fields — equality, membership, or a case-insensitive substring — with an optional limit, typed from the collection’s own schema. A storage write can disclose itself by returning a disclosure (§3.4).
Work units are read, never written. Only tasks:: creates or moves a unit, since every change to one is a post (§3.15), so a plugin tracking delegated work binds its records to a unit by reference and takes the unit’s state from the framework rather than keeping its own. A plugin tool finds a unit by reference among those the acting agent created or was assigned in the turn’s channel — another pair’s is simply absent, as it is to tasks::read — and reads its parties, outcome, criteria, context, state, when it was assigned and last changed, the unit it follows, the unit that continues it, and, once it is closed, the state it closed from and what closed it. Its turn names the unit it serves as assignee: the one whose assignment post started it, still assigned when it began, fixed for the turn so a report mid-turn does not unname it. Its turn also says which tools its agent is granted, so a text it renders offers only a tool that agent can call (§3.4), and the handle of the channel it runs in, so a record can name the lane that wrote it; a direct or group message has none. Any other turn names none, even where its agent holds a single assigned unit in the channel. The exhaustion report falls back to that unit (§3.15) because it reports on a turn already over; a running turn that a person or a trigger started may be doing something else, and rows stamped with the wrong unit mislead silently.
3.15 Work Unit
A durable record of one thing one agent handed to another, in one channel. A unit carries the outcome wanted, the criteria its creator will judge the result by, the context the assignee needs, its creator, its assignee, and its state: assigned, blocked, review, done, cancelled. It is created by tasks::assign and by nothing else. There is no unassigned unit, because work nobody handed over is work an agent took on its own initiative (A3). Humans do not create units: a human’s request is a post, and the agent that received it is accountable for it in the ordinary way. A question to a peer is a post too.
Every state change is a post, and the post comes first. The framework publishes the post under the acting agent’s account, from a fixed template carrying model-authored fields — the arrangement §3.2 permits — and only then is the record written, pointing at that post. A change the channel never saw is therefore unrepresentable. A crash between the two leaves a post with no unit, which the next report finds and says, and the creator assigns again — visible where it happened; the reverse order would leave a unit the channel never announced, which is the one disagreement a work unit exists to prevent. An assignment addresses its assignee and a report addresses its creator, so each activates its addressee as any agent’s mention does, once the turn that made it ends or parks (§5.2), and §7.4 counts them as it counts any other mention. A close addresses nobody. A continuation (below) is one post for two changes, the close of the unit it follows and the assignment of the next, and addresses the assignee as an assignment does. Each verb’s result says its post is published and that the colleague’s turn waits for this one to end. A turn may end there with no reply unless it owes one: it owes one where a post it answers (§5.2) is a person’s, a trigger’s announcement, or a plain post of a colleague other than the one its unit post addressed, or where a person steered it. The runner, not a tool, says so on the unit post’s result where the turn owes no reply and no unit a move, since only the runner knows what the turn answers. tasks::assign and tasks::report may instead ask to end the turn at their post. The runner ends it there once every other call of the same completion has returned with neither a refusal nor a failure and none waited on a person, no steer has reached the turn since, and the turn owes no reply and no unit a move. Otherwise the turn goes on, and the model is told why. The ask is the model’s and the gate the runner’s: a model that means to say nothing more says so with the call, since an empty completion is not one it can reliably produce.
The assignee cannot close. tasks::report reaches review and blocked and nothing else; tasks::close reaches done and cancelled and is refused for anyone but the unit’s creator. A report is a claim and a close is a judgement, and one agent never makes both about one unit. A unit in review its creator judges incomplete is handed back as a fresh unit with corrected criteria, or closed as cancelled with the reason; it takes no second report, and a report refused there names the creator it is with and those two ways on. The authority is read off the record: creator and assignee are fields, so no agent holds a role and there is no lead or worker anywhere in the configuration. blocked is for a blocker nobody in the channel can answer, or it is the framework’s report that the assignee ran out of context (below); an assignee held up by a question a person there can answer asks it (§3.7a) and keeps the unit.
A close rests on the report it judges. A unit in review or blocked closes only from a turn that has read its latest report — the report was in the turn’s window, or tasks::read showed it — and any other close is refused with the instruction to read it first, since a verdict on a report its judge never saw judges nothing. Done means the pass is not abandoned: it delivered, or a unit that follows it continues its work, and the verdict says which. Cancelled means the work is abandoned. A blocked unit may close either way. A unit still assigned closes only while no turn of its assignee that opened after the assignment is running in the channel, because that turn may be writing the report the close would pre-empt, and the report would then be refused. An assignee’s posts start its creator only once its turn ends or parks (§5.2), so what this catches is a creator’s turn that a person or a trigger started, or that an assignee parked on a person released. done from assigned stays, for an assignee that answered in a post rather than a report; a person can cancel at any time (§8.4). A report refused because its unit has closed names who closed it, when, and with what verdict, which the unit records on closing, and that a finding still wanted goes to the creator in a post: an assignee told only that its unit moved cannot tell a judgement from a mistake.
A continuation is a fresh unit that follows one. tasks::assign naming a unit in follows is its creator’s alone, from review or blocked, on the terms a close has — its latest report read by this turn — and to that unit’s own assignee. Its one post closes the unit as done, with a verdict naming the unit that continues it, and assigns the next, which records the unit it follows; the next one’s line under Open work and tasks::read show the link, and the unit it closes stops counting against the creator’s cap. Work done in passes is then one unit per pass, each with its own outcome and criteria, rather than one unit whose record describes only its first pass while plain mentions carry the rest; and the next pass costs one call, no more than the mention it replaces. Nothing reopens a unit, and nothing continues one in place. follows may also name a unit of the creator’s already closed as done that nothing follows yet: it links that unit without closing it again, and the unit’s verdict then names its successor, so closing a unit and then continuing it gives the state and the link a continuation alone gives, keeping the creator’s own verdict. A continuation alone records where the unit stood at the link: in review, or blocked and why. Every close records the state the unit closed from and what closed it: its creator’s close, a continuation, or a person’s cancellation. A link to a unit already closed changes neither. A unit is followed once, and a cancelled unit is not continued.
The open units are in the prompt, not in the scrollback. Each turn’s context lists, after the window (§3.8), the units the agent created or was assigned in this channel and has not closed, oldest first, one line each. This is how a supervisor remembers what it handed out and how a stalled unit surfaces, with nothing scheduling and nothing waking (A3). The list is capped by the toolset’s settings; a unit is never truncated, the list is, and the remainder is stated as a count. Because nothing wakes, the section then states, every turn and whether or not a unit is open, that nothing starts another turn of the agent’s here until a person posts, a colleague mentions it, or a trigger fires: the fact sits where a turn decides whether to hand its work on, not only in the preamble, and a close that leaves the agent no other open unit here repeats it in its result. A reference that matches no unit of the agent’s here is refused with the references this section lists, which §7.2 allows since the agent was shown them, and a reference longer than a unit’s is told how a unit is named.
Each line says where the other party stands, as the turn begins. A unit waits on its assignee’s report while assigned and on its creator’s verdict once reported, and the line names the other party by its display name (§3.1), as tasks::read and the tasks results and refusals do, keeping the @handle only where the model must type it, and describes the party awaited: awaiting the reader’s own move since the unit last changed; working here since its turn took the lane, or waiting on a person since the decision it parked on (§3.7, §3.7a), noting a turn begun before the change; its latest turn here since the change ended, and when, without the move; or no turn here since. It is computed from the channel lock, the turn record and the pending decisions whenever it is shown and never stored, so it cannot disagree with them; a time is the operator’s time of day, with the day where that is not today. It reports what the other party’s turns did, never what they called: an agent reads a colleague through its posts alone (§3.8). Why: a unit whose record says only assigned reads the same whether its assignee is mid-turn or stopped without reporting, and a supervisor that cannot tell the two apart either pokes a colleague at work or waits on one that will not answer.
tasks::read shows the unit whole: its outcome, criteria and context, the unit it follows if it continues one, its parties with the other one’s standing beside it, when it was assigned and last changed, and the post of its latest report or close, verbatim and delimited, read by the id the record holds. The report is shown whole rather than bounded like a search hit, because a verdict rests on its words; search finds the same post (§3.8), but a search hit does not count as read for the close that judges it, and a report shown this way does. A post since forgotten (§8.4) is not shown, and the result says so.
A unit whose assignee runs out of context is reported blocked by the framework. When a turn ends as context exhausted while its agent is the assignee of a unit in assigned in this channel, the framework reports that unit blocked through the path tasks::report takes, so the creator is activated. The units are the agent’s units here still assigned whose assignment post is among the posts the turn answers (§5.2), and each is reported, except one the turn handed on to a colleague other than its creator; where none of those posts is an assignment, the agent’s only assigned unit in the channel; where it holds several and none is among them, none is reported, since the framework does not guess which work a turn was doing. The reason is fixed text, not the model’s, because the model has just run out of room. Every other exit leaves the unit assigned and wakes nobody. A denial, a stop or a kill was a person’s act, and waking the creator — whose ruling may be the very action denied — would race whatever that person does next; after a failure the queue waits for a person anyway (§7.1), and a restart resumes nothing (§7.3). The denial notice names each unit the turn was working, chosen as above, and the boot notice names each unit an abandoned turn that had effects was working, and each unit whose report such a turn answered and left awaiting its creator’s verdict; both name its agents without the @.
A turn does not end silently on a unit still assigned. A final reply that mentions no colleague present and no person, read as §4.5 reads mentions and as the model wrote it, from a turn whose agent still holds in assigned a unit it was working, found among the posts the turn answers as above, is rejected once under §4.5, naming the creator and both ways on: a report, which starts the creator’s turn, or, for an interim update or a question, a post that mentions the creator. The same reply sent again posts, and the unit stays assigned. It is not rejected where the turn has addressed a colleague other than the unit’s creator already, whose turn then carries the work on and beside whom a second addressee would be refused (§4.5), nor as the rejection that ends the turn or asks a person for more attempts (§5.3), since the reply it would cost is a valid one. Likewise, a creator’s final reply that reaches nobody, in a turn among whose answered posts is a report of its own unit still awaiting its verdict, is rejected once, naming the unit and the ways on. Each side is sent back at most once per turn. A turn that completes while such a unit is still open, its reminder spent or not allowed, names the unit on its status post, whether or not it replied. A unit the turn handed on, by assigning work to a colleague other than its creator, is not open in this sense: that colleague’s report returns to it. Why: a report reaches its creator on its own, and an interim post reaches nobody unless it names someone; a worker that takes the one for the other leaves its unit assigned, starts nobody and raises no error, and the work stops until a person notices.
Whom an agent hands units to can be declared. An agent’s tasks settings may name the colleagues it assigns to; an assignment to anyone else, a continuation included, is refused before its post, and the refusal names those colleagues by display name with the handle the call takes. The preamble names them too, since they change only with the deployment (§3.8). Boot refuses a declared colleague that is the agent itself, is not a configured agent, or holds no tasks::report and so could never report back. Undeclared, any colleague present that can report may take a unit. Why: Peers lists what each colleague’s toolsets can do, and a supervisor short of a tool finds the colleague that holds it; a grant scoped to the work bounds nothing if delegation can route around it.
An assignment is refused, not stripped, at a loop limit. Where the turn stands at the §7.4 depth or chain-length limit, tasks::assign is refused before it runs and the agent is told why. Stripping the mention would leave a unit assigned to a peer that was never activated.
The unit is not a second approval surface. Assignment, report and close are ungated. A hand-off is not more consequential than the mention it replaces, the assignee’s own actions gate on their own merits, and a gate on the structured path while the plain mention stays free would make the worse path the cheaper one.
4. Activation
4.1 How Work Reaches An Agent
All activation paths are Mattermost posts.
A scheduler calling the runtime directly would look compliant — everything downstream would still land in Mattermost — but the occasion for the work would exist only in process memory: nothing to point at, and two ingestion paths forever.
4.2 Triggers
External events do not post directly. Deterministic code — cron, mail polling, webhook — evaluates its condition and, when it fires, records a trigger: source, target agent, target channel, a reference (sender, subject, ID, and where the source carries one, the full body), and status.
A schedule is declared, never created. An agent’s configuration names its schedules, each carrying a recurrence, a timezone, a channel the agent belongs to, and the text the system bot posts when it fires. Nothing about a firing is model-written: the announcement is the operator’s own words, so it is the system bot speaking as §3.2 requires, and its text is the operator’s instruction, which the agent carries out and then marks done. The turn it starts is gated as any turn is (§3.7), and an approval prompt there names the schedule rather than saying no person asked. A schedule that comes due while the channel is busy is announced when the channel next goes quiet, rather than being dropped or stacked. An agent can no more add a schedule than it can add a tool (§9).
A missed occurrence fires once. A deployment that was down for a week announces one firing rather than seven, and a schedule declared this afternoon does not announce this morning. When a schedule’s next occurrence is announced, the previous one is marked handled if the agent never did, so an outstanding list holds at most one firing per schedule.
A reference carrying a body is disclosed in full, inline while it fits the substrate’s post limit and otherwise as an attached file the post names — the §6.2 rule from the other side. A preview the reader must open the source to complete makes the channel a notification rather than a record.
A source may own part of resolution. Marking a trigger handled runs the source’s own completion first — mail marks the message read (§3.13) — and a failure there leaves the trigger outstanding rather than claiming work is done.
A trigger is posted only when the target channel is idle — no turn running, no approval pending, nothing debouncing. The system bot posts it, mentioning the agent, which starts a normal turn. If the channel is busy the trigger waits and is posted when the channel next goes idle. One post per trigger, and one more where a restart abandoned the turn the post started before it had any effect; after an unclean stop, at most one more (§7.3).
A mail or webhook announcement reports an arrival; it does not instruct. It says the item arrived and asks the agent to read it and say what it needs, then to mark it done. Nothing in it authorises acting on the item’s contents: the text below the heading is outside content (§3.8), and an action that would leave the workspace on its account — a reply to the sender, a send — waits for a person here to ask for it, which the approval prompt for such a call also says (§3.7).
A DM is never a trigger target. The system bot cannot be a member of a DM, so delivery has no route there; intake refuses loudly rather than recording a trigger that can never be posted. Triggers announce unattended work, and unattended work belongs in a channel other humans can see. Nor is a channel the target agent is not in: the trigger would announce work that cannot be done and strand as permanently unresolvable.
Nothing the system bot posts enters the queue (§5.2). Idle-gating is what makes that safe: a trigger is never posted into a channel that cannot immediately act on it. Gating on idle rather than merely on no pending approval matters — an agent mid-model-call is equally unable to take the lock.
Triggers are marked handled by the agent, through a tool. Without this, the outstanding list only grows and the same item is re-posted after every turn. The tool takes the id the announcement brackets, which is the framework’s; given none, it marks the trigger whose announcement started the turn. It never accepts the sender’s reference, which the sender chooses and nothing makes unique. An id that matches no trigger of the agent’s is refused with the ids announced to it and unresolved in that channel, so the refusal points at the right one rather than reading as someone else’s item.
Why this does not violate A1. The trigger record holds things to be announced, not work to be executed. Nothing runs because a record exists; something runs because the system bot posted. For most sources the record is a cache — ground truth is the mailbox — though for webhooks it is the only record, which is an honest exception rather than a hidden one.
4.3 What The Scheduler Decides
Trigger predicates are deterministic code. “Nothing to report” is resolved there, not by the agent — there is no path where an agent wakes, thinks, and stays silent, because that would be activity outside the substrate.
4.4 Folding Fragments Into One Turn
People type in fragments seconds apart. Context is assembled once, at turn start, so a fragment arriving after assembly is invisible to the turn it belongs to. Folding happens in two places, and the second is what makes the guarantee.
Before the turn: a short window. A brief delay, resetting on each message, precedes turn start, and further messages from the same human in the same channel during it are folded into the same turn’s context. This costs nothing: the common case — a mention, then the request a beat later — is absorbed before any work has been done. A message in the window that names the agent is queued before the turn starts, and the turn takes its row (§5.2); one that names nobody is folded and queues nothing.
During the turn: the running turn absorbs. A fragment from the same human arriving while that turn is still on its first model call is handed to the turn rather than queued. The turn discards the completion that only saw the first sentence, re-assembles its context, and calls again. A fold aborts the completion in flight, which it would discard anyway (§7.5). Waiting longer costs a human real time; absorbing costs only a discarded completion.
A fragment addressing nobody is absorbed. A post from the same person that names the agent is absorbed too, and queued (§5.2): it is queued first, and the turn that absorbs it takes its queue row, so absorbing it survives a failure. A post that names a colleague is recorded as ordinary history and costs the running turn nothing: it is a request being made of that colleague, not a sentence being finished here. Each fold is a line on the status post (§8.1), since the completion it discarded was paid for, and the approval prompt then quotes the newest fragment rather than the one the turn began on (§3.7).
Absorption ends at the first action. Once a tool has run or a post exists, discarding would throw away work that already had effects, so the turn stops absorbing and anything later takes the queue path. A fold limit bounds it from the other side, so a human typing steadily reaches an answer instead of paying for one completion per sentence. Past that point the human’s route into the running turn is /collegium steer (§7.5).
Absorption is scoped to the turn’s own author. Unrelated conversation in the channel is recorded and reaches the next turn’s window; it never costs a running turn its completion. Fragments arriving when no turn is absorbing are handled by the queue instead (§5.2), which drains them into one turn.
Folding a post that names nobody queues nothing, so a crash loses it; the post is recorded, but nothing re-activates it. Folding one that names the agent, in the window or during the turn, leaves its queue row to §5.2. The window has a ceiling for the same reason the fold limit exists — without it, a human typing steadily never gets a response.
4.5 Multi-Agent Mentions Are Refused
A post mentioning two or more agents present in the channel starts no turn and enters no queue. The system bot posts a mechanical correction: address one agent per message, and how to name an agent without addressing it — write its name without the @, as Mira rather than @mira. A handle inside code is no mention either, but code is for where the handle itself is meant, not for naming a colleague.
Mentions of agents not present in the channel are inert text: an absent agent never receives the post, so the harm this rule prevents cannot arise. A DM therefore never trips this rule.
Why: two agents working the same task produce two approval prompts for overlapping actions, and the second gets approved having half-read the first. Picking the first-mentioned agent would be arbitrary, since mention order carries no intent.
Why, more consequentially: this rule is what pins delegation width at one. §7.4 bounds the depth of an agent-to-agent chain and says nothing about its branching factor. If a turn could address three peers, ten levels would be tens of thousands of turns and the hourly ceiling would be the only brake. The depth limit bounds total work only because every turn can delegate to at most one colleague.
For agent-authored output the check happens at post time. The post is rejected and fed back to the model as a rejected post, because the model produced valid output violating a framework rule it cannot see, and one retry is cheap.
A reply longer than a post can carry is rejected the same way. A final output over the substrate’s post limit (§6.2) is fed back as a rejected post stating the limit and the reply’s length, and the model answers again more briefly, or replies with the first part and says what remains. It is never truncated: a reply cut to fit reads as a whole reply that says less, which is the quiet degradation A4 forbids. A post a tool returns (§3.15) over the limit is refused as its result.
The bound is per turn, not per post. A turn addresses at most one peer, however many posts it emits: a second post naming a different agent present in the channel is refused exactly as a post naming two of them is, since two addressing posts would produce the two concurrent turns this rule exists to prevent. The same peer addressed twice in one turn is one addressee and one activation: the peer starts once the turn ends or parks, with every post the turn addressed to it in view, each of them queued (§5.2).
A tool call written as prose is rejected the same way, whatever form it takes. A model copies back the transcript line it read in its history, or its provider fails to structure the call and leaves its own markup in the text — DeepSeek’s DSML, stray invoke and parameter tags, a <tool_call> wrapper, a bare call object. Posted, either runs nothing and reads as an action taken. Which markup leaks is the provider’s own, so the inference adapter recognises it behind the vendor seam and hands the turn a leaked call rather than text; the rule the turn applies is one rule for every provider. So is a reply with no prose in it: no letter or digit in any script outside markup tags, or one line repeated five or more times that makes up most of the reply — except an empty ending after a unit post of the turn handed work to a colleague, where the turn owes no reply (§3.15). Neither is a message anyone meant to send.
So, once in a turn, is a reply that reaches nobody while the unit its turn works is still assigned, or while a report it answers awaits its verdict (§3.15). Unlike the rest it reminds rather than forbids: the same reply sent again posts, and it is never the rejection that ends a turn.
Two rejections in a row, then the turn ends with a deterministic notice under the agent’s name. A rejected post spends one action attempt (§5.3). The dedicated count exists because the budget’s bound is the wrong one to rely on here: twenty-five rejected posts would reach the extension prompt with a number that says a turn worked hard rather than that it failed the same way repeatedly. A tool call that runs between two rejections resets the count. Every cause shares the one count — the mentions above, a tool call written as prose, a reply with no prose, output cut at the limit (§7.1), a reply longer than a post, a reply that reaches nobody while its unit is assigned, a call naming a tool the agent does not hold (§7.2) — because a model alternating between them is exhibiting one behaviour. The trace keeps every rejected post and the reason it was refused (§8.3).
Agent mentions in transient status text and in approval and ask prompts (§3.7, §3.7a) are stripped before posting. Neither ever addresses anyone: a prompt is recorded as a prompt, and no deferred hand-off, whether the running turn’s or one a restart recomputes (§7.3), is read off one.
Detection and stripping share one grammar, and that grammar is Mattermost’s. What the framework treats as a mention is what the client highlights and notifies; anything else acts on a mention the human never saw, or ignores one they did. The two paths diverging would defeat both the width bound above and the depth limit in §7.4.
5. Execution Model
5.1 One Turn At A Time, Per Channel
An agent executes strictly one turn at a time within a given channel. While a turn is live — waiting on the model, executing a tool, or blocked on approval — that channel is closed to further turns for that agent. Other channels are unaffected.
The channel is therefore the concurrency unit. It is also the intervention unit (§7.5) and the context unit (§3.8). The cross-channel exceptions are memory (§3.6), search (§3.8), and the global rate ceiling (§7.4).
Only one approval prompt can ever be live for a given agent in a given channel, which is what keeps approval resolution unambiguous.
Within a turn, a completion’s calls run in the order the model made them, except that a run of consecutive concurrent calls (§3.4) runs together. Every call of such a run is admitted to the budget before the run starts, so the extension prompt still blocks in order and never twice at once, and the results are recorded in call order. A gated call is never concurrent. A call that would park a person is not run while the context is over its ceiling (§3.8), and its line says so.
5.2 The Queue
A message addressed to a busy agent is queued, not dropped. The framework acknowledges it with a 👀 reaction on the post — not a reply, because a post per queued message would be noise in a channel where approval prompts also live.
An unaddressed fragment the running turn absorbs (§4.4) is neither queued nor acknowledged: the turn is already answering it. A post that addresses the agent is always queued, however busy the agent is — a colleague’s once its author stops acting (below) — and acknowledged unless the running turn it names absorbed it (§4.4).
Drain, not pop. When a turn ends on an exit that allows progress (§7.1) and the channel goes idle, everything queued for that channel is consumed by a single new turn. Ten fragments queued during a long turn become one turn, not ten.
The posts a turn answers are the post that started it, the posts the debounce batched into it (§4.4), the rows it took from its queue, and the posts folded into it. Every other section that reads this set cites this sentence.
A turn takes what was queued for it before it assembled its context, and answers what it took; a turn a trigger started takes nothing. Each assembly takes the rows queued for the turn’s agent and channel before it began, so a fold’s reassembly takes the folded post’s row; a drain’s first assembly is bounded instead by the moment the drain chose the post it answers, so a person’s post queued after that waits for a turn of its own. Rows left standing in an idle channel wait for a person (§7.1), and a trigger’s turn answering them would answer a person under the trigger’s authority. An exit that allows progress consumes what the turn took, since it is a completion, a person’s act or an exhaustion, which is never retried (§7.1). An exit that does not allow progress returns everything it took to the queue, and queues the post that started it. What was queued after the turn’s last assembly drains.
An agent’s post activates its addressee when the turn that wrote it stops acting, not when it lands. A colleague started by the first post of a turn reads a request its author is still writing, and judges a report before the reply that details it exists; the author’s closing mention then queues a second turn over the same work. So the turn defers the one colleague its posts address (§4.5), and when it ends, on any exit, or parks on a person (§3.7, §3.7a), each of those posts is queued for the colleague and drained like any queue — at once where it is free, after its own turn otherwise, acknowledged as any post waiting behind a busy agent is. One turn then reads everything its author addressed to it. §7.4 admission runs at that drain, on the deferred post: the ceiling, the chain limit, and the depth and return test, read off the turn that wrote it. What this costs is latency: the colleague starts when its author finishes, not when the first post lands. A deferred hand-off lives only in the turn deferring it; a restart recomputes it (§7.3).
A drain answers the newest person who addressed the agent. Where the rows a drain takes hold a post a person addressed to the agent, the newest such post is the one the turn answers, as if it had found the agent idle: the turn is human-initiated (§7.4), the approval prompt quotes that person (§3.7), and their further fragments fold into it (§4.4). Otherwise a person’s request queued behind a colleague’s post would run inside that colleague’s chain, and be refused with it at the chain limit. Only a drain of nothing but colleagues’ posts answers the earliest of them.
A draining turn never continues its own last message. Its window ends on the trace of the turn it drains behind, and a model handed its own message as the last thing said continues it rather than answering; the window ends with a line marking where the agent’s previous turn ended, or that a restart cut it off, so the completion that follows is a new message; the message after the window names the posts the turn answers (§3.8).
The drain is visible even when context is not. When the window cannot reach back as far as the earliest unprocessed post, the draining turn’s status post says how far back context actually reached — detection, not prevention, the same posture as memory-write disclosure (§3.6). What the window could not reach, the agent can search for (§3.8).
The queue holds posts, not content. It holds one row per queued post, per agent and channel: the post’s id and never its text. The content arrives through the normal context path. Delete the queue and it rebuilds from posts, which is why it does not violate A1.
Nothing from the system bot is queued. Triggers are governed by idle-gating (§4.2) instead.
The backlog is unbounded and nothing expires. If approvals are slow, work accumulates behind them indefinitely. The alternative is silently discarding work, and the failure mode of an unbounded queue is visible — a channel that has been busy for a day is obviously busy.
5.3 Action Budget
A fixed number of action attempts per turn, set for the deployment and overridable per agent. An action attempt is one model-emitted tool invocation, including invocations denied before execution. Not counted: framework transport retries, framework posting, and the tools declared budget-exempt (§3.4) — builtins::now, skills::load, memory::read and tasks::read, the exemption being for loading context the framework already holds or a fact it can answer without leaving the process.
On exhaustion the agent posts what it has and requests approval to extend. Approving grants the agent’s budget over again and preserves accumulated context. Extensions are unbounded in number, but each prompt carries the running count, the calls the turn has repeated most, and the agent’s own last words or the fact that it has written nothing, because the human in the loop is the control, and a count alone was approved into loops that had already fetched the same three pages a hundred times each.
Why the budget is per agent: the number is a statement about a role. An agent whose unit of work is a hundred ungated reads either pays that number once, in configuration, or pays it as four extension prompts answered by a human who has stopped reading them, which is the gate degradation A5 names. A declared budget is enumerable by reading config (§6.1) and buys an agent nothing it was not granted: every gated call still blocks. The default stays low because the default agent is a conversational one.
No extension prompt is raised while the turn’s context is over its ceiling (§3.8): an exhausted budget there ends the turn as context exhausted.
Denying an extension ends the turn’s actions, not its voice. A bare denial terminates. A denial with reason feeds the reason back with zero attempts remaining (§3.7), so the agent may conclude in words but not in actions. A tool call emitted after a denied extension ends the turn as budget exhausted and does not prompt to extend a second time, since a second prompt would let a denial buy an unbounded loop. Steering buys words, never budget (§5.4).
Why a ceiling at all, given every consequential call is gated: reads are ungated. An agent can execute forty searches and file reads without touching the approval gate, and the first visible sign is whatever it concluded. Frequent limit-hits are information: the tools are too fine-grained or the task is too large.
5.4 Denial Semantics
Bare denial terminates. Denial-with-reason continues the same turn under the same budget.
Why bare denial is a full stop: a bare “denied” tells the model only that a path is blocked, so it tries an adjacent path — four near-identical proposals refused in sequence, which is the attention burn that destroys the gate (A5). Terminating also makes denial loud: an event you notice and can count, rather than something the model routes around invisibly.
Why denial-with-reason stays inside the turn: a new turn would reset the action budget. Keeping it inside means a human who keeps steering still runs into the ceiling. This is why denials count against the budget.
6. Safety Model
Three layers, in order of precedence.
6.1 Capability And Confinement
Which tools exist for an agent at all, and where they may point. Most agents have no shell; some have no write access; each has only the tools its role requires. Set in configuration, immutable at runtime, fully enumerable without running anything.
For tool-only agents, confinement is enforced inside hand-written tool bodies — path rooting, read-only database connections, network allowlists. The model’s cooperation is irrelevant.
For shell-holding agents it is enforced by OS permissions (A2), which is a stronger boundary because it does not depend on our code being correct.
Shell confinement is OS permissions, not a path check. shell::run runs each command as a dedicated OS user derived from the agent’s username, never as the app’s own user, and two agents can never share one. At boot the framework probes every shell-holding agent and refuses to start if its OS user is not provisioned, or if its workspace is not readable, and only readable, by that user, so a misconfigured host fails loudly rather than on the first command. The approved bytes are the executed bytes: the command runs in an argument slot of its own that no intermediate shell parses, so nothing the approver read as a literal is expanded on the way (§6.2).
Confinement from framework code is traversal, not a path check either. The app root is not traversable by any agent OS user, and the plugin root is beneath it, so framework code and operator-supplied plugin code alike are unreadable to the accounts shell::run executes as.
6.2 Approval
The residual: irreversible, externally-visible, or shell actions block for consent (§3.7).
Known limitation: approval verifies the call, not the content. workspace::write("notes.md", <900 words of confident nonsense>) is a well-formed, in-bounds, correctly-scoped invocation. The tool has no opinion about whether the prose is true.
Two conditions therefore hold, or the gate becomes theatre:
- The read-only floor stays generous. Gating pure observation multiplies prompt volume with zero risk reduction, and volume is what kills the gate. This is why the workspace has read tools beside
workspace::write(§3.4). - The prompt shows the payload in full. A payload nobody can read is a payload nobody is checking. Where a payload exceeds what a single post can carry, the prompt shows a bounded prefix inline and the complete payload as an attachment — the approver sees the exact bytes rather than a rendering of them. A payload shown as code is fenced so that no payload can close its own fence, since text after an early close would render as Markdown, where a link shows its words and hides its target.
The payload limit belongs to the substrate, not to us. Mattermost’s post size is a server setting an administrator can change, so it is read at runtime. Hardcoding it makes the gate silently wrong the day it moves, and raising it is not a remedy: it moves the cliff rather than removing it.
Shell commands are never attached, hidden, or truncated. They are presented inline and in full, and a command too long to present is refused — a shell command that will not fit in a post is itself the signal.
6.3 Reversibility
There is no undo. workspace::write and shell::run act only inside confinement that holds nothing of independent value — the former within the agent’s workspace directory, the latter as a dedicated OS user in its own home — so nothing there can be destroyed. Purpose-built tools that write to real systems are individually reviewed with fixed write targets; recovery on those paths is the underlying system’s problem, not the framework’s.
6.4 Callback Endpoints Trust The Network
The decision, command and trigger endpoints require a shared secret on every request, in addition to trusting the network. The port must still not be publicly routable — bind the loopback or a private interface and let Mattermost reach it over that path; the secrets exist because that binding controls host exposure, not exposure within the deployment’s own network. There are two secrets, separated by the authority each confers. Mattermost’s is carried by the command route and, as a signature over one approval rather than the secret itself, by the decision route, so it is never written into a post. A trigger sender’s is accepted on the trigger route and nowhere else: a webhook integration holds a credential that can announce work and cannot approve an action or stop a turn. It is optional, and its absence closes the route rather than opening it.
7. Failure And Recovery
7.1 Failure Taxonomy
Every way a turn can stop, and what the human sees:
- Normal completion — model emits no tool call, or its unit post asked to end the turn and the runner ended it there (§3.15). Ends. Visible as the final post, or as the unit post that ended a hand-off turn that owed no reply (whether the model then wrote nothing or its unit post asked to end the turn); a hand-off turn that reaches its ceiling after that post ends completed too, unless a unit it owes a move on is still open.
- Denial — human clicks Deny. Ends. Prompt rewritten to a terminal state; the agent posts asking how to proceed, naming who denied which tool and the unit still assigned to it (§3.15).
- Budget exhausted — the agent’s action budget spent (§5.3). Blocks on an approval to extend; denial ends the turn’s actions, leaving it a final word only (§5.3).
- Transport error — timeout, 5xx, rate limit, a completion the provider itself interrupted mid-stream, or one that arrived malformed (§7.2). Retried invisibly; nothing is shown unless retries exhaust or a rate limit asks for a longer wait than the policy allows.
- Output cut at the limit — the provider stopped the completion at its output ceiling, or it ran past the agent’s completion time limit and the framework stopped it. Not output, since the model never finished, and not a failure: it is fed back as a rejected post (§4.5), spending one attempt. A cut that kept nothing says so, and says which limit cut it. A second completion in the same turn cut at the time limit ends the turn, with a notice naming the limit, except after a hand-off that owes nothing and leaves no unit to name (§3.15), where the turn closes as completed.
- Semantic error — a tool that threw, or a completion still malformed once its transport retries are spent (§7.2). Ends immediately. Error posted under the agent’s name. A call whose arguments the schema refuses, a well-formed call whose value the domain refuses, and a call naming a tool the agent does not hold are each a result the model reads instead (§7.2). A call whose argument text never parsed at all is forgiven once and ends the turn here on the second. Output the framework refuses to post, twice in a row, ends the turn here too.
- Side-effect ambiguity — a mutating call times out. Ends, with an explicit statement that completion cannot be confirmed.
- Provider outage — completion fails after retries. Ends. Failure posted under the agent’s name, naming the transport cause as one fixed phrase per cause — connecting timed out, the provider went quiet, the address did not resolve, the connection was refused, the provider rate-limited the request with the wait it asked for — never the runtime’s or the provider’s words (§3.2). The distinction between connecting and responding is the one an operator acts on.
- Provider rejection — the provider refused the request itself: a 4xx other than a rate limit, or a filtered completion. Ends. Never retried, since the same request would be refused again. Failure posted under the agent’s name, naming the class of refusal — a refused credential, an exhausted balance, a permission or moderation refusal, or the status by number — as one fixed phrase per class. A refusal for the request’s length is context exhaustion instead.
- Context exhausted — the turn’s context still exceeds its turn ceiling once every result it has read stands as its line (§3.8), or its model’s window cannot hold it, whether measured by the framework or reported by the provider. Ends. Posted under the agent’s name, stating the context’s size against the ceiling and its largest parts, and pointing at the trace — or, where the starting context itself did not fit, that the context budget and the ceiling disagree, which is configuration rather than anything the turn did. Never retried. The notice names each unit of the agent’s own whose report the turn answered and did not judge; it is not queued again. A turn working a delegated unit also reports that unit blocked (§3.15), and the report says it cannot show what the turn wrote. The queue drains, since exhaustion is a fact about one turn’s accumulation, not about the channel.
- Delivery failure — the chat substrate refused a post the turn had to make. Ends, carrying the substrate’s own reason. This is not a provider outage: naming the wrong system sends the reader to the wrong place. Where there is no post to point at at all, the failure is loud in the operational record even though the channel stays silent.
/collegium stop— human command. Ends at the next iteration boundary. Stop notice posted./collegium kill— human command. Ends immediately; an in-flight completion is aborted, an in-flight tool may still complete.- Global halt — hourly ceiling breached. All agents stop; prominent post; requires
/collegium resume. - Restart — deploy or crash. All in-flight turns abandoned; one system-bot notice in the main channel.
In every case the channel lock is released. The queue drains into a fresh turn only when the exit allows progress — normal completion, denial, budget exhaustion, context exhaustion, /collegium stop, /collegium kill. After a provider outage or rejection, semantic error, side-effect ambiguity, or delivery failure — and while a global halt stands — the queue is left standing, with what the failed turn took returned to it and the post that started it queued (§5.2): a fresh turn would inherit the same failure, and a drain loop bounded only by the hourly ceiling would halt the whole framework over one dead provider. A standing queue drains at the next human post, the next idle trigger flush, or the boot//collegium resume sweep — all human-visible moments. A human post that arrived while the failed turn ran counts as that next post, so the queue drains the moment the lock is released. One drain per human post keeps this bounded — a fresh turn that fails again has nothing new behind it and stands. A peer’s mention earns no such drain; it waits for a human. There is no retry timer: a slow retry loop is still the graceful degradation A4 rejects.
7.2 Retry Policy
Retry the transport, never the intent.
Transport errors (timeout, 5xx, rate limit, a completion interrupted mid-stream) are retried transparently: a bounded count with backoff, invisible to the model, not counted against budget. A rate limit that says how long to wait is waited out as asked, up to the policy’s longest single wait; one asking for longer ends the turn at once as a provider outage naming the wait, because a wait that long is the silent stall A4 rejects. A request whose turn has been killed is never retried.
A malformed completion is retried as transport. A stream chunk that does not parse, a call with no id or no name, a reply carrying neither text nor a call, except after a unit post that handed work on, where the runner decides it (§3.3): each is a delivery the provider got wrong, not an intent the model formed, and nothing has acted on it. A completion that is still malformed once the retries are spent ends the turn as a semantic error. The argument text of a single call is the one exception, forgiven inside the turn instead (below).
Every completion streams, under an idle timeout. The inference timeout bounds how long the provider may send nothing, not how long the completion takes. A model that thinks for ten minutes and keeps streaming is served, up to the agent’s completion time limit; one the provider has stopped serving is cut and retried. Every completion also runs under that wall-clock limit, set per agent. The idle timeout cuts a provider that has gone quiet; the limit cuts one that keeps a completion going past what a lane can afford to wait, whatever keeps it going. It is set far above a productive completion, so it bounds a lane’s stall, not the model’s thinking. A completion it cuts reports no usage, so its event carries an estimate of what it had streamed, marked as one and counted in no total.
A tool that throws ends the turn. Its body failed in a way it did not return as a result, so whether it acted is unknown.
Arguments the schema refuses are answered, not ended on. A call for a tool the agent holds, whose arguments parse but fail the tool’s schema, returns one result naming what the schema refused, and spends an attempt (§5.3). The call reaches neither its gate nor its body. Every reshaped call is validated, gated and billed as before, so no sequence of guesses reaches anything the agent was not granted, and the budget is what stops a model that never converges. An argument a tool does not declare is ignored, and the result says so, unless the tool’s schema refuses undeclared keys, as a tool whose arguments are a patch should. A refused size states the size received, not only the limit. A size bound is stated in the parameter’s description and enforced when the call is parsed; it is never sent to the provider as a schema bound, because a provider that enforces one while decoding cuts the value to fit and the cut value then passes (A4).
A name that resolves to no tool is answered, not ended on. A call naming a tool the agent does not hold, under every spelling §3.4 admits, returns one result: the name it used, and the tools it can call. It spends an attempt and counts toward §4.5’s count of rejections in a row. Misnaming a tool is the ordinary error of a long context, and ending the turn over it ended whole supervised runs over a typo; resolution is a lookup over the agent’s own grants, so no guess calls anything it was not already granted.
A value the domain refuses is a result too. A well-formed call whose value the domain refuses — a description over its cap (§3.6), a trigger not addressed to this agent (§4.2), a path outside the workspace (§6.1), a skill outside the agent’s manifest (§3.5) — is returned as the tool result, the turn continuing on its remaining budget. A refused invocation is still an invocation and still costs an attempt.
Argument text that never parsed is not a shape the model chose. A call whose arguments did not parse as JSON at all is a byte-level accident, not a model guessing at an interface. It returns one result — that the arguments were not valid JSON and the call did not run — and spends an attempt. The second such call in the same turn ends it: one is a mistake corrected; two says this call’s arguments are not surviving the trip. The tool body never sees the broken arguments, and the provider’s raw text is never read back to the model.
A refusal never enumerates accepted values the agent was not already shown. A domain refusal names what was refused and never lists what would have passed, because a valid set echoed back on failure turns a boundary into something to guess at. A schema refusal may quote the tool’s own schema back, and the unknown-name result lists the tools the agent holds: each restates what the model was already given.
Never retry a call that may have committed a side effect. If mail::send times out, we do not know whether the mail went. The turn terminates and posts the ambiguity explicitly. This is why retryable is a per-tool declaration: reads yes, mutations no.
7.3 Restart
Nothing resumes. All in-flight turns are abandoned; every agent boots idle.
Why not resume: tool calls are not idempotent, so re-running a turn risks sending the same email twice. Checkpointing instead produces intentions formed against a stale world — an approval clicked at 6pm executing a plan assembled at 9am.
On boot:
- The status post of every abandoned turn is closed, under the agent’s own name: its working line becomes an abandoned one, so the turn’s own post says which work the restart cost. A turn that called no tool leaves nothing behind.
- Pending approval prompts are invalidated. The post is edited to a dead state and its buttons removed.
- Backfill runs per channel from the last recorded post forward (§8.2).
- The roster reconciles against Mattermost (§3.11).
- The system bot posts one boot notice in the main channel, stating the downtime window, that in-flight work was abandoned, and how much of it was queued again — the turns that had no effects, their posts returned, and the hand-offs the abandoned turns deferred (below) —, naming each post not queued again, with one line for each unit still assigned to an agent whose abandoned turn had effects — the unit a report on exhaustion would have chosen (§3.15) — and one for each unit whose report an abandoned turn with effects answered and left awaiting its verdict. The line names the unit, its channel, its creator and its assignee, and activates no one.
The downtime window is measured against the process, never inferred from activity. A clean shutdown records its stop time; a crash leaves only a periodic liveness stamp, and the notice then says since last known alive — honest about which of the two it is. Deriving the window from the last observed post would report a quiet channel as a day of downtime.
Boot also verifies what it is about to trust. For every model a configured agent names, one minimal completion is issued through its provider before the framework serves Mattermost events, and boot refuses if the provider rejects it as unauthorized, naming the model and the agents it strands. An answer that is neither acceptance nor refusal within a short deadline is recorded as unverified, never as verified, because a check that could not run must not read as a check that passed (§A4).
Queue state and outstanding triggers both survive a restart, so pending work is not lost to a deploy — it drains once channels go idle.
An abandoned turn that had no effects is queued again. A turn that posted no reply, parked on no person, and made only calls that never ran or are declared safe to run again (§3.4) changed nothing, so the posts it took from its queue are returned and its triggering post queued, and the boot sweep drains them into a fresh turn that answers the same posts. A turn a trigger started sends its trigger back to be announced again, since nothing from the system bot is queued (§5.2). This is not resumption: the fresh turn assembles its context from the channel as it stands after the restart. A turn that had effects is abandoned as before, since a fresh turn could not see what it called, and what it took from its queue is consumed. Where the abandoned turn had made a completion and the process did not stop cleanly, the turn may be what took the process down, so each such post, and each such trigger, is returned at most once; a post or trigger not queued again is named in the boot notice. Without this, a deploy landing between a delegation and the worker’s first action leaves the delegation answered by nobody.
A colleague an abandoned turn deferred is queued too. The deferred hand-off is not persisted but recomputed from the store, which already records who wrote each post and when every turn started: for each abandoned turn, the colleague its posts addressed (§4.5) — its replies, work-unit posts and notices, by the rule the running turn applied, and never its prompts — read against the reconciled roster, has each of those posts that no turn of that colleague in the channel has started since queued, and the boot sweep drains it. A colleague that has run since read what it was sent. Without this, a deploy landing while a turn that had just handed work over was still running leaves the hand-off answered by nobody.
7.4 Loop Control
Agent-to-agent mentions make unbounded chains possible: each turn is individually well-behaved and under budget while the chain runs until someone notices.
Two counters bound them, both carried in the turn record and never shown to the model. Depth measures how far work has been handed down from a human; chain length measures how many turns one human post has produced at all. Both are read off the post a turn answers, which for a drain is the one §5.2 names.
Depth counter:
- Human-initiated turn → depth 0
- Trigger-initiated turn → depth 1 (a cron is not a human; unattended work is the dangerous kind)
- Agent-initiated turn → parent depth + 1, unless the mention is a return, in which case the depth of the turn being returned to
- At the configured limit, agent mentions in output are refused and the turn posts visibly that it has reached the delegation limit and someone needs to pick this up.
A return is an answer to the turn that asked. A mention is a return when the post carrying it was authored by a turn whose own triggering post was authored by a turn of the agent now being activated: Sam asks Naomi (Naomi at depth 1); Naomi mentions Sam (a return: Sam back at depth 0). Had Naomi mentioned Omar instead, that is a hand-off, and Omar is at depth 2. Depth therefore counts nesting, not exchanges — a supervisor that delegates one unit at a time and verifies each result sits at depth 0 for the whole run, its worker at depth 1. The test is mechanical and read off the store; no configuration declares a supervisor, so the shape can never be claimed by an agent or drift from what the channel shows. The same discipline governs who may close a unit of delegated work (§3.15).
Chain length:
- Human-initiated turn → 1, and the turn is the root of its chain
- Trigger-initiated turn → 1, and the turn is the root of its chain
- Agent-initiated turn, return or hand-off → parent chain length + 1, carrying the parent’s root
- At the configured limit, counted as the turns carrying one root and enforced when a turn is opened, an agent’s mention in output is refused first, as the friendly stop — and a hand-off through a work unit (§3.15) is refused before it runs — the turn posting visibly that the chain has reached its limit. A turn that would still exceed it is not opened, and the system bot says so in the channel, naming the agent that was not activated; a fresh human post starts a fresh chain.
Chain length is what returns make necessary. With returns free of depth, two agents could answer each other indefinitely at depths 0 and 1, and the only brake would be the hourly ceiling — an emergency stop, not a design. The chain-length limit is the bound on total unattended work one human post may set in motion: continuing past it costs exactly one human-visible act, which is the property every other limit in this section has. The number is deliberately far below the hourly ceiling, so that a legitimate long run is refused by its own limit rather than halting every agent in the framework.
Depth bounds the chain only because §4.5 bounds its width. Every turn may address at most one peer, so ten levels is ten turns, not ten levels of a branching tree.
Chain length is counted by root, not along a path, because a chain can still have two turns running at once — an author parked on a person has already released the colleague it addressed (§5.2), and both run once the person answers — and two checks made at output time could both pass at the limit; admission is the guarantee, and the refusal at output is the early, agent-voiced stop. A turn with no recorded root is its own root: a restart does not continue a chain (§7.3).
Enforcement is in the framework, not the prompt. Prompt-level constraints are advisory, and advisory constraints are what produced Hermes.
Global ceiling: a configured number of turns per hour, framework-wide. Breach halts all agents, posts prominently, and requires an explicit /collegium resume. A halt invalidates pending approval prompts, as a restart does. While a halt stands, queues do not drain and triggers are not flushed — a trigger posted into a halted channel would strand. On clearing, /collegium resume runs the same drain-and-flush sweep boot performs.
/collegium resume refuses only what it can objectively re-check. A §3.10 topology violation is a fact about membership, so a halt raised by one stands until membership is fixed. A ceiling halt clears on the human’s authority: a fresh allowance begins. The halt exists to stop a chain and put a human in front of it; once they have looked, their judgement is the control the system was routing to.
Accepted cost: a loop that trips the ceiling is handed a fresh allowance every time someone resumes it. The compensating controls are that the halt post is prominent, the resume is attributable, and the running count is in front of whoever clicks.
The rolling window is durable; the halt flag is not. Turn starts are counted from the store over the trailing hour, and a /collegium clear (§8.5) keeps every turn record, so a clear is not a way to reset the count. The halt not surviving a restart is acceptable only because a restart breaks the loop that raised it; the ceiling is re-evaluated at the first turn start after boot and re-raises the halt if the window is still full, so a crash-looping instance cannot grant itself a fresh allowance every boot.
7.5 Manual Intervention
All three commands are channel-scoped. Stop and kill apply to every agent in the channel; a steer reaches one.
/collegium stop aborts current turns at the next iteration boundary. The honest guarantee is no further tool calls, not nothing happened.
/collegium kill abandons current turns immediately: the turn record is closed and the channel lock released. A tool already in flight may still complete and its side effect may still land — /collegium kill is for a wedged process, and it accepts that ambiguity in exchange for immediacy.
/collegium steer [{agent}] {text} hands one instruction to one agent’s running turn in the channel, read before that turn’s next completion. It is what §4.4 folding cannot be: folding discards a completion and reassembles from scratch, which is affordable only before the turn has acted. Steering keeps the work already done and appends — the instruction arrives as the human speaking, prefixed with their name as a post would be — but discards a completion in flight, since a plan made before the correction is exactly what the human is correcting. A steer aborts a completion in flight, which it would discard anyway, so a correction reaches a long completion at once. A steer spends one action attempt, so a human who keeps steering runs into the same ceiling a denial with a reason does (§5.3). The turn names the steer and its author on its status post, and the trace keeps the text. A tool already running is not interrupted, and a turn parked on an approval is steered by denying with a reason (§5.4), not by this command.
A steer reaches one agent, and never guesses which. The first word names the agent only where it is an agent in the channel, with or without its @, so a steer may open with any other word. Named, the steer goes to that agent’s turn alone. Unnamed, it goes to the one agent running here; with more than one running it is refused, naming them, and reaches none. The response names the agent reached. With nothing to reach, whether nothing runs here or the named agent does not, the invoker is told so and the text is discarded rather than queued; §5.2 remains the one durable path for work.
Why refuse rather than reach every turn: a steer carries no addressee, since it arrives as the human speaking, so one written for one agent reads to every other as an instruction to it. A correction meant for all is steered once per agent, or is a stop.
Why a command rather than a post: §5.2 says an addressed post is always queued and acknowledged, and that rule is load-bearing. A command is not a post: its response is ephemeral, it re-activates nobody, and it creates no obligation the framework then has to honour.
/collegium stop and /collegium kill post visibly, naming the agents whose turns they reached; a steer is ephemeral to the invoker and visible through the status post. A turn a command ended closes its status post naming who issued it (§8.1), so a stopped turn is told apart from one that stopped itself, and a call the command cancelled is marked as one that did not run. Any human in the channel may issue any of the three, since stopping is always safe and steering is no less safe than posting. Neither clears the queue: whatever was pending drains into the next turn.
A command’s own response is ephemeral, so it never enters a channel as a post and cannot re-activate the agent it just interrupted. The visible notice is separate: posted by the system bot where it is present, and under the agent’s own account in a DM, which the roster knows the channel to be (§3.11); a DM the system bot could reach is still the agent’s to speak in (§3.2).
Either command also resolves a pending approval in the channel as cancelled: the prompt is rewritten to an invalidated terminal state, exactly as a restart does, and no follow-up posts. A cancellation is not a denial — denial statistics keep meaning that a human refused an action. This is how /collegium stop reaches a turn parked on an approval, which would otherwise be untouchable for days.
7.6 Stalls Are Announced
Every exit in §7.1 posts, or was caused by a human watching. Two shapes end a run with nothing said at all, and leave the channel reading exactly like work in progress or work finished. The system bot names each in the channel it happened in. It announces and never clears: a notice that acted would be the recovery A4 rejects, and the human reading it is the one who decides.
- A standing queue. An agent has had posts queued for a channel (§5.2) with no turn of its own running there for longer than a configured threshold. The notice names the agent and says what clears it: a post addressing the agent, and
/collegium queue {agent}shows what waits. - A long turn. One turn has held its channel’s lock (§5.1) for longer than a configured threshold since it started or last waited on a person. Time parked on an approval or a question is not counted: that wait is its prompt’s and its status post’s to announce (§8.1), and its remedy is an answer, not a kill. The notice names the agent and how long, says whether the turn has called a tool yet, and says that
/collegium killends a turn whose status post shows no progress. A turn that has traced nothing has no status post yet (§8.1), so the sweep opens one before it speaks: the notice then points at a post that exists, and the minutes that follow are accounted for somewhere durable.
Each is one notice per episode, re-armed only once its condition clears, so a standing condition is said once, not every minute it stands. A long turn is re-armed once more when a post queues behind it (§5.2): work waiting is a new fact about an old episode, and the second notice says so. Nothing is announced while a global halt stands (§7.4): its own post already says why nothing moves. An agent is named by its display name, never its @, as every system-bot notice names one (§3.2). In a DM the notice is posted under the agent’s own account, as §3.2 permits.
Why notices and not timeouts: each of these was a common way a long run died silently. A timeout that killed the long turn would also kill the legitimately long one, and a sweep that drained the standing queue would be the retry timer §7.1 refuses. Neither threshold is a limit the agent can see or work around: A3 is untouched, since nothing here starts a turn. The completion time limit (§7.2) is not such a timeout: it bounds one completion, set far above a productive one, and the turn goes on after it.
8. Persistence And Observability
8.1 What The Human Sees
One status post per turn, edited in place as the turn progresses. Tool calls are appended to it as they occur, so the post accumulates a readable trace of the turn rather than only showing current state. The post as a whole is bounded by the substrate’s post limit: when the trace outgrows it, the oldest lines give way behind one line that says how many. A run of identical lines is one line with a count, so a loop reads as one line rather than two hundred.
Each line names the tool and what the call is doing — the URL navigated to, the path written, the command run — because a column of bare tool names says a turn was busy without saying what it did. Each tool renders its own one-line summary; the line is capped in length and the untruncated arguments are always in /collegium trace. A line states the call’s disposition where it was not plain success — denied and by whom, refused, timed out, cancelled by a command or a halt (§7.5) — or what the call came to, such as the page a click landed on, a status a fetch got that was not success, or a search that matched nothing, since a denied write rendered like a completed one is a record that lies. A line whose mark says the call did not run states its subject and not its effect: no size, count or duration is attributed to an action nobody took. A line states a result the model was shown only in part: how much of the whole, and that the rest is by reference. The closing line also states how long the turn ran, and who ended it where a person did: who issued the command (§7.5), or who denied the action (§5.4).
While the turn waits on a person, the head says so and since when: 🔐 _waiting on a decision since 14:05 EDT_ for an approval, ❓ _waiting on an answer since 14:05 EDT_ for a question (§3.7, §3.7a), in the operator timezone. It turns back to working when the decision lands; a turn parked on several at once names the earliest still open. It is one edit, not a reminder: nothing re-posts or chases (§3.7). /collegium units and /collegium trace state the same wait (§8.4).
Why not stream every tool call as a separate post: a ten-call turn would produce ten posts of machinery around one post of substance, and approval prompts live in the same channel — noise in the supervision channel degrades the gate (A5).
A memory write, its evictions, and a memory delete appear here as ordinary tool-call lines; what was written, evicted, or removed is in /collegium trace (§3.6). Beside the calls sit the framework’s own notes on the turn: a steer and who gave it (§7.5), a fold that started the turn over (§4.4), and a drain whose earliest post the window could not reach (§5.2).
Queued messages are acknowledged with a 👀 reaction (§5.2). This and the typing indicator below are the only signals the framework emits without posting.
A typing indicator shows while the model is generating, and through the §4.4 window, which is otherwise the one stretch where an agent has committed to answering and nothing says so. It creates no post and nothing durable. It is deliberately dark during tool execution and while an approval is pending: the status post and the approval prompt already account for that time, and an indicator held through a human’s deliberation would claim work that is not happening.
Why this and not an eagerly-created status post: a turn that calls no tool should leave nothing behind but its reply (A5), yet the human who addressed the agent is owed some sign that it heard them. An indicator that expires on its own satisfies both, for a turn of ordinary length. Past the §7.6 threshold the post is opened anyway: a turn that has held a channel that long owes something durable, and its closing line then records how it ended.
8.2 What The Store Holds
The store is authoritative for conversation content. Context assembly reads only the store, never the Mattermost API on the turn path.
Why a second copy at all: an agent’s context is posts interleaved with tool calls, tool results, approval requests and decisions — its own turns’ trace, never a peer’s (§8.3). None of that exists in Mattermost except as rendered text, and split across two stores every context assembly becomes a merge across different clocks and ID spaces. One ordered store is a material simplification.
Stored: every observed post and the files it carried by name, type and size, what the post is to the turn that wrote it — a reply, a work-unit post (§3.15), a notice, a prompt or a status post (§4.5) — and whether it is pinned now; every tool call and result, with how the model read a result it did not read whole (§3.8); every approval request and decision, question and answer (§3.7a), steer (§7.5), and output refused as a post with the reason (§4.5); the memories, each with how often and when it was last revised (§3.6); the triggers, the queue state, the work units (§3.15); and per-turn metadata — what activated the turn and the post a drain began from (§5.2), depth, chain length, action count, model, token usage and cost, written after each completion so an abandoned turn still counts what it spent, and the context it last assembled: when, how far back its window reached, and what the budget charged for it (§8.3); and, on each completion’s event, what the provider reported for it and which upstream served it. The turn’s row stays the total; a completion cut at its time limit, by a steer or by a fold reports nothing, so its event carries an estimate of what it had streamed, which no total includes (§7.2, §7.5, §4.4). The file bytes are not stored: Mattermost holds them.
Backfill on boot is required. Posts made while the process was down are absent from the store. Each channel is backfilled from its last recorded post forward, per agent, using per-agent tokens — never a privileged token, or the framework would import posts from channels the agent has no membership in. Backfilled posts never trigger a turn but are context-eligible, as history rather than missed requests.
Repair path: slash commands typed in the channel rather than a SQL console, so repair stays visible and attributable.
The one edit the store follows is a pinned post’s. A pinned post is a standing instruction (§3.8), and one revised in place that read as its first draft would be the stale copy pinning exists to replace. Mattermost reports a pin, an unpin and an edit as the same edit over the socket; for a pinned post the store records the text as it stands, and for any other only that it is not pinned. A deleted post is no longer pinned. Boot and a resync read each channel’s pinned posts afresh, through each agent’s own token as backfill does, and record a pinned post the store never saw — pinned from before the agent’s history — as a post, unless it is older than the channel’s last clear (§8.5). No pin ever starts a turn.
Accepted losses: editing a post that is not pinned does not correct what an agent believes; deleting a post in the client does not redact it from context — /collegium clear (§8.5) is the one deletion the framework performs and honours; a late-joining agent has no channel history before its join point, and search (§3.8) cannot find what was never stored.
8.3 Trace
The complete tool trace — every call, arguments, and results — is retrievable via /collegium trace {post-id}, where the post is named by its id or by its permalink. The response is ephemeral, visible only to the invoker, because trace output contains file contents and email bodies verbatim, and everyone in a channel can also approve agents.
The trace opens with the turn’s own record, because a reader who cannot see why a turn started or what it read misreads what it did. What started it: a post addressing an idle agent, its queue drained as its previous turn here ended, a colleague’s turn that addressed it ending or parking (§5.2), a trigger (§4.2), posts a reconnect recovered, or the boot or resume sweep (§7.3) — with the post it answers and the post a drain began from, and every post it answers (§5.2), whether a steer reached it, and, for a turn that handed work on, whether it owed a reply (§3.15), and, where a unit post asked to end the turn and the turn went on, why. When it started, how long it ran, and over how many actions. The context it last assembled: when, how far back the window reached, and what the budget charged for it (§3.8). The tokens the provider reported — the prompt with the share its cache served, the completion with its reasoning — and the cost. Each event beneath states its offset from the start, and a completion’s event what the provider reported for it, which upstream served it, and whether a relief pass had edited the prompt since the completion before it (§3.8) — beside the event rather than on a line of its own, so no reader takes one completion for the turn. A result carries the mark its status-post line shows (§8.1), and where the model did not read it whole — shown as a view of its first part, or replaced by its line once read (§3.8), with its reference; an event from before views keeps its cut — says so beside the whole output. A post rejected under §4.5 is kept with the reason it was refused.
None of it enters the window. The record lives on the turn’s row rather than in its events, because the events are the agent’s own trace and replay into its later windows (§3.8); a rejected post is an event the window skips, since the model was told why and answered again.
A trace carrying those payloads runs into the same substrate limit §6.2 does, and takes the same answer: where it exceeds what a single post can carry, it is delivered as an attachment, still ephemeral.
Reading traces is how the tool inventory gets tightened over time.
8.4 Command Surface
Every command is a subcommand of one slash command, /collegium, so typing /collegium offers the whole surface with an argument hint and a line of help per subcommand:
/collegium trace {post-id}— full tool trace for a turn, headed by what started it, what it read and what it cost (§8.3); a turn still waiting on a person says on what, since when, and its prompt. Ephemeral./collegium forget {post-id}— remove a post from agent context, and from every queue it waits in, since a post forgotten is one not to act on. Posts./collegium clear [--memories]— delete every post in this channel and every agent’s record of them, after a confirmation;--memoriesalso deletes the memories their turns here wrote (§8.5). Posts./collegium reset {agent}— mark an episode boundary; posts pinned here still stand (§3.8). Posts./collegium stop— abort current turns in this channel at the next boundary. Posts./collegium kill— abandon current turns in this channel immediately. Posts./collegium steer [{agent}] {text}— hand one instruction to an agent’s turn running in this channel, read before its next model call; a call in flight is made again. The agent may go unnamed only while it is the one running here (§7.5). Ephemeral, naming the agent reached; the turn names it on its status post./collegium resume— clear a global halt./collegium approvals [{agent}]— every approval and every question (§3.7a) still waiting on a human, in the channels you are in, oldest first, each naming the agent, the action, the head of a question’s words, its age, and its prompt. Ephemeral. The channel filter is the same membership check that decides who may answer one (§3.7). It decides nothing — the buttons on the prompt post remain the only way to answer./collegium queue {agent}— show whether a turn holds the agent’s lane here (§5.1), since when, the post that started it, and its status post or that it has none yet — a turn waiting on a person shows as parked, on what and since when (§8.1); then how many posts wait in its queue and the oldest of them. Ephemeral. A turn that has called no tool has posted nothing (§8.1), so this is where a human learns that a post addressing the agent would now queue behind it (§5.2)./collegium queue {agent} clear— discard the posts waiting in the agent’s queue here, so the next drain does not run work a configuration change made stale. Posts. The posts themselves stay; only their queue rows go, and a running turn keeps what it took./collegium triggers {agent}— list outstanding triggers. Ephemeral./collegium memory {agent}— inspect and prune an agent’s memories, each listed with how often and how lately it was revised (§3.6). Ephemeral./collegium units {agent}— list an agent’s open work units in this channel as its prompt lists them, with age, state and where the other party stands (§3.15), then each party to them — the agent or a counterpart — whose turn here waits on a person, on what and since when (§8.1), which the lines above leave to this list. Ephemeral./collegium units {agent} cancel {reference}— close a unit as cancelled on a human’s authority, for one whose creator will never reach it. Posts./collegium inspect {agent}— show an agent’s model, its context budget, its turn ceiling with the view it implies (§3.8) and its completion time limit (§7.2), tools (marking which need a human on every call), skills, schedules with their next occurrence in the operator timezone, and the prompt a turn in this channel would be given: the system prompt, then the message that follows the window (§3.8). Ephemeral./collegium usage— show token usage per agent and model, with cached-prompt and reasoning breakdowns and the cost the provider charged where it reports them, over turns that ended in the last 24 hours in any channel, abandoned ones included, then how many completions a time limit or a steer cut short and about how many tokens they spent unreported. Ephemeral.
A bare /collegium, or a subcommand nothing declares, answers the invoker with the list above.
The command is held by a Mattermost plugin the framework ships, and the framework declares its subcommands to that plugin at boot. The plugin knows nothing of the subcommands: it holds /collegium for a team, forwards every execution to the framework, and relays the answer. The framework’s command definitions are the one source. Boot fails loudly if the plugin is not installed or the declaration is refused: a framework that starts without its stop switch is worse than one that does not start, and the failure would otherwise surface only as an opaque client error at the moment someone reaches for /collegium kill.
8.5 Clearing A Channel
/collegium clear gives a channel a fresh start: every post in it is deleted from Mattermost, and every agent’s record of it is deleted from the store. It is what clear is in a terminal, with one difference the substrate forces — a terminal forgets nothing, and this forgets everything, because the agents read the store and a human who sees an empty channel is entitled to assume the agents see one too. It applies to every agent in the channel, in a channel or a DM alike, and any member may run it, since membership is the whole authority model (§3.7). It is the one deletion the framework performs and honours.
Content goes, accounting stays. Deleted: the posts, pinned ones among them, the trace events of every turn here, their approvals and asks, the episode boundaries, the queue rows of the erased posts, the work units (§3.15), and the triggers that already posted here. Kept: the turn records, with what they cost and how long they ran, so the §7.4 hourly count and /collegium usage read exactly as before. Kept too: the files agents wrote, plugin storage, mail, schedules, the global halt, and triggers still pending, which post into the cleared channel when the idle gate opens, since they are future work rather than history. Search from any channel no longer finds what was here, because it is gone rather than hidden.
Memories are kept by default. Memory is the agents’ work product, not the channel’s record, and the path by which knowledge deliberately crosses channels (§3.6); the human saw each write when it was kept. --memories deletes those written from turns in this channel, selected by provenance — exact for an entry written and revised here, approximate at the edges, since a revision carries the revising turn’s provenance. /collegium memory {agent} remains the way to prune one on purpose.
A confirmation dialog stands before it. It states what goes and what stays, in two sentences. Every other command runs on a keystroke because every other command is safe or reversible; this one is neither.
It refuses while any turn runs here, including one parked on a decision. It does not stop or kill: §7.5 gives those two different guarantees, and the human chooses between them. On confirmation it takes every agent’s channel lock (§5.1), so no turn can start until the store is cut.
The notice is the boundary. It is posted first, by the system bot where present and under the agent’s own account in a DM (§7.5), and if it cannot be posted nothing is changed. Everything older than it goes; a post that lands during the clear is newer than the notice and survives in both stores. The store is cut before the posts are deleted from Mattermost, and a restart backfills from the notice forward (§8.2).
Failure is stated, never repaired. A failure after the store is cut leaves a channel the agents have already forgotten and the human can still see — the honest direction — and the notice says so and how many remain. What remains stays forgotten: the store never again records a post older than the channel’s last clear, so a pin on one, or its pin read afresh at a restart (§8.2), does not bring it back. A failure before it leaves everything as it was, and the notice says that. In every case /collegium clear again is the repair: it finds nothing older than its own new notice in the store and removes what remains in the channel.
Why deletion and not a boundary for everyone: a /collegium reset for every agent would hide the same content and delete nothing, and the store would hold what the channel no longer shows. “Cleared” has to be true of the database, or the command lies.
9. Deliberate Non-Goals
No self-modification of instructions. Agents cannot write skills, system prompts, tool definitions, schedules, channel configuration, or model selection, and cannot pin a post (§3.8). Memory is the sole exception. A schedule is in that list for the same reason a tool definition is: it is the agent’s own activation surface, and A2’s guarantee that an agent’s capabilities are enumerable by reading config extends to when it acts, not only to what it may do. An approval would not repair this — every other approval in the system authorises one call, where a schedule authorises an unbounded number of future turns.
Fixed model per agent, no fallback chain. A provider outage means the affected agents are dead for the duration, failing loudly.
No queue bounds and no expiry. Backlog grows without limit and nothing goes stale. Accepted for a system with a small number of internal users; revisit if a channel is ever genuinely swamped.
No plugin ecosystem, and no plugin sandbox. No registry and no discovery: what loads is named in config.json and mounted by the operator, and nothing is fetched. @collegium/sdk is published so a plugin can be written in a repository of its own, and boot refuses a plugin whose declared SDK range the deployment does not satisfy. No isolation between framework and plugin: trust is total and deliberate (§3.14). Plugins do not depend on, extend, or communicate with one another; there is no lifecycle beyond startup — no hot reload, no enable/disable at runtime. Nothing in the system lets an agent write, install, configure, or enable a plugin — the prohibition on self-modification extends here unchanged.