Part II — THE EXECUTION LAYER
4.Cogs: Accountable AI Workers
the unit at which AI work becomes assignable, governable, and auditable. Assumes: Section 3 (Frames).
Organizations do not employ "intelligence." They employ workers: bounded roles with a job description, defined tools, defined permissions, and a name that can appear in an audit trail. This is not bureaucratic habit. It is how accountability is manufactured. "Who did this, under what authority, with access to what?" only has an answer because work is assigned to defined units rather than to a general pool of capability.
AI, as commonly deployed, has no such unit. A model endpoint answers whatever it is asked by whoever reaches it, and the question "what did the AI do last quarter?" dissolves into a pile of undifferentiated API calls. The whitepaper's execution layer starts by restoring the unit. It calls the unit a Cog.
Anatomy
Cog — the atomic unit of AI work
Model
- general or specialized weights
Context
- Frames (base)
- + retrieved docs
- + memory slices
- + task instructions
Harness / Skills / Tools / APIs
- what it runs in and may call
Governance
- data access,
- allowed actions,
- approval rules
A Cog is drawn as one bounded unit containing four parts: a model (possibly specialized), a context founded on one or more Frames that orient the model to the organization, the harness, skills, tools, and APIs the Cog runs in and may call, and governance parameters covering what data it may access, what actions it may take, and what requires human approval. The harness is an implementation detail chosen by the Cog’s builder, not something its consumer has to understand. The Cog, not the bare model, is the unit at which AI work becomes auditable and governable.
View as text diagram
┌─────────────────────────────────────────────┐
│ COG │
│ ┌─────────────┐ ┌──────────────────────┐ │
│ │ Model │ │ Context │ │
│ │ (general or │ │ Frames (base) │ │
│ │ specialized │ │ + retrieved docs │ │
│ │ weights) │ │ + memory slices │ │
│ └─────────────┘ │ + task instructions │ │
│ └──────────────────────┘ │
│ ┌─────────────┐ ┌──────────────────────┐ │
│ │ Harness, │ │ Governance │ │
│ │ skills, │ │ data access, │ │
│ │ tools, APIs │ │ allowed actions, │ │
│ └─────────────┘ │ approval rules │ │
│ └──────────────────────┘ │
└─────────────────────────────────────────────┘The paper is emphatic that the second compartment is where Cog quality is actually decided. A Cog's working context is assembled at invocation time: its Frames provide the durable, governed foundation, around which the system layers retrieved documents, relevant slices of organizational history (the memory substrate Section 7 describes), recent conversation, and task-specific instructions. Deciding what the model sees, in what order, under what constraints, is a first-class engineering discipline, and the paper gives it the section's pull quote:
"A well-constructed Cog is, in large part, a well-managed context."
The harness, and why the Cog sits above it
The fastest-moving part of the agent ecosystem now gets a passage of its own, and the earlier revisions needed one: the harness, meaning the software wrapped around model weights that turns them into something able to act. The paper's inventory is "agent loops, tool-calling frameworks, graph-centric orchestration layers, memory and retrieval scaffolds, skill libraries, evaluation hooks," and it observes that new ones arrive monthly, from model labs' reference implementations to community projects.
The interpretive move is the interesting part.
So the harness sits inside the Cog, as an implementation detail the Cog's consumer does not have to understand: "Cog builders choose the harness that fits the work; Cog consumers get a versioned, installable, auditable unit with declared tools and permissions, whatever runs inside." The paper's landscape section files today's agent frameworks the same way, as "a capability, not a governance model" and "a natural layer inside a Cog" (§9). Whether the framing is generous or merely convenient depends on something not yet demonstrated: a Cog built on one harness and a Cog built on another have to be composable by the same Op without the Op caring, and no published specification yet says what a Cog must expose for that to hold. The paper names the requirement; the interface is on the list of things still being built.
The workforce argument follows from it, and it is where the paper says why any of this scales. Cogs modularize agentic capability into something installable and replaceable, specialize it (the paper's formula: "weights + harness + skills + tools + Frames tuned to a kind of work"), and let many parties build independently — "a vertical specialist can build a Contract Review Cog on one harness while a community builds an Invoice Extraction Cog on another, and an Op composes both without caring which is which." The failure case is stated as plainly: "Without the Cog abstraction, every harness is a silo and every agent is bespoke."
What the unit buys you
The payoff is the restored question. Instead of "what did the model do?", an organization can ask: what did this Cog do, with what inputs, under which Frames, and what was the outcome? Every noun in that sentence is now a versioned, inspectable thing. The paper's claim, which regulated industries will recognize as the actual bar for deployment, is that this specificity is what makes AI workers deployable where auditors live. Section 2's evidence base carries over directly: the NBER mechanism (codified best practice, distributed to every worker) is precisely what a Frame-oriented Cog operationalizes, with the addition that the codification is now governed rather than folkloric.
Where agents fit
Readers will have noticed that a Cog sounds adjacent to what the industry calls an AI agent. Earlier revisions of the paper handled the word defensively, calling it overloaded and risk-laden and then moving on. Revision 9 deletes that treatment and replaces it with what is now, on this guide's reading, the strongest section in the document. It opens by conceding the industry's win outright:
Employment is the organizing metaphor, and it earns its place because it generates questions rather than merely renaming things. Who does this agent work for? Whose context and policies govern it? Which tools and data may it touch, and under whose identity? How much autonomy has it been granted, and by whom? Who checks its work before the work has consequences? What record remains when it is done? Every one of those is a question an organization already knows how to answer about a person and mostly cannot answer about a piece of software acting on its behalf.
The three-way split. Earlier editions of this guide described the architecture as unbundling "agent" two ways, into workers and workflows. Revision 9 makes it three, and the third is the one that matters most:
- the capability is a Cog — installable, versioned, and auditable before it runs;
- the engagement is an Op — the goal, the autonomy budget, the checkpoints, the validation strategy;
- the continuing actor is held by the Hub itself: identity from the Hub's own directory, credentials brokered and scoped rather than copied, memory in Local and Organizational Memory, and history in Tracks, which together are the things "that make an agent feel like a colleague rather than a function call."
The compressed definition is the one to carry: an agent is "a Cog engaged through an Op, given identity and memory by the Hub." And the paper is unusually direct about why the industry fuses the three into one word: "fused, they are how vendors capture instance state and why governance cannot be tailored. Kept separable, agents stay ownable, auditable, and portable." That is a commercial accusation, not just an architectural preference, and it is the load-bearing reason the paper gives for keeping the three apart.
The employment mapping then falls out cleanly. Frames are the context the agent is oriented by and the policies it is accountable to. A Cog is the qualified hire, vetted before deployment. An Op is the assignment, with deliverable and checkpoints defined. Guards check the work. Gates define what the paper calls the autonomy budget, "the precise boundary between what the agent may do alone and what requires a human." Tracks are the personnel record. Guards, Gates, and Tracks are named here because the employment mapping needs them; Section 6 is where they are taught. The Hub is the workplace.
Autonomy as a grant, not a property. The consequence is that autonomy is "granted rather than assumed — and granted per engagement, not per platform." A drafting Op may run nearly unattended; a payments Op may require approval at every consequential step. The paper reads Gartner's warning that uniform governance across agents leads to failure as support for exactly this, and calls the Gate "autonomy tailoring made operational." It is a genuinely useful reframing: most enterprise agent policy today is written at the platform level, where it is either too loose for the payments case or too tight for the drafting one.
The evidence, named. The section is also where the paper stops arguing structurally and starts naming things. Cloudflare OS (an agent workspace with organizational context and skills, policy-enforcing "Gatekeepers," and a persistent app platform) is described as validating the category, and then used to draw the distinction this section turns on: openness and sovereignty are different properties. The code is Apache-2.0 and can run on an open-source runtime, "but the system is designed for one vendor's network, with self-hosting an escape hatch whose tooling Cloudflare itself describes as still in progress" (Cloudflare, August 2026). Alongside it sits the session-portability evidence Section 1 already borrowed. The paper's conclusion from the pair:
"Agents are the hands; the Hub is the employer of record."
The argument is now falsifiable in a way the older competitive framing was not: if a vendor ships production-grade self-hosting with portable session state, the paper's distinction narrows to a preference. And the guide's own earlier position, that agent frameworks are components of this architecture rather than competitors to it, is now the paper's position too, stated more precisely, which is the correct outcome and worth recording as such rather than quietly absorbing.
The governance context makes the design choice legible:
The unit is deliberately incomplete: deciding what happens with a Cog's output, sequencing it alongside other work, and placing humans at the right checkpoints all live one layer up, in the Op.