## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
39 KiB
Paperclip Specification
Target specification for the Paperclip control plane. Living document — updated incrementally during spec interviews.
1. Company Model [DRAFT]
A Company is a first-order object. One Paperclip instance runs multiple Companies. A Company does not have a standalone "goal" field — its direction is defined by its set of Initiatives (see Task Hierarchy Mapping).
Fields (Draft)
| Field | Type | Notes |
|---|---|---|
id |
uuid | Primary key |
name |
string | Company name |
createdAt |
timestamp | |
updatedAt |
timestamp |
Board Governance [DRAFT]
Every Company has a Board that governs high-impact decisions. The Board is the human oversight layer.
V1: Single human Board. One human operator.
Board Approval Gates (V1)
- New Agent hires (creating new Agents)
- CEO's initial strategic breakdown (CEO proposes, Board approves before execution begins)
- [TBD: other governance-gated actions — goal changes, firing Agents?]
Connection tool reviews also appear in task history, with a composer takeover for human approval, decline, or scoped remembered permission. Connections and task views resolve the same review, and the agent continues with the server-recorded outcome. See the implementation contract.
Model authentication failures also surface a provider-specific Connections card on the task immediately after failure. Users can reconnect inline and resume; legacy agents keep their authentication until an explicit, validated adoption. See AI Connections.
AI connections also expose an on-demand, credential-scoped usage probe. It reports provider windows, scoped exhaustion, reset times and overage observations for downstream decisions, with explicit unknown/unsupported/error states. Probing does not change runner routing or enforce provider limits automatically.
Board Powers (Always Available)
The Board has unrestricted access to the entire system at all times:
- Set and modify Company budgets — the Board sets top-level token/LLM cost budgets
- Pause/resume any Agent — stop an Agent's heartbeat immediately
- Pause/resume any work item — pause a task, project, subtask tree, milestone. Paused items are not picked up by Agents.
- Full project management access — create, edit, comment on, modify, delete, reassign any task/project/milestone through the UI
- Override any Agent decision — reassign tasks, change priorities, modify descriptions
- Manually change any budget at any level
The Board is not just an approval gate — it's a live control surface. The human can intervene at any level at any time.
A Board status inquiry does not itself pause unfinished task execution. Native ordinary tasks that report blocking remaining work must continue, register a real wait, or surface a bounded recovery failure. Recorded approvals, questions, dependencies, and pauses remain authoritative; obsolete requests must not replay.
Budget Delegation
The Board sets Company-level budgets. The CEO can set budgets for Agents below them, and every manager Agent can do the same for their reports. How this cascading budget delegation works in practice is TBD, but the permission structure supports it. The Board can manually override any budget at any level.
Future governance models (not V1):
- Hiring budgets (auto-approve hires within $X/month)
- Multi-member boards
- Delegated authority (CEO can hire within limits)
Open Questions
- External revenue/expense tracking — future plugin. Token/LLM cost budgeting is core.
- Company-level settings and configuration?
- Company lifecycle (pause, archive, delete)?
- What governance-gated actions exist beyond hiring and CEO strategy approval?
2. Agent Model [DRAFT]
Every employee is an agent. Agents are the workforce.
Agent Identity (Adapter-Level)
Concepts like SOUL.md (identity/mission) and HEARTBEAT.md (loop definition) are not part of the Paperclip protocol. They are adapter-specific configurations. For example, an OpenClaw adapter might use SOUL.md and HEARTBEAT.md files. A Claude Code adapter might use CLAUDE.md. A bare Python script might use command-line args.
Paperclip doesn't prescribe how an agent defines its identity or behavior. It provides the control plane; the adapter defines the agent's inner workings.
Agent Configuration [DRAFT]
Each agent has an adapter type and an adapter-specific configuration blob. The adapter defines what config fields exist.
Paperclip Protocol (What Paperclip Knows)
At the protocol level, Paperclip tracks:
- Agent identity (id, name, role, title)
- Org position (who they report to, who reports to them)
- Adapter type + adapter config
- Status (active, paused, terminated)
- Cost tracking data (if the agent reports it)
Adapter Configuration (Agent-Specific)
Each adapter type defines its own config schema. Examples:
- OpenClaw adapter: SOUL.md content, HEARTBEAT.md content, OpenClaw-specific settings
- Process adapter: command to run, environment variables, working directory
- HTTP adapter: endpoint URL, auth headers, payload template
Exportable Org Configs
A key goal: the entire org's agent configurations are exportable. You can export a company's complete agent setup — every agent, their adapter configs, org structure — as a portable artifact. This enables:
- Sharing company templates ("here's a pre-built marketing agency org")
- Version controlling your company configuration
- Duplicating/forking companies
Context Delivery
Configurable per agent. Two ends of the spectrum:
- Fat payload — Paperclip bundles relevant context (current tasks, messages, company state, metrics) into the heartbeat invocation. Suited for simple/stateless agents that can't call back to Paperclip.
- Thin ping — Heartbeat is just a wake-up signal. Agent calls Paperclip's API to fetch whatever context it needs. Suited for sophisticated agents that manage their own state.
Minimum Contract
The minimum requirement to be a Paperclip agent: be callable. That's it. Paperclip can invoke you via command or webhook. No requirement to report back — Paperclip infers basic status from process liveness when it can.
Integration Levels
Beyond the minimum, Paperclip provides progressively richer integration:
- Callable (minimum) — Paperclip can start you. That's the only contract.
- Status reporting — Agent reports back success/failure/in-progress after execution.
- Fully instrumented — Agent reports status, cost/token usage, task updates, and logs. Bidirectional integration with the control plane.
Paperclip ships default agents that demonstrate full integration: progress tracking, cost instrumentation, and a Paperclip skill (a Claude Code skill for interacting with the Paperclip API) for task management. These serve as both useful defaults and reference implementations for adapter authors.
Export Formats
Two export modes:
- Template export (default) — structure only: agent definitions, org chart, adapter configs, role descriptions. Optionally includes a few seed tasks to help get started. This is the blueprint for spinning up a new company.
- Snapshot export — full state: structure + current tasks, progress, agent status. A complete picture you could restore or fork.
The usual workflow: export a template, create a new company from it, add a couple initial tasks, go.
3. Org Structure [DRAFT]
Hierarchical reporting structure. CEO at top, reports cascade down.
Agent Visibility
Full visibility across the org. Every agent can see the entire org chart, all tasks, all agents. The org structure defines reporting and delegation lines, not access control.
Visibility settings on an agent profile (where supported) do not alter company-level visibility for tasks, projects, issues, comments, costs, or activity. Those work-object privacy controls are not a V1 feature until centralized scoped authorization is in place.
Each agent publishes a short description of their responsibilities and capabilities — almost like skills ("when I'm relevant"). This lets other agents discover who can help with what.
Cross-Team Work
Agents can create tasks and assign them to agents outside their reporting line. This is the mechanism for cross-team collaboration. These rules are primarily encoded in the Paperclip SKILL.md which is recommended for all agents. Paperclip the app enforces the tooling and some light governance, but the cross-team rules below are mainly implemented by agent decisions.
Task Acceptance Rules
When an agent receives a task from outside their team:
- Agrees it's appropriate + can do it → complete it directly
- Agrees it's appropriate + can't do it → mark as blocked
- Questions whether it's worth doing → cannot cancel it themselves. Must reassign to their own manager, explain the situation. Manager decides whether to accept, reassign, or escalate.
Manager Escalation Protocol
It's any manager's responsibility to understand why their subordinates are blocked and resolve it:
- Decide — as a manager, is this work worth doing?
- Delegate down — ask someone under them to help unblock
- Escalate up — ask the manager above them for help
Request Depth Tracking
When a task originates from a cross-team request, track the depth as an integer — how many delegation hops from the original requester. This provides visibility into how far work cascades through the org.
Billing Codes
Task detail keeps hierarchy separate from creation provenance: the Tasks tab shows all subtasks and, independently, work created from the current task grouped by project or No project. A created subtask may appear in both sections. Creation provenance follows the originating run equally for legacy and native runners.
Tasks carry a billing code so that token spend during execution can be attributed upstream to the requesting task/agent. When Agent A asks Agent B to do work, the cost of B's work is tracked against A's request. This enables cost attribution across the org.
Open Questions
- Is this a strict tree or can agents report to multiple managers?
- Can org structure change at runtime? (agents reassigned, teams restructured)
- Do agents inherit any configuration from their manager?
- Billing code format — simple string? Hierarchical?
4. Heartbeat System [DRAFT]
The heartbeat is a protocol, not a runtime. Paperclip defines how to initiate an agent's cycle. What the agent does with that cycle — how long it runs, whether it's task-scoped or continuous — is entirely up to the agent.
Execution Adapters
Agent configuration includes an adapter that defines how Paperclip invokes the agent. Built-in adapters include:
| Adapter | Mechanism | Example |
|---|---|---|
process |
Execute a child process | python run_agent.py --agent-id {id} |
http |
Send an HTTP request | POST https://openclaw.example.com/hook/{id} |
claude_local |
Local Claude Code process | Claude Code heartbeat worker |
codex_local |
Local Codex process | Codex CLI heartbeat worker |
opencode_local |
Local OpenCode process | OpenCode heartbeat worker |
pi_local |
Local Pi process | Pi CLI heartbeat worker |
cursor |
Cursor API/CLI bridge | Cursor-integrated heartbeat worker |
openclaw_gateway |
OpenClaw gateway API | Managed OpenClaw agent via gateway |
hermes_local |
Local Hermes process | Hermes agent heartbeat worker |
The process and http adapters ship as generic defaults. Additional built-in adapters cover common local coding runtimes (see list above), and new adapter types can be registered via the plugin system (see Plugin / Extension Architecture).
An adapter's selected execution engine is part of its permission and session contract. Missing prerequisites or engine failures must be surfaced without silently launching a different engine. A default local engine must support normal task work and control-plane coordination; explicit operator restrictions remain authoritative.
Adapter Interface
Every adapter implements three methods:
invoke(agentConfig, context?) → void // Start the agent's cycle
status(agentConfig) → AgentStatus // Is it running? finished? errored?
cancel(agentConfig) → void // Graceful stop signal (for pause/resume)
This is the full adapter contract. invoke starts the agent, status lets Paperclip check on it, cancel enables the board's pause functionality. Everything else (cost reporting, task updates) is optional and flows through the Paperclip REST API.
What Paperclip Controls
- When to fire the heartbeat (schedule/frequency, per-agent)
- How to fire it (adapter selection + config)
- What context to include (thin ping vs. fat payload, per-agent)
What Paperclip Does NOT Control
- How long the agent runs
- What the agent does during its cycle
- Whether the agent is task-scoped, time-windowed, or continuous
Pause Behavior
When the board (or system) pauses an agent:
- Signal the current execution — send a graceful termination signal to the running process/session
- Grace period — give the agent time to wrap up, save state, report final status
- Force-kill after timeout — if the agent doesn't stop within the grace period, terminate
- Stop future heartbeats — no new heartbeat cycles will fire until the agent is resumed
This is "graceful signal + stop future heartbeats." The current run gets a chance to land cleanly.
Open Questions
- Heartbeat frequency — who controls it? Fixed? Per-agent? Cron-like?
- What happens when a heartbeat invocation fails? (process crashes, HTTP 500)
- Health monitoring — how does Paperclip distinguish "stuck" from "working on a long task"?
- Can agents self-trigger their next heartbeat? ("I'm done, wake me again in 5 min")
- Grace period duration — fixed? configurable per agent?
5. Inter-Agent Communication [DRAFT]
All agent communication flows through the task system.
Model: Tasks + Comments
- Delegation = creating a task and assigning it to another agent
- Coordination = commenting on tasks
- Status updates = updating task status and fields
Low-trust agents can create self-assigned tasks and subtasks within their
permitted scope, subject to assignment permissions and the responsible user's
authority. Created work retains containment. An authorized user's direct message
in their own Agent Chat may authorize an agent to edit its own AGENTS.md;
outside work and subtasks do not inherit this authority. Permission failures
should name the rejected action and the specific restriction.
There is no separate messaging or chat system. Tasks are the communication channel. This keeps all context attached to the work it relates to and creates a natural audit trail.
Experimental Agent Chat presents one persistent task per person and agent as a simplified conversation. Chat has a searchable secondary sidebar with agent avatars that lists every eligible agent, conversations first; adding an agent starts or reopens their single conversation. It retains the task composer, transcript, tools, attachments, documents, and existing Subtasks panel, with ordinary company visibility. New execution tasks are ordinary project tasks, not children of the conversation. Idle conversations wait for a message without entering execution-task work queues. Agents clarify goals here and create assigned tasks for substantial execution. /new resets provider context at an ordered session boundary within the same task while preserving visible history. enableAgentChat is disabled by default; the V1 lifecycle and rollout contract is specified in SPEC-implementation.md.
Question recipients
Ordinary Agent Chat questions use the server-owned conversation recipient.
Task questions may optionally name a particular user or agent. Explicit user
recipients must be valid and authorized in the company before a question is
saved. See SPEC-implementation.md §9.8.1 for the resolver contract.
Implications
- An agent's "inbox" is: tasks assigned to them + comments on tasks they're involved in
- A human's Mine inbox and its badge include failed runs attributed to that human, not another user's runs. All retains company-wide failure visibility. Historical unattributed runs remain in the local single-user board's Mine view; see
SPEC-implementation.mdfor the routing contract. - The CEO delegates by creating tasks assigned to the CTO
- The CTO breaks those down into sub-tasks assigned to engineers
- Discussion happens in task comments, not a side channel
- If an agent needs to escalate, they comment on the parent task or reassign
Task Hierarchy Mapping
Full hierarchy: Initiative (company goal) → Projects → Milestones → Issues → Sub-issues. Everything traces back to an initiative, and the "company goal" is just the first/primary initiative.
6. Cost Tracking [DRAFT]
Token/LLM cost budgeting is a core part of Paperclip. External revenue and expense tracking is a future plugin.
Cost Reporting
Fully-instrumented Agents report token/API usage back to Paperclip. Costs are tracked at every level:
- Per Agent — how much is this employee costing?
- Per task — how much did this unit of work cost?
- Per project — how much is this deliverable costing?
- Per Company — total burn rate
Costs should be denominated in both tokens and dollars.
Billing codes on tasks (see Org Structure) enable cost attribution across teams — when Agent A requests work from Agent B, B's costs roll up to A's request.
Budget Controls
Three tiers:
- Visibility — dashboards showing spend at every level (Agent, task, project, Company)
- Soft alerts — configurable thresholds (e.g. warn at 80% of budget)
- Hard ceiling — auto-pause the Agent when budget is hit. Board notified. Board can override/raise the limit.
Budgets can be set to unlimited (no ceiling).
Open Questions
- Cost reporting API — what's the schema for an agent to report costs?
- Dashboard design — what metrics matter most at each level?
- Budget period — per-day? per-week? per-month? rolling?
7. Default Agents & Bootstrap Flow [DRAFT]
Bootstrap Sequence
How a Company goes from "created" to "running":
- Human creates a Company and its initial Initiatives
- Human defines initial top-level tasks
- Human creates the CEO Agent (using the default CEO template or custom)
- CEO's first heartbeat: reviews the Initiatives and tasks, proposes a strategic breakdown (org structure, sub-tasks, hiring plan)
- Board approves the CEO's strategic plan
- CEO begins execution — creating tasks, proposing hires (Board-approved), delegating
Default Agents
Paperclip ships default Agent templates:
- Default Agent — a basic Claude Code or Codex loop. Knows the Paperclip Skill (SKILL.md) so it can interact with the task system, read Company context, report status.
- Default CEO — extends the Default Agent with CEO-specific behavior: strategic planning, delegation to reports, progress review, Board communication.
These are starting points. Users can customize or replace them entirely.
Default Agent Behavior
The default agent's loop is config-driven. The adapter config contains the instructions that define what the agent does on each heartbeat cycle. There is no hardcoded standard loop — each agent's config determines its behavior.
This means the default CEO config tells the CEO to review strategy, check on reports, etc. The default engineer config tells the engineer to check assigned tasks, pick the highest priority, and work it. But these are config choices, not protocol requirements.
Paperclip Skill (SKILL.md)
A skill definition that teaches agents how to interact with Paperclip. Provides:
- Task CRUD (create, read, update, complete tasks)
- Status reporting (check in, report progress)
- Company context (read goal, org chart, current state)
- Cost reporting (log token/API usage)
- Inter-agent communication rules
This skill is adapter-agnostic — it can be loaded into Claude Code, injected into prompts, or used as API documentation for custom agents.
8. Architecture & Deployment [DRAFT]
Deployment Model
Single-tenant, self-hostable. Not a SaaS. One instance = one operator's companies.
Development Path (Progressive Deployment)
- Local dev — One command to install and run. Embedded Postgres. Everything on your machine. Agents run locally.
- Hosted — Deploy to Vercel/Supabase/AWS/anywhere. Remote agents connect to your server with a shared database. The UI is accessible via the web.
- Open company — Optionally make parts public (e.g. a job board visible to the public for open companies).
The key constraint: it must be trivial to go from "I'm trying this on my machine" to "my agents are running on remote servers talking to my Paperclip instance."
Agent Authentication
When a user creates an Agent, Paperclip generates a connection string containing: the server URL, an API key, and instructions for how to authenticate. The Agent is assumed to be capable of figuring out how to call the API with its token/key from there.
Flow:
- Human creates an Agent in the UI
- Paperclip generates a connection string (URL + key + instructions)
- Human provides this string to the Agent (e.g. in its adapter config, environment, etc.)
- Agent uses the key to authenticate API calls to the control plane
Tech Stack
| Layer | Technology |
|---|---|
| Frontend | React + Vite |
| Backend | TypeScript + Express (REST API, not tRPC — need non-TS clients) |
| Database | PostgreSQL (see doc/DATABASE.md for details — PGlite embedded for dev, Docker or hosted Supabase for production) |
| Auth | Better Auth |
Concurrency Model: Atomic Task Checkout
Tasks use single assignment (one agent per task) with atomic checkout:
- Agent attempts to set a task to
in_progress(claiming it) - The API/database enforces this atomically — if another agent already claimed it, the request fails with an error identifying which agent has it
- If the task is already assigned to the requesting agent from a previous session, they can resume
No optimistic locking or CRDTs needed. The single-assignment model + atomic checkout prevents conflicts at the design level.
Agent @-mentions provide context without waking agents or changing task ownership. New work requires explicit assignment, delegation, or a review request; ordinary issue comments can still wake the current assignee.
Releasing a terminal task clears execution locks while preserving its assigned
owner and final status. Assignment remains part of the work history after Done
or Cancelled. Releasing unfinished work still relinquishes the agent assignment;
only an active in_progress task returns to todo.
Human in the Loop
Agents can create tasks assigned to humans. The board member (or any human with access) can complete these tasks through the UI.
When a human completes a task, if the requesting agent's adapter supports pingbacks (e.g. OpenClaw hooks), Paperclip sends a notification to wake that agent. This keeps humans rare but possible participants in the workflow.
The agents are discouraged from assigning tasks to humans in the Paperclip SKILL, but sometimes it's unavoidable.
API Design
Single unified REST API. The same API serves both the frontend UI and agents. Authentication determines permissions — board auth has full access, agent API keys have scoped access (their own tasks, cost reporting, company context).
No separate "agent API" vs. "board API." Same endpoints, different authorization levels.
Work Artifacts
Paperclip manages task-linked work artifacts: issue documents (rich-text plans, specs, notes attached to issues) and file attachments. Agents read and write these through the API as part of normal task execution. Full delivery infrastructure (code repos, deployments, production runtime) remains the agent's domain — Paperclip orchestrates the work, not the build pipeline.
Users may start a task with only a prompt. Paperclip uses a short prompt slice as its initial title and asks the assigned agent to name the task early. Explicit user titles are preserved, and naming does not change task execution state.
Task work mode is explicit persisted state. Requesting a plan in a title or description does not switch the task into planning mode. Standard execution may produce a plan as its requested deliverable; explicit planning mode separately governs plan-only execution and its approval transition.
Open Questions
- Real-time updates to the UI — WebSocket? SSE? Polling?
- Agent API key scoping — what exactly can an Agent access? Only their own tasks? Their team's? The whole Company?
Crash Recovery: Manual, Not Automatic
When an agent crashes or disappears mid-task, Paperclip does not auto-reassign or auto-release the task. Instead:
- Paperclip surfaces stale tasks (tasks in
in_progresswith no recent activity) through dashboards and reporting - Paperclip may perform bounded continuity repair with the same assigned agent; when that is exhausted or unsafe, it opens a board-owned recovery action without waking a substitute agent
- Paperclip does not fail silently — the auditing and visibility tools make problems obvious
- Recovery is handled by humans or by emergent processes (e.g. a project manager agent whose job is to monitor for stale work and surface it)
Principle: Paperclip reports problems, it doesn't silently fix them. Automatic recovery hides failures. Good visibility lets the right entity (human or agent) decide what to do.
Plugin / Extension Architecture
The core Paperclip system must be extensible. Features like knowledge bases, external revenue tracking, and new Agent Adapters should be addable as plugins without modifying core. This means:
- Well-defined API boundaries that plugins can hook into
- Event system or hooks for reacting to task/Agent lifecycle events
- Agent Adapter plugins — new Adapter types can be registered via the plugin system
- Plugin-registrable UI components (future)
The plugin framework has shipped. Plugins can register new adapter types, hook into lifecycle events, and contribute UI components (e.g. global toolbar buttons). A plugin SDK and CLI commands (paperclipai plugin) are available for authoring and installing plugins.
9. Frontend / UI [DRAFT]
Primary Views
Each is a distinct page/route:
- Org Chart — the org tree with live status indicators (running/idle/paused/error) per agent. Real-time activity feed of what agents are doing.
- Task Board — Task management. Kanban and list views. Filter by team, agent, project, status.
- Dashboard — high-level metrics: agent count, active tasks, costs, goal progress, burn rate. The "glance" view from GOAL.md.
- Agent Detail — deep dive on a single agent: their tasks, activity, costs, configuration, status history.
- Project/Initiative Views — progress tracking against milestones and goals.
- Cost Dashboard — spend visualization at every level (agent, task, project, company).
Board Controls (Available Everywhere)
- Pause/resume agents (any view)
- Pause/resume tasks/projects (any view)
- Approve/reject pending actions (hiring, strategy proposals)
- Direct task creation, editing, commenting
10. V1 Scope (MVP) [DRAFT]
Full loop with one adapter. V1 must demonstrate the complete Paperclip cycle end-to-end, even if narrow.
Must Have (V1)
- Company CRUD — create a Company with Initiatives
- Agent CRUD — create/edit/pause/resume Agents with Adapter config
- Org chart — define reporting structure, visualize it
- Process adapter — invoke(), status(), cancel() for local child processes
- Task management — full lifecycle with hierarchy (tasks trace to company goal)
- Atomic task checkout — single assignment, in_progress locking
- Board governance — human approves hires, pauses Agents, sets budgets, full PM access
- Cost tracking — Agents report token usage, per-Agent/task/Company visibility
- Budget controls — soft alerts + hard ceiling with auto-pause
- Default agent — basic Claude Code/Codex loop with Paperclip skill
- Default CEO — strategic planning, delegation, board communication
- Paperclip skill (SKILL.md) — teaches agents to interact with the API
- REST API — full API for agent interaction (Express)
- Web UI — React/Vite: org chart, task board, dashboard, cost views
- Agent auth — connection string generation with URL + key + instructions
- One-command dev setup — embedded PGlite, everything local
- Multiple Adapter types (HTTP, OpenClaw gateway, and local coding adapters)
Not V1
- Knowledge base - a future plugin
- Advanced governance models (hiring budgets, multi-member boards)
- Revenue/expense tracking beyond token costs - a future plugin
- Public job board / open company features
11. Knowledge Base
Anti-goal for core. The knowledge base is not part of the Paperclip core — it will be a plugin. The task system + comments + agent descriptions provide sufficient shared context.
The architecture must support adding a knowledge base plugin later (clean API boundaries, hookable lifecycle events) but the core system explicitly does not include one.
12. Anti-Requirements
Things Paperclip explicitly does not do:
- Not an Agent runtime — Paperclip orchestrates, Agents run elsewhere
- Not a knowledge base — core has no wiki/docs/vector-DB (plugin territory)
- Not a SaaS — single-tenant, self-hosted
- Not opinionated about Agent implementation — any language, any framework, any runtime
- Not automatically self-healing — surfaces problems, doesn't silently fix them
- Does not manage delivery infrastructure — no repo management, no deployment, no file systems (but does manage task-linked documents and attachments)
- Does not auto-reassign work — stale tasks are surfaced, not silently redistributed
- Does not track external revenue/expenses — that's a future plugin. Token/LLM cost budgeting is core.
13. Principles (Consolidated)
- Unopinionated about how you run your Agents. Any language, any framework, any runtime. Paperclip is the control plane, not the execution plane.
- Company is the unit of organization. Everything lives under a Company.
- Tasks are the communication channel. All Agent communication flows through tasks + comments. No side channels.
- All work traces to the goal. Hierarchical task management — nothing exists in isolation.
- Board governs. Humans retain control through the Board. Conservative defaults (human approval required).
- Surface problems, don't hide them. Good auditing and visibility. No silent auto-recovery.
- Atomic ownership. Single assignee per task. Atomic checkout prevents conflicts.
- Progressive deployment. Trivial to start local, straightforward to scale to hosted.
- Extensible core. Clean boundaries so plugins can add capabilities (Adapters, knowledge base, revenue tracking) without modifying core.
Agent visual identity
Agent appearances are stable, versioned ClipLab end-cap personas, separate from behavioral instructions. Compact surfaces use on-demand cached PNG URLs; larger placements may use a lazy live character. See agent-personas.md for persistence, migration, rendering, and integration contracts.
Agent chat project handoff (2026-09-11)
Chat supports research and full plan drafting/revision in its existing plan document. On handoff, each ordinary assigned task receives the relevant plan in its own plan document, committed with task creation before execution is scheduled. The source plan remains in the conversation. Plan acceptance hands off execution; it never switches the conversation into implementation.
Chat instructions require selecting a suitable project, reusing an existing one where appropriate. The project requirement is prompt-only; ordinary projectless tasks remain supported. New parent relationships beneath conversation tasks are rejected by task services, including direct API creation and reparenting. Existing children remain readable/editable and can be moved elsewhere. The Subtasks panel is unchanged.
The create_project runtime tool uses the normal project API with durable idempotency. list_projects and list_project_repositories support selection. Multiple repositoryIds select authorized catalog entries; multiple HTTPS GitHub repositoryUrls register existing repositories absent from the catalog. IDs and URLs may be combined, but cannot accompany an explicit workspace. URLs do not create repositories on GitHub or grant credentials. Execution uses normal repository access rules. Repository IDs are revalidated against the authenticated run's responsible user and connection grants. Agents should consider proper available repositories, clarify material ambiguity, and use repository-free projects when appropriate for non-code work.
Confirmed project creation appears as a durable card in the shared task transcript, including selected repository links. Tasks are linked inline. Failed creation never produces a success card. Tool evals cover planning/handoff, project/repository selection, retries, permission and mode denials, and ordinary delegation regressions using the production chat directive.
Paused task messages
A paused task takes over the composer with an amber notice and a Resume action. Operators must release the effective task or ancestor pause before sending a new message. The draft stays intact. This applies to both task interfaces and to board comment API requests; an agent may still report interrupted work.
Experimental iMessage Photon channel
A Photon Cloud project can represent one agent through the existing experimental channel subsystem. DMs and explicitly enabled groups create or continue task-bound conversations. Linked sender identity is the default; telephone numbers, email addresses, names, and group membership do not grant Paperclip authority. Photos/files and ordinary questions/confirmations use the existing attachment, interaction, continuation, and publication contracts. Pause and Disconnect govern runtime behavior independently of the UI gate. Local Mac access, unsolicited conversations, and SMS/RCS fallback are excluded. Live qualification is required before release readiness. Pro shared allocation supports DMs only, with sender enrollment in Photon and separate identity linking in Paperclip. Shared channels reserve one project, not a pool phone number; group admission and publication are disabled. Dedicated allocation retains one selected number and individually enabled groups.
See iMessage Photon for the implementation contract, setup, recovery, boundaries, and qualification status.
Task search relevance
Task discovery uses PostgreSQL and the existing search indexes, with no external search service or background indexing job. The task-list quick search and full company search share lexical matching and ranking. Known identifiers and direct title matches lead; current conversation and document content supplies supporting evidence. See Task search relevance for the evaluation rubric, matching contract and reproducible quality tests.
Keyboard shortcuts
Keyboard shortcuts are always enabled for every signed-in user. There is no instance setting and no personal preference that turns them off.
Managed agents own a persistent file directory across tasks and sessions. The Instructions Editor and stopped agent execution synchronize the same current files, including AGENTS.md and its supporting files. Task working directories and provider home directories remain separate concepts. Concurrent runs synchronize only the files they change, with the last sync winning for the same file. Temporary copies are cleaned up; this storage does not add a revision-history system. See agent-files.md for lifecycle and upgrade compatibility.
Full agent storage produces a run warning without stopping current or future work. Storage limits constrain saved file changes, not the agent's ability to run and remove files to recover space.
Unsafe native workspace exports
An unsafe workspace link does not fail an accepted native task result. Retry
export automatically with confined entries only and keep archive confinement in
place. If the export remains unsafe, omit it and finish the saved result under
normal completion rules. Record diagnostics only in run logs; do not add a task
warning or manual repair action. This also applies to historical unsafe failures:
omit the already-rejected export, clear stale repair notices, and finalize the
accepted result without another provider turn, even when its old sandbox is
unavailable. Preserve current ownership and newer-work fences. See
native-workspace-finalization-recovery.md.
Git-backed skill library sources
Repositories can supply read-only company skills independently of project repositories. A source tracks repository identity, branch, selected paths, and installed commits; an existing GitHub connection supplies caller-authorized reads. Manual refresh publishes complete local immutable versions, preserving skill identity and assignments. Selection operates on whole skill packages, with inspectable included files, declared runtime requirements, and advisory warnings for missing or external references. New upstream skills require reviewed selection; removed or deselected skills remain installed. Editing starts with an independent copy. Write-back and PR publication are a later milestone; exact path and commit provenance provide their base.
Public assistant connection (opt-in)
The user-authorized MCP surface connects assistants to an explicitly selected company as the consenting person. It exposes first-party task reads, additive task creation and comments, durable documents and approval links. It reuses existing domain authorization and scheduling; OAuth does not grant agent identity, native run ownership, approval decisions or third-party credentials. See Public MCP for the implemented instance-side boundary, configuration, plugin packages and outstanding hosted release gates. The delivery plan separates external agent participation and granted third-party tools into later releases.
Experimental connection routing
A virtual AI connection can rotate new task/agent allocations through an authorized pool while preserving session affinity. Admission, credentials and durable recovery remain host responsibilities; policy can be supplied by an opt-in plugin. See the experimental contract.