DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00
…
…

Paperclip is the app people use to manage AI agents for work.

Quickstart · Docs · GitHub · Discord · Twitter · Website

MIT License Stars Star History Rank Discord


Sign up for the Paperclip Cloud waitlist →


Paperclip is the app people use to manage AI agents for work.

Open-source orchestration for teams of AI agents.

If OpenClaw is an employee, Paperclip is the company.

Paperclip is a Node.js server and React UI that orchestrates a team of AI agents to run a business. Bring your own agents, assign goals, and track work and costs from one dashboard. Choose models and harnesses per agent while keeping your team's tasks, skills, permissions, and history in one place.

It looks like a task manager. Under the hood: org charts, budgets, governance, goal alignment, and agent coordination.

Manage business goals, not pull requests.

Step Example
01 Define the goal "Build the #1 AI note-taking app to $1M MRR."
02 Hire the team CEO, CTO, engineers, designers, marketers — any bot, any provider.
03 Approve and run Review strategy. Set budgets. Hit go. Monitor from the dashboard.

Works
with
OpenClaw
OpenClaw
Claude Code
Claude Code
Codex
Codex
Cursor and Cursor Cloud
Cursor
+ Cloud
Gemini CLI
Gemini CLI
OpenCode
OpenCode
Pi
Pi
Hermes and Hermes Gateway
Hermes
+ Gateway
Grok Build
Grok Build
Kimi Code
Kimi Code

If it can receive a heartbeat, it's hired.

Custom processes, HTTP endpoints, and external adapter packages extend the roster. See the adapter overview for setup and capabilities.


Paperclip is right for you if

  • ✅ You want to build autonomous AI organizations
  • ✅ You coordinate many different agents (OpenClaw, Codex, Claude, Cursor) toward a common goal
  • ✅ You have 20 simultaneous Claude Code terminals open and lose track of what everyone is doing
  • ✅ You want agents running autonomously 24/7, but still want to audit work and chime in when needed
  • ✅ You want to monitor costs and enforce budgets
  • ✅ You want a process for managing agents that feels like using a task manager
  • ✅ You want to manage your autonomous businesses from your phone

The four pillars

Four things have to work for an organization of AI agents to actually produce: the tasks, the org, the training, and the infrastructure. Paperclip is built around exactly those four pillars.

The four pillars of Paperclip
Pillar Built for What it covers
Agentic Task Manager — Declare intent. Agents work. You verify the output. Everyone, daily Tasks, approvals & review gates · proactive agent coworkers · auditable routines & workflows · verify from diffs, screenshots & tests
Org Chart for Agents — Roles, permissions & boundaries for humans and agents. Managers Mixed human + agent org chart · responsibilities, delegation, specialization · governance: who can do what · scoped secrets & company boundaries · connection permissions & responsible-user identities
Agent Employee Training — Design, train & evaluate your AI employees. Enablers Skill Studio & shared org-wide skills · evals & saved test runs · active learning loops & quality metrics · performance reviews for agents · saved test inputs · skill version history & restore · reusable team templates
Agentic OS — The infrastructure that makes the work run. IT & platform Cross-provider runtime: any model, any agent · sandboxing, integrations & MCP servers · SSO, GRC, RBAC & cost controls · data privacy, internal trace collection, compounding data value · personal & shared app connections · run history & opt-in tracing

Features

🔌 Bring Your Own Agent

Any agent, any runtime, one org chart. If it can receive a heartbeat, it's hired.

🎯 Goal Alignment

Link tasks and projects to your organization goals. Agents receive the goal context behind their work.

💓 Heartbeats

Agents wake for assigned work, follow-up messages, or configured schedules. Delegation flows up and down the org chart.

💰 Cost Control

Company, agent, and project budgets. Track reported spend, get threshold alerts, and pause work at configured limits.

🏢 Multi-Organization

One deployment, many organizations. Separate tasks, agents, permissions, and activity histories for each.

🎫 Task Threads

Keep conversations, plans, blockers, files, and run history attached to the work. Assign tasks to agents or people.

🛡️ Governance

Configure review and approval stages, approve hires, and pause, reassign, or stop work when needed.

📊 Org Chart

Hierarchies, roles, reporting lines. Your agents have a boss, a title, and a job description.

📱 Mobile Ready

Monitor and manage your autonomous businesses from anywhere.

🔗 Apps & Connections

Connect services such as GitHub, Notion, and Railway, or your own MCP server. Set gateway actions to Allowed, Ask first, or Off.

👥 Shared Agents, Personal Accounts

Choose who can use a connection and which agents can access it. Managed GitHub operations can use the account of the person directing the work.

🧠 Skills & Skill Studio

Install or write shared skills, test them with saved inputs, inspect results, and restore earlier versions.

📅 Scheduled Routines

Run recurring work on a schedule or trigger it through an API or webhook. Each run has a task, an owner, and a history.

📎 Artifacts & Feedback

Find the files and documents agents produce. Preview supported formats and leave comments on specific passages in documents.

📦 Ready-Made Teams

Preview and install teams with roles, skills, projects, and routines. Choose their runtimes and make the setup your own.

Experimental Agent Chat and chat/email connectors add conversations with agents in Paperclip and through configured services such as Slack, Discord, Telegram, and AgentMail. Enable the relevant instance settings to try them.


Problems Paperclip solves

Without Paperclip With Paperclip
❌ You have 20 Claude Code tabs open and can't track which one does what. On reboot you lose everything. ✅ Tasks are ticket-based, conversations are threaded, sessions persist across reboots.
❌ You manually gather context from several places to remind your bot what you're actually doing. ✅ Context flows from the task up through the project and company goals — your agent always knows what to do and why.
❌ Folders of agent configs are disorganized and you're re-inventing task management, communication, and coordination between agents. ✅ Paperclip gives you org charts, ticketing, delegation, and governance out of the box — so you run a company, not a pile of scripts.
❌ Runaway loops waste hundreds of dollars of tokens and max your quota before you even know what happened. ✅ Spend tracking, budget alerts, and automatic pauses help you control the cost of ongoing work.
❌ You have recurring jobs (customer support, social, reports) and have to remember to manually kick them off. ✅ Routines create assigned tasks on a schedule, with outputs and run history you can inspect.
❌ You have an idea, you have to find your repo, fire up Claude Code, keep a tab open, and babysit it. ✅ Add a task in Paperclip. Your coding agent works on it until it's done. Management reviews their work.

Why Paperclip is special

Paperclip handles the hard orchestration details correctly.

Atomic task checkout. A single assignee and execution locks prevent competing runs from claiming the same task.
Persistent work context. Tasks, comments, and documents stay in Paperclip. Supporting adapters resume saved sessions across runs.
Runtime skill injection. Agents can learn Paperclip workflows and project context at runtime, without retraining.
Governance with rollback. Approval gates are enforced, config changes are revisioned, and bad changes can be rolled back safely.
Accountable connections. Human access, agent eligibility, and gateway action permissions are separate controls. Approve a call once or save a revocable rule.
Goal-aware execution. Linked tasks and projects carry goal ancestry so agents see the "why," not just a title.
Portable company templates. Export/import orgs, agents, and skills with secret scrubbing and collision handling.
Organization boundaries. Company-scoped access checks keep each organization's work, agents, and activity separate within one deployment.

What's Under the Hood

Paperclip is a full control plane, not a wrapper. Before you build any of this yourself, know that it already exists:

┌──────────────────────────────────────────────────────────────┐
│                       PAPERCLIP SERVER                       │
│                                                              │
│  ┌───────────┐  ┌───────────┐  ┌───────────┐  ┌───────────┐  │
│  │Identity & │  │  Work &   │  │ Heartbeat │  │Governance │  │
│  │  Access   │  │   Tasks   │  │ Execution │  │& Approvals│  │
│  └───────────┘  └───────────┘  └───────────┘  └───────────┘  │
│                                                              │
│  ┌───────────┐  ┌───────────┐  ┌───────────┐  ┌───────────┐  │
│  │ Org Chart │  │Workspaces │  │  Plugins  │  │  Budget   │  │
│  │ & Agents  │  │ & Runtime │  │           │  │ & Costs   │  │
│  └───────────┘  └───────────┘  └───────────┘  └───────────┘  │
│                                                              │
│  ┌───────────┐  ┌───────────┐  ┌───────────┐  ┌───────────┐  │
│  │ Routines  │  │ Secrets & │  │ Activity  │  │  Company  │  │
│  │& Schedules│  │  Storage  │  │ & Events  │  │Portability│  │
│  └───────────┘  └───────────┘  └───────────┘  └───────────┘  │
└──────────────────────────────────────────────────────────────┘
         ▲              ▲              ▲              ▲
   ┌─────┴─────┐  ┌─────┴─────┐  ┌─────┴─────┐  ┌─────┴─────┐
   │  Claude   │  │   Codex   │  │   CLI     │  │ HTTP/web  │
   │   Code    │  │           │  │  agents   │  │   bots    │
   └───────────┘  └───────────┘  └───────────┘  └───────────┘

The Systems

Identity & Access — Two deployment modes (trusted local or authenticated), human roles and permissions, agent API keys, short-lived run JWTs, company memberships, and invite flows. Responsible-user attribution follows work through delegation and supported managed connections.

Org Chart & Agents — Agents have roles, titles, reporting lines, permissions, and budgets. Adapter examples match the diagram: Claude Code, Codex, CLI agents such as Cursor/Gemini/bash, HTTP/webhook bots such as OpenClaw, and external adapter plugins. If it can receive a heartbeat, it's hired.

Work & Task System — Issues carry company/project/goal/parent links, atomic checkout with execution locks, first-class blocker dependencies, comments, documents, attachments, work products, labels, and inbox state. Search across work, review document revisions, and leave anchored feedback.

Heartbeat Execution — DB-backed wakeup queue with coalescing, budget checks, workspace resolution, secret injection, skill loading, and adapter invocation. Runs produce logs, usage records, and adapter-specific session state. Bounded recovery handles supported failures and surfaces cases that need human action.

Workspaces & Runtime — Project repositories and workspaces, optional isolated execution workspaces (git worktrees, operator branches), and runtime services (dev servers, preview URLs). Sandbox providers extend execution beyond the local host; availability depends on the configured environment and adapter.

Governance & Approvals — Board approval workflows, execution policies with review/approval stages, decision tracking, budget hard-stops, and agent pause/resume/terminate. Configured task reviews govern completion; connection action approvals govern calls through the tool gateway.

Budget & Cost Control — Token and cost tracking by company, agent, project, goal, issue, provider, and model. Scoped budget policies with warning thresholds and hard stops. Enforcement uses recorded spend; usage reporting and in-flight work can delay a stop.

Routines & Schedules — Recurring tasks with cron, webhook, and API triggers. Concurrency and catch-up policies. Each routine execution creates a tracked issue and wakes the assigned agent — no manual kick-offs needed.

Plugins — Instance-wide plugin system with out-of-process workers, capability-gated host services, job scheduling, tool exposure, and UI contributions. Extend Paperclip without forking it.

Secrets & Storage — Company secrets and per-person secret values, encrypted credential storage, local or S3-compatible file storage, attachments, and work products. Secret references supply credentials to authorized runs without copying values into ordinary agent configuration.

Activity & Events — Mutating actions, heartbeat state changes, cost events, approvals, comments, and work products are recorded as durable activity so operators can audit what happened and why.

Company Portability — Preview, export, and import organization packages with agents, skills, and optional projects, routines, tasks, and attachments. Referenced secret values are omitted; review packages before sharing because plain environment values and local paths can remain. Packages share an operating setup; full-instance recovery uses backups.


What Paperclip is not

Not just a chatbot. Conversations stay attached to tasks, plans, decisions, and outputs. Experimental Agent Chat can hand work off to assigned tasks.
Not an agent framework. We don't tell you how to build agents. We tell you how to run a company made of them.
Not just a workflow builder. Routines and experimental pipelines operate within an organization, with roles, goals, budgets, and governance.
Not a prompt manager. Agents bring their own prompts, models, and runtimes. Paperclip manages the organization they work in.
Not limited to one agent. Start with one agent and grow into a team with shared skills, delegation, and review.
Not only for code review. Coding and PR review fit alongside research, operations, content, and other work.

Quickstart

Open source. Self-hosted. No Paperclip account required. Follow the guided quickstart to set up your first agent.

Just ask your agent to install Paperclip

Share the installation guide with your agent.

Or install it yourself

With Node.js 24.11 or newer installed:

npx paperclipai@latest onboard --yes

The CLI runs from npm's cache; your instance configuration and data persist locally.

See the installation guide for managed installs, pinned versions, canary and git-ref installs, updates, rollback, service management, and uninstalling.

For an isolated manual test instance that is already initialized with a CEO agent, use test-drive. It stays in the foreground, never installs a service or creates a first task, and opens the browser only after setup succeeds:

ANTHROPIC_API_KEY=... npx paperclipai test-drive
OPENAI_API_KEY=... npx paperclipai test-drive --harness codex
OPENROUTER_API_KEY=... npx paperclipai test-drive \
  --harness opencode \
  --model openrouter/anthropic/claude-sonnet-4.5

Each run without --data-dir gets a unique, retained temporary directory; its absolute path is printed at startup. Pass --data-dir to reuse one, or --no-browser to leave the initialized instance unopened. When invoked from a linked Git worktree, test-drive also enables task execution in that worktree. See the test-drive guide for credential and reuse behavior.

Troubleshooting: private npm registry .npmrc

If this fails with an E404 for paperclipai (or similar) and you use a private npm registry (for example GitHub Packages) via a global ~/.npmrc, npx may be resolving paperclipai against that private registry instead of the public npm registry.

Diagnostic:

npm config get registry

Workaround (cross-platform; force the public npm registry for this command):

npx --registry https://registry.npmjs.org paperclipai@latest onboard --yes

That quickstart path now defaults to trusted local loopback mode for the fastest first run. To start in authenticated/private mode instead, choose a bind preset explicitly:

npx paperclipai@latest onboard --yes --bind lan
# or:
npx paperclipai@latest onboard --yes --bind tailnet

If you already have Paperclip configured, rerunning onboard keeps the existing config in place. Use npx paperclipai configure to edit settings.

Or manually:

git clone https://github.com/paperclipai/paperclip.git
cd paperclip
pnpm install
pnpm dev

This starts the UI and API at http://localhost:3100. An embedded PostgreSQL database is created automatically — no setup required.

Requirements: Node.js 24.11+, pnpm 9.15+

Source development also builds the native Paperclip Runner when enabled (the self-hosted default). Install a Rust toolchain, or set PAPERCLIP_RUNNER_BINARY to a compatible prebuilt runner.


FAQ

Q: Is this project maintained or just slop?

A: Paperclip is maintained by the Paperclip team. We've merged over 2,700 pull requests.


Q: What does a typical setup look like?

A: Locally, a single Node.js process manages an embedded Postgres and local file storage. For production, point it at your own Postgres and deploy however you like. Configure projects, agents, and goals — the agents take care of the rest.

For remote access, use authenticated mode with a private-network bind such as Tailscale, or deploy the persistent server with Docker. See deployment modes and the Docker guide.


Q: Can I run multiple companies?

A: Yes. A single deployment can host multiple organizations with company-scoped data and access checks.


Q: How is Paperclip different from agents like OpenClaw or Claude Code?

A: Paperclip uses those agents. It orchestrates them into a company — with org charts, budgets, goals, governance, and accountability.


Q: Why should I use Paperclip instead of just pointing my OpenClaw to Asana or Trello?

A: Agent orchestration has subtleties in how you coordinate who has work checked out, how to maintain sessions, monitoring costs, establishing governance - Paperclip does this for you.

(Bring-your-own-ticket-system is on the Roadmap)


Q: Do agents run continuously?

A: Agents wake for assigned work and follow-up messages. Optional timer heartbeats let them check for work periodically; routines create recurring tasks on their own schedules. You can also connect externally running agents such as OpenClaw. A mention alone does not assign work or wake another agent.


Development

pnpm dev              # Full dev (API + UI, watch mode)
pnpm dev:once         # Full dev without file watching
pnpm dev:server       # Server only
pnpm dev:mobile       # Serve prebuilt UI on :3101 for phones/tablets (proxies /api → :3100)
pnpm dev:both         # Run `pnpm dev` and `pnpm dev:mobile` together
pnpm build            # Build all
pnpm typecheck        # Type checking
pnpm test             # Cheap default test run (Vitest only)
pnpm test:watch       # Vitest watch mode
pnpm test:e2e         # Playwright browser suite
pnpm db:generate      # Generate DB migration
pnpm db:migrate       # Apply migrations

pnpm test does not run Playwright. Browser suites stay separate and are typically run only when working on those flows or in CI.

See doc/DEVELOPING.md for the full development guide.


Roadmap

  • ✅ Plugin system (e.g. add a knowledge base, custom tracing, queues, etc)
  • ✅ Get OpenClaw / claw-style agent employees
  • ✅ companies.sh - import and export entire organizations
  • ✅ Easy AGENTS.md configurations
  • ✅ Skills Manager, Skill Studio & Skills Store
  • ✅ Scheduled Routines
  • ✅ Better Budgeting
  • ✅ Agent Reviews and Approvals
  • ✅ Multiple Human Users
  • ✅ Cloud / Sandbox agents (e2b, Cloudflare, Daytona, Modal, Novita, self-hosted Kubernetes)
  • ✅ Artifacts & Work Products
  • ✅ Deep Planning (planning mode, revisioned plans, plan approvals)
  • ✅ Enforced Outcomes (watchdogs, recovery actions, review gates)
  • ✅ MCP Tool Gateway & Apps (governed tool access)
  • ✅ Secrets Manager with per-agent access
  • ✅ Activity log & action attribution
  • ✅ Self-healing runs & automatic recovery
  • ✅ Agent evals & feedback
  • ✅ Connected Apps
  • ✅ Personal & Shared AI Accounts
  • ✅ Shared Agents Use Personal GitHub Identities
  • ✅ Skill Version History & Restore
  • ✅ Document Comments & Revision History
  • ✅ Company-Wide Search
  • ✅ Multi-Model & Multi-Harness Teams
  • 🟡 Memory / Knowledge
  • ⚪ MAXIMIZER MODE
  • ⚪ Work Queues
  • ⚪ Self-Organization
  • ⚪ Automatic Organizational Learning
  • 🟡 Agent Chat
  • 🟡 Cloud deployments
  • ⚪ Desktop App
  • ⚪ Bring-your-own-ticket-system (Asana / Linear / Jira as on-ramps)

This is the short roadmap preview. See the full roadmap in ROADMAP.md.


Community & Plugins

Find Plugins and more at awesome-paperclip

Observability

Paperclip ships with opt-in OpenTelemetry auto-instrumentation for the server (traces only). It activates when OTEL_EXPORTER_OTLP_ENDPOINT is set and supports grpc, http/protobuf, and http/json via the standard OTEL_EXPORTER_OTLP_PROTOCOL env var. @opentelemetry/api is a normal server dependency; the SDK, auto-instrumentation, and exporter packages are optional peer dependencies — install them only if you want tracing. See doc/observability.md for install commands and the full env-var reference.

Paperclip also ships with opt-in Sentry error monitoring for the server and the browser. Set SENTRY_DSN_FRONTEND to activate it for the browser and SENTRY_DSN_BACKEND to activate it for the server — each variable is optional, and the legacy SENTRY_DSN variable still works as a fallback for either component. The supported server SDK version is @sentry/node@10.71.0; it is an optional peer dependency for the server, so install it only if you want error monitoring. The browser SDK, @sentry/browser, is pinned to the same exact version. See doc/observability.md for the install command, the privacy settings, and the full default capture set.

Telemetry

Paperclip collects anonymous usage telemetry to help us understand how the product is used and improve it. No personal information, issue content, prompts, file paths, or secrets are ever collected. Private repository references are hashed with a per-install salt before being sent.

Contributors changing emitted telemetry events should follow the Telemetry Data Contract. For proposed first-party events that are not in the generated contract yet, follow Telemetry Workflow.

Telemetry is enabled by default and can be disabled with any of the following:

Method How
Environment variable PAPERCLIP_TELEMETRY_DISABLED=1
Standard convention DO_NOT_TRACK=1
CI environments Automatically disabled when CI=true
Config file Set telemetry.enabled: false in your Paperclip config

Contributing

We welcome contributions. See the contributing guide for details.

We're hiring


Community


License

MIT © 2026 Paperclip Labs, Inc

Star History

Star History Chart

Open source under MIT. Built for people who want to get work done, not babysit agents.

S
Description
No description provided
Readme MIT
1.7 GiB
0 Stars 1 Watchers 0 Forks
Languages
TypeScript 93.3%
Rust 3%
JavaScript 2.7%
Shell 0.5%
CSS 0.3%