Files
PaperClipAI/doc/PRODUCT.md
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00

17 KiB
Raw Permalink Blame History

Paperclip — Product Definition

What It Is

Paperclip is the control plane for autonomous AI companies. One instance of Paperclip can run multiple companies. A company is a first-order object.

Core Concepts

Company

A company has:

  • A goal — the reason it exists ("Create the #1 AI note-taking app that does $1M MRR within 3 months")
  • Employees — every employee is an AI agent
  • Org structure — who reports to whom
  • Revenue & expenses — tracked at the company level
  • Task hierarchy — all work traces back to the company goal

Employees & Agents

Every employee is an agent. When you create a company, you start by defining the CEO, then build out from there.

Each employee has:

  • Adapter type + config — how this agent runs and what defines its identity/behavior. This is adapter-specific (e.g., an OpenClaw agent might use SOUL.md and HEARTBEAT.md files; a Claude Code agent might use CLAUDE.md; a bare script might use CLI args). Paperclip doesn't prescribe the format — the adapter does.
  • Role & reporting — their title, who they report to, who reports to them
  • Capabilities description — a short paragraph on what this agent does and when they're relevant (helps other agents discover who can help with what)

Example: A CEO agent's adapter config tells it to "review what your executives are doing, check company metrics, reprioritize if needed, assign new strategic initiatives" on each heartbeat. An engineer's config tells it to "check assigned tasks, pick the highest priority, and work it."

Then you define who reports to the CEO: a CTO managing programmers, a CMO managing the marketing team, and so on. Every agent in the tree gets their own adapter configuration.

Agent Execution

Paperclip supports several ways to run an agent's heartbeat:

  1. Local CLI/session adapters — Paperclip starts or resumes local coding-tool sessions such as Claude Code, Codex, Gemini, OpenCode, Pi, and Cursor, then tracks the run.
  2. Run a command — Paperclip kicks off a process (shell command, Python script, etc.) and tracks it. The heartbeat is "execute this and monitor it."
  3. Fire and forget a request — Paperclip sends a webhook/API call to an externally running agent. The heartbeat is "notify this agent to wake up." OpenClaw-style hooks work this way.
  4. External adapter plugins — Paperclip loads adapter packages through the plugin/adapter flow so self-hosted installs can add runtimes without hardcoding them in core.

Agent runs can use project and execution workspaces, managed runtime services such as preview/dev servers, adapter-specific session state, and HTTP/webhook-style execution. We provide sensible defaults, but the adapter is still the boundary: if a runtime can be invoked, observed, and authorized, Paperclip can coordinate it.

Task Management

Task management is hierarchical. At any moment, every piece of work must trace back to the company's top-level goal through a chain of parent tasks:

I am researching the Facebook ads Granola uses (current task)
  because → I need to create Facebook ads for our software (parent)
    because → I need to grow new signups by 100 users (parent)
      because → I need to get revenue to $2,000 this week (parent)
        because → ...
          because → We're building the #1 AI note-taking app to $1M MRR in 3 months

Tasks have parentage. Every task exists in service of a parent task, all the way up to the company goal. This is what keeps autonomous agents aligned — they can always answer "why am I doing this?"

The current issue model includes stable issue identifiers, parent/sub-issues, blockers, a single assignee, comments, issue documents, attachments and work products, and review/approval handoffs. That structure keeps work inspectable by both the board and agents while still allowing agents to decompose work into smaller tasks.

Company Skills and Policy

Company skills are shared operating capabilities, not privileged objects by default. Every authenticated agent in a company can create, import, install, edit, update, test, reset, and remove that company's skills unless the company has configured an explicit restriction.

The governing rule is: skill permissions are opt-in restrictions, not opt-in capabilities. Missing skill grants never create a denial in an otherwise unconfigured company, and ordinary skill work does not require board confirmation, a draft-only workflow, or an activation approval.

Core Paperclip owns the skill runtime, company-boundary enforcement, policy evaluation contract, API denials, validation, path containment, secret redaction, and activity logging. Those safety invariants cannot be disabled by policy. Open-by-default skill work never authorizes arbitrary host-path reads, unsafe executable content, or policy edits: local imports and scans must stay within Paperclip-known workspace or managed-skill roots, remote sources must resolve to validated immutable content, and platform safety denials must stay distinct from optional administrative restrictions. Paperclip EE may provide detailed administration for per-agent, per-role, per-action, per-source, and protected-skill rules, but EE is not required to use skills and is not an enforcement boundary. Without EE, companies remain open by default and any already-configured restrictions continue to be enforced by core.

GitHub repositories are first-class sources within Skills. People import selected skills through their existing GitHub connection and manually refresh installed snapshots. The source records repository, branch, path, and installed commit; the connection supplies caller-authorized access. Originals remain viewable and testable, with independent editable copies. New upstream skills require selection, and deselection, source disconnection, or upstream removal retains installed content.

An explicit restricted policy may deny selected operations or switch to a default-deny preset with explicit allow rules. Core exposes a stable versioned policy API so EE and other administrative clients configure and simulate the same evaluator used by skill mutation routes. Core Skill Studio only needs to perform normal skill work, explain an explicit denial, and point administrators to EE when its richer policy UI is available; it must not recreate a partial enterprise permission editor.

Principles

  1. Unopinionated about how you run your agents. Your agents could be OpenClaw bots, Python scripts, Node scripts, Claude Code sessions, Codex instances — we don't care. Paperclip defines the control plane for communication and provides utility infrastructure for heartbeats. It does not mandate an agent runtime.

  2. Company is the unit of organization. Everything lives under a company. One Paperclip instance, many companies.

  3. Adapter config defines the agent. Every agent has an adapter type and configuration that controls its identity and behavior. The minimum contract is just "be callable."

  4. All work traces to the goal. Hierarchical task management means nothing exists in isolation. If you can't explain why a task matters to the company goal, it shouldn't exist.

  5. Control plane, not execution plane. Paperclip orchestrates. Agents run wherever they run and phone home.

User Flow (Dream Scenario)

  1. Open Paperclip, create a new company
  2. Define the company's goal: "Create the #1 AI note-taking app, $1M MRR in 3 months"
  3. Create the CEO
    • Choose an adapter (e.g., process adapter for Claude Code, HTTP adapter for OpenClaw)
    • Configure the adapter (agent identity, loop behavior, execution settings)
    • CEO proposes strategic breakdown → board approves
  4. Define the CEO's reports: CTO, CMO, CFO, etc.
    • Each gets their own adapter config and role definition
  5. Define their reports: engineers under CTO, marketers under CMO, etc.
  6. Set budgets, define initial strategic tasks
  7. Hit go — agents start their heartbeats and the company runs

Guidelines

There are two runtime modes Paperclip must support:

  • local_trusted (default): single-user local trusted deployment with no login friction
  • authenticated: login-required mode that supports both private-network and public deployment exposure policies

Canonical mode design and command expectations live in doc/DEPLOYMENT-MODES.md.

Further Detail

See SPEC.md for the full technical specification and TASKS.md for the task management data model.


Paperclip’s core identity is a control plane for autonomous AI companies, centered on companies, org charts, goals, issues/comments, heartbeats, budgets, approvals, and board governance. The public docs are also explicit about the current boundaries: tasks/comments are the built-in communication model, Paperclip is not a chatbot, and it is not a code review tool. The roadmap already points toward easier onboarding, cloud agents, easier agent configuration, plugins, better docs, and ClipMart/ClipHub-style reusable companies/templates.

What Paperclip should do vs. not do

Do

  • Stay board-level and company-level. Users should manage goals, orgs, budgets, approvals, and outputs.
  • Make the first five minutes feel magical: install, answer a few questions, see a CEO do something real.
  • Keep work anchored to issues/comments/projects/goals, even if the surface feels conversational.
  • Treat agency / internal team / startup as the same underlying abstraction with different templates and labels.
  • Make outputs first-class: files, docs, reports, previews, links, screenshots.
  • Provide hooks into engineering workflows: worktrees, preview servers, PR links, external review tools.
  • Use plugins for edge cases like rich chat, knowledge bases, doc editors, custom tracing.

Do not

  • Do not make the core product a general chat app. The current product definition is explicitly task/comment-centric and “not a chatbot,” and that boundary is valuable.
  • Do not build a complete Jira/GitHub replacement. The repo/docs already position Paperclip as organization orchestration, not focused on pull-request review.
  • Do not build enterprise-grade RBAC first. Paperclip now has authenticated mode, company memberships, instance roles, and permission grants, but fine-grained enterprise governance should remain secondary to the core company control plane.
  • Do not interpret agent-level privacy flags as a project/issue privacy feature in V1; work visibility stays company-scoped.
  • Do not lead with raw bash logs and transcripts. Default view should be human-readable intent/progress, with raw detail beneath.
  • Do not force users to understand provider/API-key plumbing unless absolutely necessary. There are active onboarding/auth issues already; friction here is clearly real.

Specific design goals

  1. Time-to-first-success under 5 minutes A fresh user should go from install to “my CEO completed a first task” in one sitting.

  2. Board-level abstraction always wins The default UI should answer: what is the company doing, who is doing it, why does it matter, what did it cost, and what needs my approval.

  3. Conversation stays attached to work objects “Chat with CEO” should still resolve to strategy threads, decisions, tasks, or approvals.

  4. Progressive disclosure Top layer: human-readable summary. Middle layer: checklist/steps/artifacts. Bottom layer: raw logs/tool calls/transcript.

  5. Output-first Work is not done until the user can see the result: file, document, preview link, screenshot, plan, or PR.

  6. Execution visibility without log worship Active runs, recovery issues, blockers, and work products should be first-class surfaces. Raw transcripts are available when needed, but they are not the primary product surface.

  7. Local-first, cloud-ready The mental model should not change between local solo use and shared/private or public/cloud deployment.

  8. Safe autonomy Auto mode is allowed; hidden token burn is not.

  9. Thin core, rich edges Put optional chat, knowledge, and special surfaces into plugins/extensions rather than bloating the control plane.

Experimental iMessage Photon channel

A Photon Cloud project can represent one agent through the existing experimental channel subsystem. DMs and explicitly enabled groups create or continue task-bound conversations. Linked sender identity is the default; telephone numbers, email addresses, names, and group membership do not grant Paperclip authority. Photos/files and ordinary questions/confirmations use the existing attachment, interaction, continuation, and publication contracts. Pause and Disconnect govern runtime behavior independently of the UI gate. Local Mac access, unsolicited conversations, and SMS/RCS fallback are excluded. Live qualification is required before release readiness. Pro shared allocation supports DMs only, with sender enrollment in Photon and separate identity linking in Paperclip. Shared channels reserve one project, not a pool phone number; group admission and publication are disabled. Dedicated allocation retains one selected number and individually enabled groups.

See iMessage Photon for the implementation contract, setup, recovery, boundaries, and qualification status.

Experimental persistent agent conversations

Agent Chat is an opt-in core task presentation (enableAgentChat, off by default). Each person has one persistent task-backed conversation per agent and company, with ordinary company task visibility. The shared task composer, transcript, tools, files, and document panel remain the interaction surface. Agents clarify goals and hand substantial execution to linked, assigned tasks; a reply ends a turn without completing the conversation. /new starts fresh provider context in the same conversation while preserving visible history and artifacts. Healthy idle conversations wait for a message and do not count as unfinished execution work. See doc/plans/2026-09-10-agent-chat.md for the implementation contract.

Agent chat project handoff (2026-09-11)

Chat supports research and full plan drafting/revision in its existing plan document. On handoff, each ordinary assigned task receives the relevant plan in its own plan document, committed with task creation before execution is scheduled. The source plan remains in the conversation. Plan acceptance hands off execution; it never switches the conversation into implementation.

Chat instructions require selecting a suitable project, reusing an existing one where appropriate. The project requirement is prompt-only; ordinary projectless tasks remain supported. New parent relationships beneath conversation tasks are rejected by task services, including direct API creation and reparenting. Existing children remain readable/editable and can be moved elsewhere. The Subtasks panel is unchanged.

The create_project runtime tool uses the normal project API with durable idempotency. list_projects and list_project_repositories support selection. Multiple repositoryIds select authorized catalog entries; multiple HTTPS GitHub repositoryUrls register existing repositories absent from the catalog. IDs and URLs may be combined, but cannot accompany an explicit workspace. URLs do not create repositories on GitHub or grant credentials. Execution uses normal repository access rules. Repository IDs are revalidated against the authenticated run's responsible user and connection grants. Agents should consider proper available repositories, clarify material ambiguity, and use repository-free projects when appropriate for non-code work.

Confirmed project creation appears as a durable card in the shared task transcript, including selected repository links. Tasks are linked inline. Failed creation never produces a success card. Tool evals cover planning/handoff, project/repository selection, retries, permission and mode denials, and ordinary delegation regressions using the production chat directive.

In-app announcements

An optional announcement card shares product news with board users on opening or returning to Paperclip. Dismissals persist per user across companies and browsers within an instance. Operators can disable fetching independently of telemetry. See Announcements.

Agent chat discovery

With Agent Chat enabled, the Chats sidebar always includes the company's earliest-created agent, plus personal starred agents and up to four other recent conversations. First use has the same compact rows as returning use. The compose icon shares a column with stars and appears on hover or keyboard focus (always on touch). It opens a company-wide name/role search, independent of sidebar membership. Selecting an agent opens their persistent conversation; it does not reset history or create a task until the existing first-write flow requires one.