Files
1c07b5903b feat: Chat leads the left nav, agent work beside chats, and a Combined Inbox + Task List flag (#15100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The left nav is the main way people move between tasks, the inbox,
and Agent Chat
> - The nav has separate Inbox and Tasks rows that show overlapping
work, and Chat is one row among many
> - The side panel beside a chat shows the conversation's own artifacts,
not the work the agent did
> - People want Chat to be easy to find, and they want one place for
their task views
> - This pull request moves Chat to the top of Work, shows the agent's
tasks and artifacts beside each chat, and adds an experimental flag that
folds Inbox into Tasks
> - The benefit is a shorter nav and a chat view that shows what the
agent is working on. Both changes stay off until an operator enables
them

## Linked Issues or Issue Description

Refs #14706 (the secondary Agent Chat navigation this change builds on)
Refs #14848 (reopen the last visited agent chat)

**Subsystem affected**
UI navigation (left nav, mobile tab bar), Agent Chat side panel, task
list and inbox, and the company artifacts API.

**Problem or motivation**
Inbox and Tasks are two nav rows for overlapping work. Chat sits in the
top group with no clear home. The chat rail lists only agents you
already talked to, so you cannot see your other teammates there. The
side panel beside a chat shows only the conversation's own artifacts. It
does not show the tasks and files the agent made.

**Proposed solution**
With Agent Chat on, Chat leads the Work section and the rail lists every
eligible agent. The chat side panel opens on the agent's tasks as cards,
and the agent's artifacts are available from +. A new experimental flag,
Combined Inbox + Task List, makes Inbox a set of views inside Tasks.

**Alternatives considered**
Rebuilding the inbox inside the task list. Instead, `/issues` hosts the
existing Inbox component for inbox views and the existing task list for
status views, so all inbox behaviour stays the same.

**Roadmap alignment**
Agent Chat (ROADMAP.md, "Agent Chat (including CEO Chat)"). All changes
are behind experimental flags that are off by default.

## What Changed

- **Agent Chat nav (streamlined shell):** Chat is the first row of Work,
not a top-group row. Workspaces leaves the nav while Agent Chat is on.
The mobile tab bar is Home · Chat · + · Tasks · Agents. The legacy shell
keeps master's top-group Chat row.
- **Chat rail:** `AgentConversationsSidebar` lists every eligible agent.
The open chat is first, then conversations by recent activity, then the
rest of the roster alphabetically. Terminated agents and agents you left
are omitted unless you have history with them. The picker still marks
only real conversations as "Open chat".
- **Chat side panel:** a new default Tasks tab shows one card per task
the agent created, was assigned, commented on, or acted on, newest
first. It has the task list's filter popover and a sort control. **+ →
Artifacts** shows the agent's artifacts as cards. Cards open in a new
tab. Agent Chat off keeps the old Artifacts tab.
- **Artifacts API:** `GET /api/companies/:companyId/artifacts` accepts
`agentId`. The filter applies to documents, work products, and
attachments by the agent each result is attributed to. The shared
validator and the UI client carry the new parameter, and the OpenAPI
entry picks it up from the shared schema.
- **Combined Inbox + Task List flag (`enableCombinedInboxTasks`, off by
default):** new card in Settings > Experimental. The Inbox row goes away
and its badge moves to Tasks. A Views menu on `/issues` covers Mine,
Unread, Blocked, Recent, Everything, All, Active, Backlog, and Done.
Bare `/issues` opens the last-used view (default Mine). Links that carry
`assignee`, `workspace`, `participantAgentId`, or `q` open All so the
filter is kept. `/inbox/*` and
`/issues/{all,active,backlog,done,recent}` redirect to the matching
view. `/inbox/requests` stays its own page.
- **Task detail breadcrumb:** the view key now decides the source, so
quick-archive still works after a reload from an inbox view.
- **Docs:** `doc/PRODUCT.md` and `doc/SPEC.md` describe the chat rail,
the chat side panel, and the new flag.

## Verification

- `cd ui && npx vitest run --no-file-parallelism src/components/chat
src/components/task-side-panel/TaskSidePanel.test.tsx
src/components/AgentConversationsSidebar.test.tsx
src/components/Sidebar.test.tsx
src/components/SidebarCompanyMenu.test.tsx
src/components/Layout.test.tsx src/pages/AgentChats.test.tsx
src/pages/InstanceExperimentalSettings.test.tsx
src/lib/task-views.test.ts src/lib/issueDetailBreadcrumb.test.ts
src/pages/Inbox.test.tsx src/pages/Issues.test.tsx src/App.test.tsx
src/App.activity-routing.test.tsx
src/components/MobileBottomNav.test.tsx
src/components/CommandPalette.test.tsx`: 20 files, 356 tests pass.
- `cd server && npx vitest run
src/__tests__/company-artifacts-service.test.ts`: 13/13 pass, including
the new agent-filter test across all three artifact sources.
- The new rail test fails against the unmodified rail.
- `pnpm check:token-gates`: all gates clean.
- Manual: enable Agent Chat in Settings > Experimental. Open Chat. The
rail lists all agents. Open a chat. The side panel shows the agent's
tasks. Use **+ → Artifacts** to see the agent's artifacts. Then enable
Combined Inbox + Task List. The Inbox row goes away, and Tasks shows a
Views menu.
- Snapshot baselines are intentionally not updated. See
`doc/design/DECISION-SHEET.md`, "Per-change snapshot verification
demoted to dormant (Jul 13 2026)".

## Risks

- With both flags off, the app behaves like master. The only exception
is the API: it accepts a new optional query parameter.
- With Agent Chat on, the rail can list many agents in a large company.
It uses the agent list the app already loads, and search filters it.
- The Tasks panel reads at most 200 recently updated tasks per agent and
says so when it reaches the limit. The Artifacts panel reads at most 500
of the agent's artifacts.
- Combined Inbox + Task List changes what bare `/issues` opens for
people who enable it. Deep links with a task filter still open All.

## Model Used

- Claude (Anthropic), model ID `claude-opus-5-5`, through Claude Code
with tool use (shell, file edit, test runs). Extended thinking was
enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 02:13:46 -07:00

235 lines
18 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Paperclip — Product Definition
## What It Is
Paperclip is the control plane for autonomous AI companies. One instance of Paperclip can run multiple companies. A **company** is a first-order object.
## Core Concepts
### Company
A company has:
- A **goal** — the reason it exists ("Create the #1 AI note-taking app that does $1M MRR within 3 months")
- **Employees** — every employee is an AI agent
- **Org structure** — who reports to whom
- **Revenue & expenses** — tracked at the company level
- **Task hierarchy** — all work traces back to the company goal
### Employees & Agents
Every employee is an agent. When you create a company, you start by defining the CEO, then build out from there.
Each employee has:
- **Adapter type + config** — how this agent runs and what defines its identity/behavior. This is adapter-specific (e.g., an OpenClaw agent might use SOUL.md and HEARTBEAT.md files; a Claude Code agent might use CLAUDE.md; a bare script might use CLI args). Paperclip doesn't prescribe the format — the adapter does.
- **Role & reporting** — their title, who they report to, who reports to them
- **Capabilities description** — a short paragraph on what this agent does and when they're relevant (helps other agents discover who can help with what)
Example: A CEO agent's adapter config tells it to "review what your executives are doing, check company metrics, reprioritize if needed, assign new strategic initiatives" on each heartbeat. An engineer's config tells it to "check assigned tasks, pick the highest priority, and work it."
Then you define who reports to the CEO: a CTO managing programmers, a CMO managing the marketing team, and so on. Every agent in the tree gets their own adapter configuration.
### Agent Execution
Paperclip supports several ways to run an agent's heartbeat:
1. **Local CLI/session adapters** — Paperclip starts or resumes local coding-tool sessions such as Claude Code, Codex, Gemini, OpenCode, Pi, and Cursor, then tracks the run.
2. **Run a command** — Paperclip kicks off a process (shell command, Python script, etc.) and tracks it. The heartbeat is "execute this and monitor it."
3. **Fire and forget a request** — Paperclip sends a webhook/API call to an externally running agent. The heartbeat is "notify this agent to wake up." OpenClaw-style hooks work this way.
4. **External adapter plugins** — Paperclip loads adapter packages through the plugin/adapter flow so self-hosted installs can add runtimes without hardcoding them in core.
Agent runs can use project and execution workspaces, managed runtime services such as preview/dev servers, adapter-specific session state, and HTTP/webhook-style execution. We provide sensible defaults, but the adapter is still the boundary: if a runtime can be invoked, observed, and authorized, Paperclip can coordinate it.
### Task Management
Task management is hierarchical. At any moment, every piece of work must trace back to the company's top-level goal through a chain of parent tasks:
```
I am researching the Facebook ads Granola uses (current task)
because → I need to create Facebook ads for our software (parent)
because → I need to grow new signups by 100 users (parent)
because → I need to get revenue to $2,000 this week (parent)
because → ...
because → We're building the #1 AI note-taking app to $1M MRR in 3 months
```
Tasks have parentage. Every task exists in service of a parent task, all the way up to the company goal. This is what keeps autonomous agents aligned — they can always answer "why am I doing this?"
The current issue model includes stable issue identifiers, parent/sub-issues, blockers, a single assignee, comments, issue documents, attachments and work products, and review/approval handoffs. That structure keeps work inspectable by both the board and agents while still allowing agents to decompose work into smaller tasks.
### Company Skills and Policy
Company skills are shared operating capabilities, not privileged objects by default. Every authenticated agent in a company can create, import, install, edit, update, test, reset, and remove that company's skills unless the company has configured an explicit restriction.
The governing rule is: **skill permissions are opt-in restrictions, not opt-in capabilities**. Missing skill grants never create a denial in an otherwise unconfigured company, and ordinary skill work does not require board confirmation, a draft-only workflow, or an activation approval.
Core Paperclip owns the skill runtime, company-boundary enforcement, policy evaluation contract, API denials, validation, path containment, secret redaction, and activity logging. Those safety invariants cannot be disabled by policy. Open-by-default skill work never authorizes arbitrary host-path reads, unsafe executable content, or policy edits: local imports and scans must stay within Paperclip-known workspace or managed-skill roots, remote sources must resolve to validated immutable content, and platform safety denials must stay distinct from optional administrative restrictions. Paperclip EE may provide detailed administration for per-agent, per-role, per-action, per-source, and protected-skill rules, but EE is not required to use skills and is not an enforcement boundary. Without EE, companies remain open by default and any already-configured restrictions continue to be enforced by core.
GitHub repositories are first-class sources within Skills. People import selected
skills through their existing GitHub connection and manually refresh installed
snapshots. The source records repository, branch, path, and installed commit;
the connection supplies caller-authorized access. Originals remain viewable and
testable, with independent editable copies. New upstream skills require selection,
and deselection, source disconnection, or upstream removal retains installed content.
An explicit restricted policy may deny selected operations or switch to a default-deny preset with explicit allow rules. Core exposes a stable versioned policy API so EE and other administrative clients configure and simulate the same evaluator used by skill mutation routes. Core Skill Studio only needs to perform normal skill work, explain an explicit denial, and point administrators to EE when its richer policy UI is available; it must not recreate a partial enterprise permission editor.
## Principles
1. **Unopinionated about how you run your agents.** Your agents could be OpenClaw bots, Python scripts, Node scripts, Claude Code sessions, Codex instances — we don't care. Paperclip defines the control plane for communication and provides utility infrastructure for heartbeats. It does not mandate an agent runtime.
2. **Company is the unit of organization.** Everything lives under a company. One Paperclip instance, many companies.
3. **Adapter config defines the agent.** Every agent has an adapter type and configuration that controls its identity and behavior. The minimum contract is just "be callable."
4. **All work traces to the goal.** Hierarchical task management means nothing exists in isolation. If you can't explain why a task matters to the company goal, it shouldn't exist.
5. **Control plane, not execution plane.** Paperclip orchestrates. Agents run wherever they run and phone home.
## User Flow (Dream Scenario)
1. Open Paperclip, create a new company
2. Define the company's goal: "Create the #1 AI note-taking app, $1M MRR in 3 months"
3. Create the CEO
- Choose an adapter (e.g., process adapter for Claude Code, HTTP adapter for OpenClaw)
- Configure the adapter (agent identity, loop behavior, execution settings)
- CEO proposes strategic breakdown → board approves
4. Define the CEO's reports: CTO, CMO, CFO, etc.
- Each gets their own adapter config and role definition
5. Define their reports: engineers under CTO, marketers under CMO, etc.
6. Set budgets, define initial strategic tasks
7. Hit go — agents start their heartbeats and the company runs
## Guidelines
There are two runtime modes Paperclip must support:
- `local_trusted` (default): single-user local trusted deployment with no login friction
- `authenticated`: login-required mode that supports both private-network and public deployment exposure policies
Canonical mode design and command expectations live in `doc/DEPLOYMENT-MODES.md`.
## Further Detail
See [SPEC.md](./SPEC.md) for the full technical specification and [TASKS.md](./TASKS.md) for the task management data model.
---
Paperclip’s core identity is a **control plane for autonomous AI companies**, centered on **companies, org charts, goals, issues/comments, heartbeats, budgets, approvals, and board governance**. The public docs are also explicit about the current boundaries: **tasks/comments are the built-in communication model**, Paperclip is **not a chatbot**, and it is **not a code review tool**. The roadmap already points toward **easier onboarding, cloud agents, easier agent configuration, plugins, better docs, and ClipMart/ClipHub-style reusable companies/templates**.
## What Paperclip should do vs. not do
**Do**
- Stay **board-level and company-level**. Users should manage goals, orgs, budgets, approvals, and outputs.
- Make the first five minutes feel magical: install, answer a few questions, see a CEO do something real.
- Keep work anchored to **issues/comments/projects/goals**, even if the surface feels conversational.
- Treat **agency / internal team / startup** as the same underlying abstraction with different templates and labels.
- Make outputs first-class: files, docs, reports, previews, links, screenshots.
- Provide **hooks into engineering workflows**: worktrees, preview servers, PR links, external review tools.
- Use **plugins** for edge cases like rich chat, knowledge bases, doc editors, custom tracing.
**Do not**
- Do not make the core product a general chat app. The current product definition is explicitly task/comment-centric and “not a chatbot,” and that boundary is valuable.
- Do not build a complete Jira/GitHub replacement. The repo/docs already position Paperclip as organization orchestration, not focused on pull-request review.
- Do not build enterprise-grade RBAC first. Paperclip now has authenticated mode, company memberships, instance roles, and permission grants, but fine-grained enterprise governance should remain secondary to the core company control plane.
- Do not interpret agent-level privacy flags as a project/issue privacy feature in V1; work visibility stays company-scoped.
- Do not lead with raw bash logs and transcripts. Default view should be human-readable intent/progress, with raw detail beneath.
- Do not force users to understand provider/API-key plumbing unless absolutely necessary. There are active onboarding/auth issues already; friction here is clearly real.
## Specific design goals
1. **Time-to-first-success under 5 minutes**
A fresh user should go from install to “my CEO completed a first task” in one sitting.
2. **Board-level abstraction always wins**
The default UI should answer: what is the company doing, who is doing it, why does it matter, what did it cost, and what needs my approval.
3. **Conversation stays attached to work objects**
“Chat with CEO” should still resolve to strategy threads, decisions, tasks, or approvals.
4. **Progressive disclosure**
Top layer: human-readable summary. Middle layer: checklist/steps/artifacts. Bottom layer: raw logs/tool calls/transcript.
5. **Output-first**
Work is not done until the user can see the result: file, document, preview link, screenshot, plan, or PR.
6. **Execution visibility without log worship**
Active runs, recovery issues, blockers, and work products should be first-class surfaces. Raw transcripts are available when needed, but they are not the primary product surface.
7. **Local-first, cloud-ready**
The mental model should not change between local solo use and shared/private or public/cloud deployment.
8. **Safe autonomy**
Auto mode is allowed; hidden token burn is not.
9. **Thin core, rich edges**
Put optional chat, knowledge, and special surfaces into plugins/extensions rather than bloating the control plane.
### Experimental iMessage Photon channel
A Photon Cloud project can represent one agent through the existing
experimental channel subsystem. DMs and explicitly enabled groups create or
continue task-bound conversations. Linked sender identity is the default;
telephone numbers, email addresses, names, and group membership do not grant
Paperclip authority. Photos/files and ordinary questions/confirmations use the
existing attachment, interaction, continuation, and publication contracts.
Pause and Disconnect govern runtime behavior independently of the UI gate.
Local Mac access, unsolicited conversations, and SMS/RCS
fallback are excluded. Live qualification is required before release readiness.
Pro shared allocation supports DMs only, with sender enrollment in Photon and
separate identity linking in Paperclip. Shared channels reserve one project, not
a pool phone number; group admission and publication are disabled. Dedicated
allocation retains one selected number and individually enabled groups.
See [iMessage Photon](connections/IMESSAGE-PHOTON.md) for the implementation
contract, setup, recovery, boundaries, and qualification status.
### Experimental persistent agent conversations
Agent Chat is an opt-in core task presentation (`enableAgentChat`, off by default). Each person has one persistent task-backed conversation per agent and company, with ordinary company task visibility. The shared task composer, transcript, tools, files, and document panel remain the interaction surface. Agents clarify goals and hand substantial execution to linked, assigned tasks; a reply ends a turn without completing the conversation. `/new` starts fresh provider context in the same conversation while preserving visible history and artifacts. Healthy idle conversations wait for a message and do not count as unfinished execution work. See `doc/plans/2026-09-10-agent-chat.md` for the implementation contract.
### Agent chat project handoff (2026-09-11)
Chat supports research and full plan drafting/revision in its existing plan document. On handoff, each ordinary assigned task receives the relevant plan in its own `plan` document, committed with task creation before execution is scheduled. The source plan remains in the conversation. Plan acceptance hands off execution; it never switches the conversation into implementation.
Chat instructions require selecting a suitable project, reusing an existing one where appropriate. The project requirement is prompt-only; ordinary projectless tasks remain supported. New parent relationships beneath conversation tasks are rejected by task services, including direct API creation and reparenting. Existing children remain readable/editable and can be moved elsewhere. The Subtasks panel is unchanged.
The `create_project` runtime tool uses the normal project API with durable idempotency. `list_projects` and `list_project_repositories` support selection. Multiple `repositoryIds` select authorized catalog entries; multiple HTTPS GitHub `repositoryUrls` register existing repositories absent from the catalog. IDs and URLs may be combined, but cannot accompany an explicit `workspace`. URLs do not create repositories on GitHub or grant credentials. Execution uses normal repository access rules. Repository IDs are revalidated against the authenticated run's responsible user and connection grants. Agents should consider proper available repositories, clarify material ambiguity, and use repository-free projects when appropriate for non-code work.
Confirmed project creation appears as a durable card in the shared task transcript, including selected repository links. Tasks are linked inline. Failed creation never produces a success card. Tool evals cover planning/handoff, project/repository selection, retries, permission and mode denials, and ordinary delegation regressions using the production chat directive.
### In-app announcements
An optional announcement card shares product news with board users on opening
or returning to Paperclip. Dismissals persist per user across companies and
browsers within an instance. Operators can disable fetching independently of
telemetry. See [Announcements](ANNOUNCEMENTS.md).
### Agent chat discovery
With Agent Chat enabled, Chat is the first row of the Work section and opens a
secondary sidebar beside the primary nav. It lists every agent you can chat
with: the open conversation first, then your other conversations by recent
activity, then the rest of the roster alphabetically. Terminated agents and
agents you have left are omitted unless you have history with them. Search
filters by name, title, or role; **+** starts or reopens a conversation.
Selecting an agent opens their persistent conversation; it does not reset history
or create a task until the existing first-write flow requires one.
Beside a conversation, the side panel opens on the agent's tasks: one card per
task the agent created, was assigned, commented on, or acted on, newest first,
with the task list's filters and a sort control. The agent's artifacts are a
second card stack available from the panel's **+** menu. Both open in a new tab
so the conversation stays open.
### Combined Inbox + Task List
An opt-in experimental setting (`enableCombinedInboxTasks`, off by default)
folds Inbox into Tasks. The Inbox nav row goes away and its unread badge moves
to Tasks. A Views menu on the task list covers the inbox views (Mine, Unread,
Blocked, Recent, Everything) and the task-status views (All, Active, Backlog,
Done). Bare `/issues` opens the last-used view, defaulting to Mine; links that
carry a task filter open All. Old `/inbox` links redirect to the matching view.