Commit Graph
95 Commits
Author SHA1 Message Date
DottaandPaperclip a9a20fb5c6 feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a persistent identity and a verified active run for
governed access.
> - Customer-success inspection needs broad reads without tenant writes
or secret access.
> - Ordinary board login and database credentials give more authority
than this task needs.
> - This pull request adds a dedicated inspection API and strict
managed-run authority.
> - Cloud owns short grants, human approval, replay protection, and
audit records.
> - The benefit is inspectable access that an operator can disable
immediately.

## Linked Issues or Issue Description

Refs: #15352. This change reuses the persistent Ed25519 identity from
that PR.

**Subsystem affected**

Server authentication, pure resource readers, and the shared wire
contract.

**Problem or motivation**

One internal Paperclip agent must inspect customer onboarding work. It
must not receive owner login or database credentials. Reads must not
create customer sessions, memberships, activity, or read receipts.

**Proposed solution**

Add disabled-by-default run authority and versioned tenant inspection
endpoints. Require strict instance-bound managed-run JWTs on the home
instance. Require exact-operation, single-use Cloud permits on tenants.
Execute a reviewed company-scoped catalog in read-only transactions.
Cloud applies seven-day stack-age eligibility and human exceptions.

**Roadmap alignment**

This is access support for Cloud deployments and governed agent
identities. Bot creation, scheduling, scoring, and reports are separate
work. The maintainer requested this implementation.

## What Changed

- Reuse existing public identity reads and managed private-key
injection. Reject unprovisioned keys, paused agents, ended runs, legacy
signatures, and wrong instances.
- Mount `/api/customer-success/v1` before actor/session synchronization.
Verify Cloud permits and consume them centrally before reading.
- Add explicit company-scoped database readers and bounded instruction,
skill snapshot, run log, workspace, and asset reads. Preserve existing
redactions and file protections.
- Add protocol, security, database immutability, and managed-agent
qualification tests. Add deployment and rollback documentation.

## Verification

- Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck
passed after review fixes.
- The broad local `pnpm test:run` recorded 14,277 passes and four
failures in unchanged suites: two timeouts and two PR-metadata mock
assertions. All three affected suites passed on isolated reruns (36
tests). The complete CI matrix passes at the final head, including every
test lane, typecheck, build, runner checks, canary dry run, and the
security scan.
- Focused inspection, JWT, and existing identity tests pass. The catalog
test compares every public database table before and after reads.
- Inspection and route-contract tests: 22 passed. The coordinated test
runs a real managed process agent against separate home/customer
PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake
through the existing controller, bounded binary file reads, single
challenge consumption across replicas, concurrent grants with a
two-connection pool, scoped SQL audits, append-only runtime auditing,
one-year retention, and unchanged tenant data/files.
- Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST`
pointing at the sibling Cloud build. Normal unit runs skip that optional
private integration.
- Final-head Greptile is 5/5 with no unresolved findings.
- No production deployment or customer inspection occurred.

## Risks

- This adds an authentication boundary. Keep both feature flags disabled
until coordinated staging and canary qualification.
- Cloud support must deploy after this API. Unsupported tenants fail
closed. There is no owner-login or database fallback.
- Existing redactions remain the content boundary. Arbitrary pasted
secrets in readable prose or files may remain.
- Remote files and suppressed provider traces remain unavailable. Wake
can cause normal startup/background writes; test those separately.
- Disable Cloud policy first during rollback. Preserve existing identity
material and Cloud audit history.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, and browser
testing. The session does not expose a more specific deployment ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; broad-run flakes are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:38:15 -05:00
DottaandPaperclip 2d0c138122 Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - The existing assistant connection can read work and create tasks or
comments.
> - It cannot edit tasks, exchange files, or manage normal agent and
project settings.
> - These operations must retain the person's permissions and
Paperclip's execution rules.
> - This pull request adds an explicit operation registry and separately
consented configuration access.
> - Assistants can manage work without receiving credentials or runner
authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: assistant MCP, domain routes, consent UI, storage, and
Product E2E.

**Problem or motivation**

A connected assistant cannot update tasks, maintain documents, attach
files, or configure existing agents, projects, and skills. Users must
leave the assistant for these routine actions.

**Proposed solution**

Add named tools and a restricted API registry to direct connections.
Require separate configuration consent. Reuse domain routes and retry
receipts. File uploads save the attachment when the byte transfer
succeeds.

**Alternatives considered**

Arbitrary REST forwarding would expose administration and credential
operations. Runner impersonation would bypass execution ownership.
Separate upload completion calls add unnecessary client state.

**Roadmap alignment**

Checked ROADMAP.md and related MCP pull requests. This extends the
human-authorized connection from #14933. It does not replace the runner
or introduce agent impersonation.

Companion Cloud routing and directory isolation:
https://github.com/paperclipai/paperclip-cloud/pull/678.

## What Changed

- Add task editing, finish/block, documents/revisions, deliverables,
agent settings/instructions, projects/repositories, and skills/files.
- Add an allowlisted API search/call registry with identical field
restrictions, scopes, and retry identities.
- Add unchecked configuration consent. Existing write grants retain
their current authority.
- Add hashed, expiring file transfer tickets and atomic upload receipts.
No completion call is required.
- Preserve company boundaries, human attribution, active native
execution ownership, and execution review gates.
- Add protocol/domain tests, consent stories, and eight paid Product E2E
workflows.
- Repair two CI fixture races: await cold route setup before assertions,
and wait for asynchronously loaded connection copy. Both fixture suites
pass (24 + 48 tests).

## Verification

- Consent revision: one write-access checkbox controls requested work
and configuration permissions in browser and device flows. All 16
consent tests, UI typecheck/build and token gates pass. Updated
interactive stories cover default approval, opt-out and viewer
restrictions. The paid browser helper uses the new exact label. Real
GPT-5.4 Mini Product E2E passes 2/2 at
`64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission
denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic
retries, cleanup passed; $0.04149375 estimated assistant cost plus
unpriced worker usage. Raw results, usage and source fingerprints are
retained in the worktree. UI and Product E2E typechecks pass.

- Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks
pass; two optional Storybook checks skip. Greptile 5/5 on that head, no
unresolved review threads. Final consent head
`64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing
and Greptile 5/5 with no unresolved threads. The unchanged Cursor
sandbox test had one 10-second timeout, passed in local isolation, and
passed its single CI rerun; the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37554106934).
The existing chat retry-denial browser test had one visibility failure;
its single rerun passes, and the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37542735691).

- Full workspace `pnpm -r typecheck` and `pnpm build` pass at final
runtime source `b2196fae1`. UI token gates pass.
- 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests
pass, including one-connection PostgreSQL OAuth and concurrent upload
retries.
- Paid Product E2E: all eight expanded cases qualified across Mini,
Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell
timed out before application startup. Final affected-case qualification
passes 9/9 on all three models with grader v16, including the failed
cell. Automatic retries disabled; failures, costs, source hashes and
independent durable-state/file assertions are retained in [the
verification
record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md).
- Actual Codex CLI, Claude Code and OpenCode clients completed local
reads/mutations. Codex wrote a report, Claude updated it in a later
conversation, and OpenCode uploaded/downloaded a file with matching
SHA-256 and registered the attachment. Revoking the CLI grant rejects
subsequent bridge initialization.
- Butter staging is verified on final runtime `b2196fae1`
([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)).
A fresh OpenCode workspace fetched the copied invitation, configured
remote MCP, started OAuth and reached real consent with configuration
unchecked. Invalid transfer tickets return 403 through Cloud. Human
approval for the new persistent staging grant is pending; hosted
task/file success is not yet claimed. The final transaction fix is
deployed.
- Full local `pnpm test:run` passed 15,614 general-server tests but
stopped on two macOS timeouts. The heartbeat test passed in isolation;
the existing 40,000-file Git stress fixture timed out again. Its Linux
CI lane passes. Later local full-suite phases did not run after the
timeout; this is not an all-green local full-suite claim.
- Instructions and security limits are in `doc/public-mcp.md`; the saved
plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`.

## Risks

- This expands the experimental direct MCP surface. Explicit schemas and
domain permissions must stay synchronized.
- Migration 0311 adds transfer tickets and upload receipts. Expired
orphan cleanup must not remove committed attachments.
- Configuration requires a new consent request containing that scope;
the single write-access choice controls it alongside work mutations.
Refreshing an old grant does not add it.
- The public directory keeps its original ten tools through the
companion Cloud change.
- Hosted consent/work proof remains the final delivery gate. The PR
stays draft while approval of the new staging grant is pending; code
checks and review are green. Merging is a separate action.

## Model Used

OpenAI Codex (GPT-6, tool use and code execution). The exact serving
model ID and context window are not exposed in this session. Paid
evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001;
claude-sonnet-4-6.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
macOS limitation disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:29:58 -05:00
DottaandPaperclip b508a05c43 feat: add internal agent complaints and suggestions (#15367)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use legacy skills or native runner tools to work on tasks.
> - Those agents can encounter friction that does not belong in the task
thread.
> - A complaint should preserve the raw reaction. A suggestion should
describe an improvement.
> - This pull request adds attributed local storage and both submission
paths.
> - Agents can submit feedback once and continue their primary work.

## Linked Issues or Issue Description

**Subsystem affected**

Server, database, shared contracts, runtime skills, and native runner
tools.

**Problem or motivation**

Agents have no default internal channel for incidental complaints and
suggestions. Sending this feedback through task comments adds noise and
can alter task workflows.

**Proposed solution**

Store free-form feedback in the current instance database. Derive agent,
run, company, and task attribution from active authority. Provide
default legacy skills and provider-neutral native actions. Keep the
instructions close to Warp's MIT-licensed originals.

**Alternatives considered**

Task comments and external Slack delivery add unwanted side effects.
Mandatory suggestion fields and short editorial limits would discard
useful feedback. This release has no listing API, UI, read tool,
automatic triage, or external forwarding.

**Roadmap alignment**

This is a maintainer-requested addition to the existing runtime skills
and runner tool paths. It does not duplicate a listed roadmap milestone.
Searches for complaint tooling, suggestion-box, and agent commentary
found no overlapping public PR or issue.

## What Changed

- Add the company-scoped `agent_commentary` table, shared validation,
and idempotent migration `0310`.
- Add one transactional service and the agent-only POST route. Validate
active authority before writes or replay. Redact known credentials.
Commit a content-free audit with each new record.
- Add `submit_complaint` and `submit_suggestion` to standard, ask, and
planning modes. Keep review, revocation, and completion restrictions.
Store replay identity on the commentary row.
- Mount `complain` and `suggestion-box` by default for legacy agents.
Bundle a dependency-free Node.js stdin helper in the operational skill
and allow its POST through the sandbox bridge.
- Preserve Warp's complaint voice and suggestion guidance, with
attribution and local transport adaptations. Keep source attribution and
MIT notices in each skill's LICENSE, outside runtime instructions.
- Document custom-runtime HTTP use and database inspection. Add
real-database tests and a repeatable live Codex smoke for local and
Daytona execution.
- Pin the lagging-source migration fixture before the identity-repair
migration so later migrations preserve its regression coverage.

## Verification

- Personally ran real Codex submissions in all four environments on
2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at
20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two
rows with company, agent, run, and task attribution, wrote the
continuation marker, exited zero, created no task comments, and left
task status unchanged. Each recorded two content-free activity entries.
- Daytona used production provider hooks, real remote execution and file
transfer, the legacy queue callback bridge, and native private WebSocket
ingress. The current Linux runner was built from `abf47b595`, staged,
and verified against controller contracts. Both sandboxes were confirmed
deleted. This is a focused feedback transport smoke; it does not claim
full Runner E2E catalog or browser qualification.
- The immutable base image and Linux binary digest are recorded in [the
verification
documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification).
The smoke script can save content-free JSON evidence. No credentials or
feedback bodies are in these reports.

| Environment | Runner | Complaint row | Suggestion row |
| --- | --- | --- | --- |
| local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` |
`3db2364d-3e15-4f47-846f-875d3902999d` |
| local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` |
`2d5abe17-dd41-403c-a5ee-4729f2d58921` |
| daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` |
`209681c9-d1e9-4ce1-999e-48fa07692389` |
| daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` |
`77c383d8-a997-49e5-a33e-25c70e15c0b2` |

- Run the local check with `node cli/node_modules/tsx/dist/cli.mjs
server/scripts/verify-agent-commentary-live.ts`. The documentation gives
the Daytona invocation. Both use disposable instance databases and
normal Codex provider usage.
- Repository `pnpm -r typecheck` and `pnpm build` passed after the test
extension. The build includes runner generation, contracts, and replay
checks. The smoke scripts also passed a separate TypeScript check. The
lagging-source migration regression passed. All equivalent current-head
Vitest CI shards passed. The local monolithic `pnpm test:run` invocation
was stopped after CI supplied that coverage; it did not complete
locally.
- Focused tests cover company isolation, spoofing, revoked credentials,
stale ownership, post-finish rejection, concurrent replay, conflicting
keys, atomic rollback, and deletion through existing services. Boundary
tests cover empty text, Unicode, text beyond 8,000 characters, and the
524,288-character ceiling without truncation. Mounting tests cover
Codex, Claude, and sandbox staging. Helper tests cover standalone Node
execution, stdin, invalid UTF-8, redirects, HTTP failure, and its
deadline. Privacy and bridge tests cover successful and rejected
requests.
- Instructions were compared with Warp's originals. MIT notices and
source credits live only in LICENSE files. Native tools preserve
truthful disclosure when asked, without routine announcements.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190)
and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The
later head found the migration-fixture assumption fixed in this update.
On `8a4965164`, all 55 check contexts passed after one browser shard
rerun. Its initial reviewer signoff failure also passed an isolated
local browser run (1 test). Greptile scored that head 5/5 and identified
one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with
four passing tests and a passing smoke-script typecheck. Fresh CI is
pending for this final test-only fix. No commentary production code
changed during verification.

## Risks

- Feedback is internally attributed. It is not anonymous. Existing
redaction removes known credentials, but agents must still omit
sensitive content. Normal provider transcripts can include their
submitted arguments.
- Default skill availability changes for existing legacy agents. Runtime
policy filtering still applies. The helper uses the existing Node.js
runtime with no extra dependencies; custom runtimes can call the HTTP
endpoint.
- Feedback is removed with its run, agent, or company. Task deletion
clears only the issue pointer. Normal database backups include the
table.
- The migration is additive and has no backfill. Writes serialize on the
active run for replay consistency. No server suggestion quota is
imposed.

## Model Used

OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a
reported 258,400-token context window. Capabilities used: repository
inspection, code execution, and live runtime verification. No subagents
were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:14:06 -05:00
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:35:16 -05:00
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00
Devin FoleyandPaperclip 202c2d307e fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Managed deployments prepare empty applications before an owner
claims them.
> - Those applications start database pollers even though no company can
have work.
> - Health probes also query SQL, so idle databases cannot remain
suspended.
> - This pull request adds an explicit standby marker for empty,
unclaimed Cloud apps.
> - The existing signed, durable claim resumes normal processing without
restarting the app.

## Linked Issues or Issue Description

**What happened?**

An empty, unclaimed warm application runs recurring chat, email, plugin,
heartbeat, cleanup, and reconciliation queries. Its health route also
opens the database. This prevents idle database compute from suspending.

**Expected behavior**

An explicitly marked unclaimed application should keep its HTTP process
and sandbox provider plugins ready while leaving the database idle. A
successful signed claim should resume normal API behavior and background
processing. Claimed and self-hosted instances should keep their current
behavior.

**Steps to reproduce**

Start an empty Cloud-managed application and leave it unclaimed. Observe
database activity while repeatedly requesting `/api/health`. Before this
change, periodic queries continue without company data.

Related: #15153 reduces allocation during chat polling. This change
suppresses polling only for explicitly marked, empty, unclaimed Cloud
apps.

## What Changed

- Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once
after restoring the persisted Cloud runtime identity. Missing Cloud
configuration, existing data, or a persisted claim leaves normal
processing active.
- Gate recurring database pollers with an in-memory predicate. Keep
startup preparation and sandbox provider plugin loading intact.
- Serve unclaimed health probes without session or database reads and
report `warmStandby: true`. Serve standby pages/assets directly from the
UI router, bypassing session, bearer, tenant, and dynamic handlers.
Refuse API requests and all WebSocket upgrades before authentication can
query SQL or seed company data.
- Exit standby after the existing signed identity assertion commits.
Normal timers resume at their next tick; a restart restores the claim
even with stale provider variables.
- Document the marker, readiness semantics, rollout checks, and
rollback.

## Verification

- `pnpm -r typecheck` passed. A final server typecheck also passed after
adding tests.
- `pnpm build` passed.
- Focused standby, signed claim, restart, health, static/Vite routing,
hostname, HMR, and live-events suites: 72 passed after the review fixes.
Includes real HTTP upgrade admission before/after claim.
- `pnpm test:run` was attempted locally; both superseded runs were
stopped after encountering checkout/platform failures. A clean-checkout
rerun eliminated ancestor skill-directory lookup failures. The
company-skills/runtime-cache families encounter macOS
read-only-directory rename failures (`EACCES`); all three company-skills
failures reproduce on unmodified base `bf14f803d5`. The initial full run
also reported one native runner API test failure; an isolated comparison
on both revisions was blocked by local embedded PostgreSQL startup
failures. The full [Linux CI
run](https://github.com/paperclipai/paperclip/actions/runs/37423263993)
passed on final commit `5016c415ea`, including all server and workspace
test shards, browser suites, typecheck, build, and release canary. This
is not a claim that the full local suite passed.
- Isolated full server with local PostgreSQL: after startup and
connection expiry, 70 health probes, 70 page requests carrying valid
synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70
seconds observed zero app database connections. The signed claim
completed in 62 ms and normal polling resumed (475 database transactions
over 12 seconds). Restart with stale provider variables restored the
durable claim. The latency is local-only, not a provider wake
measurement.
- Apex review: **5/5** on `5016c415ea`, both earlier threads resolved,
no open recommendations.
- No live-provider test or production deployment was performed. An
actual database suspension/resume canary remains required before
enabling the control-plane switch.

## Risks

- Standby health reports HTTP readiness rather than current database
connectivity. The signed claim still requires a durable database write;
claimed health checks retain the SQL probe and 503 failure behavior.
- Pollers resume at their usual intervals. A suspended database may add
claim latency. Validate the real provider before enabling the marker.
- Startup preparation and sandbox plugins remain loaded. New plugins or
background loops must respect the same standby contract.
- The marker is off by default. Remove it or set it to `0` and restart
to roll back. No schema migration or claimed-workspace inactivity policy
changes.

## Model Used

OpenAI Codex, based on GPT-6. The exact serving snapshot and configured
context-window size are not exposed in this session. Assistance included
source review, TypeScript changes, command execution, and PostgreSQL
tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted tests pass; full
local suite limitations are documented above, and full Linux CI is green
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:11:25 -07:00
Devin FoleyandPaperclip 4a999089ef Retry transient failures in dashboard reads (#14873)
## Thinking Path

> - Paperclip shows company activity in the dashboard.
> - The dashboard reads company, task, approval, and cost data.
> - A pooled database connection can close during one of these reads.
> - The driver must reject ambiguous statements because writes may have
committed.
> - These four dashboard queries are known to be read-only.
> - This change retries only the failed read and preserves completed
work.

## Linked Issues or Issue Description

Refs #14773, which correctly removed automatic replay of ambiguous
database statements. Searched existing database and dashboard PRs.
Related #8780 changes pool recycling and logging; #13925 adds dashboard
consistency coverage. Neither provides these per-query retries.

**What happened?**

A dashboard request returns a server error when its company lookup, task
count, approval count, or monthly spend query loses its database
connection. Drizzle wraps the driver's connection error in `cause`.

**Expected behavior**

A transient connection failure gets a bounded retry of the specific
read. Completed reads and budget processing are not replayed. Persistent
outages and non-connection errors still fail the request.

**Steps to reproduce**

Inject a typed `CONNECTION_CLOSED` error into one of these reads. Then
allow the next query to succeed. Before this change, the dashboard
request fails immediately.

## What Changed

- Apply independent retries to the company lookup, task counts, pending
approval count, and monthly spend query. Each callback rebuilds its own
query with the same company scope.
- Extract the existing authentication retry helper as
`retryIdempotentDatabaseOperation`. Preserve authentication behavior and
its existing exports.
- Retain the existing limit of three total attempts with 50 ms and 100
ms pauses. Match typed connection codes through the error cause chain.
- Test later-read failures, unchanged query parameters, retry
exhaustion, missing companies, and errors that must not retry. Document
the boundary.

## Verification

- Final focused dashboard, authentication, and real database wire
suites: 38 tests passed. Six initial recovery regressions failed before
the implementation.
- `pnpm -r typecheck`: passed on the final source.
- Independent review: no actionable findings. The reviewer separately
passed all 38 focused tests and checked the code allowlist, attempt
bounds, pauses, and final error identity.
- `pnpm test:run`: the general-server group completed with 14,738 tests
passed, 13 failed, and 87 skipped. All 13 failures match the previously
reproduced clean-base macOS skill-cache failures. The two test files and
their implementations are unchanged from that baseline. The runner
exited after this group, so the remaining local workspace and serialized
groups did not run. All corresponding Linux CI groups passed on this
commit.
- `pnpm build`: passed on the final source.
- Full CI: 53 successful checks and 2 skips on `6a113529c0`. Greptile:
5/5 on that commit, with no review threads or remaining findings.
- Merge compatibility with master `f2e0f19630`, including #14866: no
conflicts. The five reviewed files are unchanged in the resulting merge
tree.

## Risks

A persistent outage adds at most two retries per covered query. Each
connection attempt retains the configured driver timeout. The change
does not repair the underlying network or database failure. Agent
counts, run-activity queries, and the budget workflow stay outside these
retry boundaries. General database statements and disconnected
transactions are not replayed. There is no schema or authorization
change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and test execution. The runtime does not expose a more specific
serving model identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (38 focused tests; the full
local run has the baseline limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 19:23:47 -07:00
Devin FoleyandPaperclip 6f40e23536 fix(logging): redact cloud authentication headers (#14413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records HTTP requests to help operators diagnose
failures.
> - Cloud requests carry tenant credentials and signed assertions in
headers.
> - The HTTP logger did not redact four of these headers.
> - This pull request adds those headers to the existing redaction list.
> - Operators retain the route, method, and response status without
recording these values.

## Linked Issues or Issue Description

**What happened?**

HTTP request logs could contain cloud tenant credentials, session
identifiers, runtime identity assertions, and cloud control assertions.

**Expected behavior**

The logger must redact these header values for successful requests and
failed requests.

**Steps to reproduce**

1. Create an Express app with the production HTTP logger and redaction
configuration.
2. Send a request with the four cloud headers and distinct test values.
3. Inspect the serialized request headers for responses with status 200,
403, and 500.

**Paperclip version or commit**

Reproduced on `0f14d26123` before this fix.

**Deployment mode**

Server with cloud proxy authentication. The regression test uses an
in-process Express server.

## What Changed

- Redact `x-paperclip-cloud-tenant-token`,
`x-paperclip-cloud-session-id`, `x-paperclip-cloud-runtime-identity`,
and `x-paperclip-cloud-control` in HTTP request logs.
- Test real logger output for HTTP 200, 403, and 500 with mixed-case
request header names. Route 403 and 500 through the production error
handler. Check response bodies, log levels, and server error context.
- Check that the method, route, and status remain available.

## Verification

The `server/src/__tests__/http-log-redaction.test.ts` suite passed,
including all three new cloud-header cases.

- Rebased onto master at `14795136f5`.
- `pnpm exec vitest run server/src/__tests__/http-log-redaction.test.ts`
passed: 59 tests, including all three new response-status cases. The
suite and server typecheck also passed after the error-handler coverage
update.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `373d29e2f1` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced
company-skill cache permission failures, the terminal-workspace cleanup
assertion, a heartbeat feedback timeout, and one process-conversation
timing failure. It was stopped during the general-server stage after
these failures. Remaining general-server tests, other workspace groups,
and serialized-server stages did not complete locally. Earlier
clean-master checks reproduced the cache failures and isolated cleanup
retries passed. The latest-head GitHub suite passed all of these groups.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- These four values will no longer be available in HTTP logs. Route,
method, and status remain available.
- This change applies to new log entries. It does not remove old entries
or rotate credentials.
- No schema, API, or authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:59:15 -07:00
Devin FoleyandPaperclip b2e9e82f05 fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The execution service owns each run and saves messages sent while it
runs.
> - Interrupt must stop the current executor before it delivers those
messages.
> - Remote Grok commands did not register the host cancellation control.
> - A cancelled task run could still write Done and prevent queue
recovery.
> - This pull request connects remote cancellation and revokes cancelled
run writes.
> - Saved input can use the existing queue admission rules after
verified cleanup.

## Linked Issues or Issue Description

**What happened?**

Interrupting a queued message marked a remote Grok run cancelled before
its sandbox stopped. The old run could still post a reply and mark the
task Done. Its saved follow-up remained deferred behind execution
recovery.

**Expected behavior**

Stop revokes run write authority and waits for verified termination.
Saved messages remain durable and enter one successor through normal
admission after cleanup.

**Steps to reproduce**

1. Run a task with `grok_local` in a remote sandbox.
2. Send a follow-up and use Interrupt while the command runs.
3. Let the old command attempt a task status update after cancellation.
4. Observe the task disposition and the saved message queue.

**Paperclip version or commit**

The gap is present in master at `d3e0f0a238`.

**Deployment mode**

Authenticated server with a Daytona sandbox.

Related work: #14028 and #14046 handle bounded continuation. #13291
covers infrastructure interruption and verified remote cleanup. #13332
addresses atomic recovery holds. This change handles direct Grok
operator cancellation and stale task writes.

## What Changed

- Register remote Grok cancellation before preparation. Keep command
ownership until the host confirms sandbox termination.
- Reuse the sandbox cancellation boundary for the direct CLI invocation.
Reject fresh attempts after cancellation and preserve workspace restore
failure evidence.
- Reject writes from cancelled task JWTs and runs with a pending stop.
Preserve diagnostic reads and existing conversation error codes.
- Recheck run authority under a database lock before task updates and
interaction responses commit.
- Preserve authorized handoffs that stop their own run. Only the
server-issued stop receipt for that request permits the final task
update.
- Add tests for hung commands, unverified stops, early cancellation,
copy-back failures, late Done, late interaction responses, authorized
handoffs, exact lease receipts, and one queue successor across
concurrent restart sweeps.
- Document the cancellation and write-authority contract.

## Verification

- Targeted adapter, cancellation-boundary, authentication,
queued-message, interaction-service, and activity-route tests passed.
The expanded run passed 214 tests; one new test had an incomplete
fixture. After correcting the fixture, all 8 selected follow-up cases
passed.
- `pnpm -r typecheck`: passed on
`179c86caf1bf0d89914a503d46e24af7e4b8c557`.
- `pnpm build`: passed on the same commit.
- `pnpm test:run`: the general-server group completed with 13,521
passed, 99 skipped, and 18 failed tests. It then stopped, so the
remaining local groups did not run. Five Slack, email, and wake-batching
failures passed on focused reruns after correcting the local
environment. The remaining 13 failures reproduce as `EACCES` on rename
in unchanged skill-cache code on macOS. Two custom-image suite setup
hooks also failed to start embedded PostgreSQL after the machine
exhausted shared-memory slots; all 31 tests in that file passed on rerun
after the local resource issue was resolved. CI covers all test groups.
- CI: 53 checks passed and 2 were skipped on the latest commit,
including the aggregate verification gate. The last server shard passed
on its single rerun after a preview-server startup timeout. The affected
file also passed locally with 28 passed and 3 skipped.
- Greptile: 5/5 on the latest commit. Both review threads are resolved.
- No live deployment or staging task mutation has been performed.

## Risks

- Stopping the sandbox can prevent file copy-back. The result preserves
workspace restore failure evidence; termination does not imply restored
files.
- If provider termination fails, the adapter keeps ownership of its
outstanding command and does not acknowledge Stop.
- The write restriction now applies to ordinary cancelled tasks. Reads
remain allowed. Task and interaction checks add a shared run-row lock to
agent mutations. An exact server-issued receipt permits the task request
that stopped its own run to complete its handoff.
- Existing terminal tasks are not reopened automatically. An operator
must correct a historical late Done before its saved queue can continue.
- No schema migration or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 17:07:07 -07:00
Devin Foley 8c6cc7dccf feat(cloud): sync primary-company archive state with the Cloud control plane (#13837)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud runs managed instances and controls their lifecycle
from a control plane.
> - Inside the app, a person can archive a company. On a managed
instance, the Cloud-pinned primary company is the whole organization.
> - Today that archive stays local. The control plane does not learn
about it, so it keeps the instance running and shows the organization as
live.
> - This pull request notifies the control plane when the primary
company crosses the archived boundary, and adds two control endpoints so
the control plane can verify the state and undo the archive during a
restore.
> - The benefit is that an archived organization stops running and shows
as archived, and a restore brings the company back without manual steps.

## Linked Issues or Issue Description

No public issue exists. Description follows the enhancement template:

**What existing behavior does this improve?**

Company archive on Cloud-managed instances. Archiving the primary
company pauses agents and cancels runs, but the hosting control plane
never learns about it.

**Subsystem affected**

Server: company service, cloud-control middleware, instance routes.

**Current behavior**

A person archives the primary company. The instance keeps running. The
Cloud portfolio still shows the organization as live. Unarchive after a
Cloud restore requires manual steps inside the product.

**Proposed behavior**

The company service rings a Cloud lifecycle doorbell when the primary
company is archived or unarchived. The control plane verifies the state
through `GET /api/instance/lifecycle` before it acts. During a restore,
the control plane calls `POST /api/instance/lifecycle/unarchive-primary`
to bring the company back. Self-hosted instances are not affected.

**Reason and benefit**

An archived organization should not keep running, and its hosting
console should show it as archived. The verified read-back keeps the
doorbell a hint: a forged or duplicated ring cannot change state.

## What Changed

- New `services/cloud-lifecycle-sync.ts`: fire-and-forget doorbell `POST
{cloudOrigin}/v1/tenant/lifecycle-changed` with bounded retries. It runs
only on cloud-managed instances, and only for the derived primary
company id. It never blocks or fails the company mutation.
- `services/companies.ts`: ring the doorbell after a committed archive
or unarchive transition, in both the `update()` status-patch path and
`archive()`.
- New `GET /api/instance/lifecycle`: reports the primary company id, its
status, and how many other companies are not archived. Bound to a
`lifecycle:read` Cloud control assertion; board members can also read
it.
- New `POST /api/instance/lifecycle/unarchive-primary`: idempotent
unarchive of the primary company as a system actor. Bound to
`lifecycle:unarchive-primary`; requires instance admin otherwise.
- `middleware/cloud-control.ts`: a closed endpoint→method→action table
replaces the single hardcoded endpoint. Each assertion still authorizes
exactly one action on one endpoint.
- `services/cloud-instance.ts`: the primary-company id derivation moves
here as the single definition; `middleware/auth.ts` delegates to it. The
derivation itself is unchanged.

## Verification

- `pnpm exec tsc --noEmit` in `server/` is clean.
- `pnpm exec vitest run src/__tests__/cloud-lifecycle-sync.test.ts
src/__tests__/cloud-control-task-drain.test.ts
src/__tests__/instance-settings-routes.test.ts
src/__tests__/companies-service.test.ts
src/__tests__/cloud-tenant-company-provisioning.test.ts` — all pass.
- New tests cover: doorbell env-gating, retry bounds, no-throw contract;
the read-back status and sibling count; idempotent unarchive and the
admin gate; non-cloud 404s; control-assertion action binding for the new
endpoints, including cross-action and wrong-method rejection; and a
companies-service test that proves the doorbell rings exactly on
archived-boundary transitions.

## Risks

- Self-hosted instances see no behavior change: without a Cloud signal
the doorbell is a no-op and the endpoints answer 404.
- The doorbell is advisory by design. The control plane verifies through
the read-back before it acts, so a lost or duplicated ring cannot
corrupt state.
- The unarchive endpoint reuses the existing `companyService.update`
path, so agent reactivation and activity logging behave exactly like an
in-product unarchive.

## Model Used

- Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended
thinking and tool use enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-22 17:49:15 -07:00
Devin FoleyandPaperclip 92d4868e79 fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.

## Linked Issues or Issue Description

Refs #13446 and #13719.

**What happened?**

After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.

**Expected behavior**

Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.

**Steps to reproduce**

1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.

## What Changed

- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.

## Verification

- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed c9db03bcab at 5/5. Its only thread is resolved. Full
PR CI passed on that same head:
https://github.com/paperclipai/paperclip/actions/runs/35774449020

## Risks

Small change to error attribution. Unrelated errors may now form their
correct Sentry groups instead of reopening a prior run group. No errors
are filtered or suppressed. No tracing is enabled and no new event
fields are added. No schema or runtime-execution changes. HTTP logs
retain their request and status diagnostics while masking the capability
value. The new SDK job has read-only permissions, no secrets, and an
in-memory Sentry transport.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 12:53:31 -07:00
Devin FoleyandPaperclip c65fc9e3c8 fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures

Include connect timeouts in the bounded retry policy for idempotent actor
synchronization. Handle WebSocket constructor failures through existing
reconnect paths and preserve HTTP polling while realtime is unavailable.
Refresh visible company queries until the socket recovers and clear all
fallback timers on hiding or unmount.

Verify 172 focused tests, server/UI typechecks, UI build, and design token
gates. Full workspace build/typecheck require the unavailable Rust toolchain;
the full test run is tracked separately.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 12:43:25 -07:00
DottaandPaperclip f589660ec0 feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Routines turn scheduled work and external events into tasks for an
assigned agent.
> - Webhook setup was disabled, and actor authentication rejected valid
webhook bearer keys.
> - Operators need to connect and test a sending app before events can
start work.
> - This pull request adds a guided setup with durable connection tests
that cannot dispatch a task.
> - It keeps trigger management, execution tasks, and activity within
the routine.
> - The benefit is a webhook that can be configured, verified, and
operated from one place.

## Linked Issues or Issue Description

Fixes #11937.

Related: #13216 adds provider-specific Sentry support. This PR addresses
general routine setup and ingress. #6841 addresses legacy secret
bindings; this PR retains the existing secret service.

**Current behavior**

Webhook creation is disabled. Bearer deliveries can fail in agent
authentication before the routine checks its key. Setup has no safe
connection test. Runs and Activity send the operator away from the
routine.

**Proposed behavior**

Choose a schedule or a webhook. Follow the setup steps, copy credentials
or complete agent instructions, and test delivery without creating work.
Finish setup to allow future events to start tasks. Edit or remove
compact trigger cards, undo removal, and inspect tasks and activity
inside the routine.

**Reason and benefit**

An operator can verify credentials and delivery before enabling
automatic work. Durable setup state survives refreshes and restarts.
Retry receipts prevent an old test event from starting work after
activation.

## What Changed

- Add a production trigger wizard using reusable Slack setup navigation
and footer components.
- Add schedule and webhook choices, one-time credentials, agent
instructions, and live connection feedback.
- Persist pending setup, test delivery receipts, connection status, and
reversible trigger removal.
- Keep setup checks free of routine runs, tasks, and agent wakeups.
Preserve delivery idempotency after activation.
- Add compact trigger cards, inline editing, key rotation, pause
controls, removal, and Undo.
- Keep Runs and Activity in the routine. Use the shared task list and
compact activity rows.
- Permit only exact public delivery POSTs through actor authentication.
Retain webhook authentication, JSON-object validation, and log
redaction.
- Add production-backed Storybook states and focused server, database,
and UI coverage.
- Document signing modes, setup checks, retries, rotation, HTTPS
ingress, and navigation.

## Verification

- Full workspace typecheck, build, and token gates passed on the rebased
branch. Storybook also builds.
- Focused routine, middleware, logging, shared wizard, and UI coverage
passes on the rebased branch: 195 tests across 14 files. The migration
passed on a fresh PostgreSQL database and on two repeated applications.
- Browser testing used the real app, database, and a deterministic
process worker through Tailscale HTTPS and the current Cloud proxy code.
- Verified rejected keys, safe setup deliveries, persisted state after
restart, activation, retry deduplication, key rotation, schedule
editing, removal, and Undo.
- Fresh bearer and GitHub-signed deliveries created tasks that the
worker checked out and completed. Runs and Activity stayed within the
routine.
- Current Cloud ingress tests passed. Public delivery POSTs passed
through without a browser session; management routes remained gated.
- All 54 current-head PR checks pass, including general and serialized
tests, all eight browser E2E shards, typecheck, build, runner checks,
security checks, and the canary dry run. Two optional Storybook jobs are
skipped by workflow conditions.
- Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review
threads. The stale connection-status finding is fixed and covered by a
regression test.
- No production deployment was performed.

## Risks

- Migration 0281 adds three trigger columns and a test-receipt table. It
is additive and safe to reapply. Apply it before running the new server.
Existing triggers remain live by default.
- Requests without delivery IDs are new events after activation. Senders
must reuse an event's delivery ID for retries.
- Completed webhooks keep normal dispatch behavior. Their management
connection check can start work; the UI states this.
- Removing a trigger archives it. Undo restores the URL and credentials.
Permanent deletion remains available through the existing API.
- Public ingress must remain restricted to the delivery POST route. The
tenant verifies credentials. Cloud sleeping-stack behavior is unchanged.
- Shared setup components also serve Slack. Existing setup contracts and
navigation tests cover that integration.
- Senders must use application/json with an object. Other media types
receive 415.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 10:49:05 -05:00
3ff3b34e15 Return safe client errors for malformed JSON requests (#13660)
Malformed JSON requests currently reach generic crash handling and return 500 before any route handler runs. Classify the specific Express parser error as a 400 with a constant response, preserving unrelated server error reporting.

Carry forward the original three commits from #7410, preserve contributor credit, and add request privacy and negative regression checks. Document the API response. Fixes #7390.

Validation: 80 focused tests, direct server typecheck, and complete Linux CI passed. Greptile 5/5. Full local typecheck/build require missing Rust tooling; local test environment failures are documented in the PR.

Co-Authored-By: developers-universe-1 <madelynreyes2026@gmail.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 22:12:51 -07:00
Devin Foley a59f5a8adc fix(server): stop reporting expected managed-cloud transients to Sentry (#13323)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server reports crashes to Sentry so operators can find real
faults
> - Three expected conditions report as crashes: a client that closes
the connection mid-request, one stale pooled database socket after a
pooled endpoint recycles, and the short boot window where a supervised
cloud stack runs a new app image before its migration runner has caught
up
> - These events arrive in the hundreds and bury real errors
> - This pull request classifies each condition as expected and stops
the Sentry capture for exactly that condition, with behavior unchanged
everywhere else
> - The benefit is a Sentry feed where each event is a real fault

## Linked Issues or Issue Description

**What happened?**

Three noise classes fill the backend Sentry project on managed cloud
fleets:

1. `Error: aborted` (ECONNRESET) reports as a 500 crash when a client
closes the tab or loses its network mid-request. Observed 18 times in
one week from routine client disconnects.
2. `Error: write CONNECTION_CLOSED ...` reports from many query paths
after a pooled Postgres endpoint suspends. The existing single retry in
cloud actor resolution still fails, because a suspended endpoint kills
every pooled socket at once and the one replay draws another dead
socket.
3. `Error: PostgreSQL has pending migrations (...). Refusing to start`
reports from every supervised stack during a fleet upgrade. The
supervisor delivers the new app image before it runs the migration
runner, so each stack crash-loops briefly by design. One fleet roll
produced 329 events (11 per container).

**Expected behavior**

A client disconnect ends the request quietly. A transient dead socket is
replayed until a live socket answers. A supervised mid-upgrade boot
refusal logs and exits nonzero without a Sentry capture, while the same
refusal on a self-hosted deployment keeps reporting.

**Steps to reproduce**

1. Abort an HTTP request mid-flight: the error handler reports a crash
to Sentry.
2. Suspend a pooled Postgres endpoint under an idle server, then issue
two quick requests: the first replay can draw a second dead socket and
surface `CONNECTION_CLOSED`.
3. On a deployment with `PAPERCLIP_CLOUD_API_ORIGIN` set, add a
migration file without running the migration runner and boot: the
refusal reports to Sentry.

**Paperclip version or commit**

master (0e14c61da)

## What Changed

- `server/src/middleware/error-handler.ts`: a request abort (`Error:
aborted` with `ECONNRESET`) ends the response with status 499 and skips
crash reporting and telemetry.
- `server/src/middleware/auth.ts`: `retryOnTransientDbConnectionError`
replays up to twice with a short pause, so a pool-wide recycle does not
defeat the retry.
- `server/src/startup-refusals.ts`: pending migrations on a database
with applied history classify as a supervised-transient refusal
(`schema-migration-pending`). The capture skip applies only when
`PAPERCLIP_CLOUD_API_ORIGIN` is set. A wiped journal beside real tables
keeps reporting. Self-hosted behavior is unchanged.
- Tests updated and added for all three behaviors.

## Verification

- `pnpm vitest run src/__tests__/startup-refusals.test.ts
src/__tests__/cloud-tenant-transient-db-retry.test.ts
src/__tests__/error-handler.test.ts` in `server/` — 25 tests, all pass.
- The abort test asserts no `captureException` and no telemetry crash
track.
- The refusal tests pin all three classifications: never-migrated,
pending-with-history, wiped journal.

## Risks

- Low risk. The abort path only triggers on the exact `aborted` +
`ECONNRESET` pair; every other error keeps reporting.
- The refusal reclassification suppresses a capture only under a cloud
supervisor. If an operator breaks a migration runner, the supervisor's
own deploy gates surface it; self-hosted deployments still report.
- The retry widening adds at most ~150 ms before a genuine connection
fault surfaces.

## Model Used

Claude (Anthropic) — claude-fable-5 (Claude Fable 5), Claude Code
harness, extended thinking with tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-12 11:15:43 -07:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
Devin Foley daea92b647 feat(server): accept a Cloud control assertion on the task-drain endpoint (#13125)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server has a task-drain admission hold so operators can stop new
agent work and wait for quiescence before maintenance
> - Cloud deploys restart tenant containers, but the Cloud control plane
has no sanctioned credential for the drain routes, so agent runs are
killed mid-restart
> - The only Cloud credential this server trusts is the runtime identity
assertion, deliberately scoped to the one-time bootstrap health call
> - This pull request adds a disjoint, action-bound Cloud control
assertion accepted only on the task-drain endpoint
> - The benefit is that Cloud can hold new work and drain a stack before
it restarts the container, through the same authorization and audit
paths a human operator uses

## Linked Issues or Issue Description

Refs #12485 (the task-drain admission hold this makes reachable for the
Cloud control plane).

**Problem or motivation**

Cloud deploys restart the container without stopping agent work first.
The task-drain hold from #12485 exists for exactly this, but its routes
require instance-admin board authority. The Cloud control plane holds no
such credential: the runtime identity assertion is accepted only on `GET
/api/health`, by design. So in-flight runs die at every deploy.

**Proposed solution**

A second, deliberately disjoint use of the same Cloud signing key
(`PAPERCLIP_CLOUD_RUNTIME_IDENTITY_JWKS`): a control assertion with its
own JWS type (`paperclip-cloud-control+jwt`), its own audience, an
`action` claim, a request id, and a short maximum lifetime. A new
middleware accepts the `x-paperclip-cloud-control` header only on
`/api/instance/task-drain`, binds each method to one exact action
(`task-drain:read` / `task-drain:start` / `task-drain:stop`), verifies
the assertion against the configured JWKS and
`PAPERCLIP_CLOUD_STACK_ID`, and installs a synthetic instance-admin
board actor so the existing route authorization, validation,
transactional audit, and activity publishing run unchanged (audit rows
record actor id `paperclip-cloud`). The header is rejected with 400
anywhere else, so it can never become an ambient credential. The board
mutation guard exempts the new `cloud_control` source exactly like the
other non-browser lanes.

**Alternatives considered**

Widening the existing runtime identity middleware would conflate a
one-time bootstrap claim with a repeatable management credential and
weaken both. A per-stack minted instance-admin API key would work with
no auth change but adds a long-lived privileged credential per tenant to
store and rotate. The action-bound short-lived assertion keeps
authorization per-call and stateless.

**Additional context**

Self-hosted instances have no `PAPERCLIP_CLOUD_STACK_ID` and reject
every assertion — the feature is inert off Cloud. A runtime identity
token cannot replay as a control token or vice versa (disjoint `typ` and
`aud`, covered by tests). The Cloud-side caller (drain before deploy,
bounded quiescence wait) lands separately in the Cloud control plane.

## What Changed

- `server/src/services/cloud-runtime-identity.ts`:
`verifyCloudControlAssertion` plus the control
header/audience/type/action constants, reusing the existing JWKS
resolution, JWS parsing, and lifetime discipline.
- `server/src/middleware/cloud-control.ts` (new): accepts the header
only on the task-drain endpoint, per-method action binding, installs the
synthetic instance-admin actor on success, 401 on invalid assertions,
400 anywhere else.
- `server/src/app.ts`: mounts the middleware directly after the actor
middleware, so a valid assertion replaces whatever actor the request
otherwise resolved to.
- `server/src/middleware/board-mutation-guard.ts`: `cloud_control` joins
the non-browser exemptions.
- `server/src/types/express.d.ts`,
`server/src/services/authorization.ts`: `"cloud_control"` added to the
actor source unions.

## Verification

- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/cloud-control-task-drain.test.ts
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/heartbeat-task-drain.test.ts
server/src/__tests__/heartbeat-scheduling-suppression.test.ts
server/src/__tests__/cloud-runtime-identity.test.ts` — 87 tests, all
passing.
- `pnpm --filter @paperclipai/server exec tsc --noEmit` reports no new
errors against the base commit's known pre-existing set.
- The new suite covers: acceptance per method, cross-action rejection,
unknown-action rejection, runtime-identity-token replay rejection,
wrong-audience rejection, wrong-stack and self-hosted rejection, expiry
and oversized-lifetime rejection, unknown-key rejection, request id
validation, endpoint containment (400 elsewhere, 400 on unbound
methods), pass-through without the header, and the mutation-guard
exemption.

## Risks

Low risk, additive. No behavior changes without the header; the header
grants nothing outside the one endpoint; each assertion authorizes one
action for at most five minutes; the existing route-level validation,
queued transitions, and audit writes are unchanged. The browser-facing
Cloud proxy strips Cloud headers, and possession of the shared
tenant-session token cannot mint an assertion (signing key never leaves
Cloud).

## Model Used

Claude (Anthropic) — Fable 5 (`claude-fable-5`), extended thinking,
agentic tool use via Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(module doc comments carry the contract)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-10 11:56:28 -07:00
DottaandPaperclip 889947c238 feat: add experimental native chat connectors (#13038)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also ask agents for work in their existing chat tools.
> - Each external conversation needs one task and a current authorized
source.
> - Retries, Stop, and provider failures must not duplicate work or
expose private data.
> - The first chat PR establishes the opt-in provider and data
contracts.
> - This PR adds experimental channel integration and its durable
control plane.
> - Users can request work from connected channels and inspect delivery
in Paperclip.

## Linked Issues or Issue Description

Refs #13100 and #13092. This is the second of exactly two chat PRs.
Foundation #13100 is merged and changed 143 files. Runner prerequisite
#13092 is also merged. This PR changes 400 files against master, below
the 500-file review limit. It contains no wireframe images or HTML
galleries.

## What Changed

- Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat
connections. Keep chat disabled unless the operator enables experimental
chat connectors. Preserve the production GitHub tool connection and its
normal setup path.
- Bind each provider bot identity to one immutable Paperclip agent. Bind
each admitted external conversation to one task. Paperclip owns tasks,
runs, permissions, and audit records.
- Add durable admission, per-conversation queues, questions, task
controls, progress, final replies, images, files, and delivery receipts.
Board comments remain internal unless explicitly sent to the channel.
- Check current identity, provider reach, resource access, credentials,
runtime generation, and exact source before provider effects. Keep
private responses private. Never send raw reasoning, private logs,
credentials, or tool arguments.
- Hold uncertain sends for explicit audited resolution. Make Board
Send-to-channel atomic and idempotent. Keep reconnect and setup
credentials in Paperclip secret storage.
- Preserve current native-runner authority across retries, lost
acknowledgements, and recovery. Keep immutable input and completion
contracts separate from newer user input. Receipt reconciliation cannot
launch a provider.
- Reconcile chat close/new ordering and provider-effect lock order.
Audit resource access changes in the same transaction. Submit only the
selected resource from each UI toggle so stale pages cannot undo
unrelated access changes.
- Drain Codex stdout before certifying process exit. Bound the drain
with the existing shutdown grace. Preserve observed terminal authority
without treating an undrained process as successful or reusable.
- Incorporate master `018ca5da` with its ACP Stop, mobile task layout,
runner packaging, and official lock changes. Preserve dedicated
chat-answer continuations in both directions when ordinary queued
comments are adopted after Stop.
- Fence late adapter readiness behind an earlier Stop for the same run.
Preserve verified cleanup for registered adapters. Handle single Stop,
agent pause, duplicate Stops, and failure release without creating a
false cancellation receipt.
- Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact
failed-chat retry authorization and lineage, retired question-source
suppression, and the block on generic recovery that would discard the
admitted source. Fresh deferred input retains its separate promotion
path.
- Incorporate master `2a05b5ed3` and its queue-admission extraction,
simplified transaction ports, and separate runner CI job. Preserve exact
durable receipts, actor separation, and dedicated-answer isolation
through the new module. A failed receipt insert rolls back the
accompanying deferred-wake merge.

## Verification

Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating
master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are
resolved. This successor fixes two test-harness boundaries exposed by
CI: per-case route-module preparation and actual durable-save completion
before intentional runner termination. Production code and all existing
test/turn deadlines are unchanged. [Exact-head Greptile
review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594)
is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable
findings or open review threads. [Fresh exact-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341)
passes **all 24 jobs**, including Build and both required aggregates.
Normal exact-head guarded merge was attempted and rejected by the
remaining branch approval policy: CODEOWNER review is required and no
human approval is present. Normal **squash auto-merge is enabled** as of
September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified;
no approval bypass or self-approval was used. Earlier-head results below
remain historical evidence, not qualification of this successor.

- Final exact-head Linux evidence: 995/995 chat integration cases; 36/36
agent-skills routes; 35/35 runner live-session cases, including real
process kill/resume; 1948 runner Vitest cases with three existing
benchmark/platform guards; 870/870 API-authority cases; and 104 browser
cases with four existing optional skips. Rust, conformance/replay, full
repository build, typecheck, canary, all server/workspace shards, and
both required aggregates pass with normal CI concurrency. Earlier failed
attempts remain recorded below.

- Latest test-only qualification: 141/141
route/permissions/authentication cases pass in separate cold forks, with
plain server types and independent review clear. The real-runner suite
passes 35/35, with plain runner types and independent review clear. A
controlled premature-save acknowledgement fails as expected; matching
ownership/effect/process evidence, rejected saves, real turn outcome,
test abort, and pre-kill liveness are covered. No local reproduction of
the original CI scheduling failure is claimed. The preceding [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34479680858)
passes 21/24 jobs, including all 995 Linux chat cases and browser
aggregate (104 passed, four existing optional skips); only Build, the
skills serialized shard, and the required verification aggregate fail.
Its exact-head Greptile review was 5/5. Both failed job logs are
retained.

- Final fixture qualification: all eight focused Discord cases and all
995 chat integration cases pass. The exact modal statement/PID is
observed before taking the real connection lock; the test then proves
its actual blocking relationship before mutation. Original SQL
execution, provider behavior, negative assertions, and 1s/15s timeouts
remain unchanged. Independent review is clear and test/production hashes
remain frozen. The preceding [CI
attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777)
passed 22 jobs, including Build/runner, typecheck, canary, all other
test shards, and browser aggregate (104 passed, four existing optional
skips); the two fixture failures and failed verification aggregate
remain recorded, not relabeled as a pass.

- Current queue-module composition: 308/308 recovery/batching/queue/Stop
tests; 995/995 full chat integration; 89/89 module tests, including real
PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary
tests; plain server and UI types. All four actual local process/ACP
browser paths pass in 1.4 minutes. Fresh databases, no skips or retries,
stable reviewed source hashes. The initial boundary failure is retained;
its no-op service wrapper was removed without changing recovery context
or weakening the check. An exploratory standalone test-directory
typecheck fails because its new upstream transformation config is not a
standalone typechecking project; standard CI/build does not invoke it,
and no configuration was weakened to suppress those diagnostics.

- The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed
[all 24 CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958)
and exact-head Greptile review at 5/5. Required CODEOWNER review
prevented its normal merge before master advanced again.

- Final extracted-module composition: 307/307 recovery, batching, queue
and Stop-control tests; 995/995 full chat integration; 49/49 module
tests including eight PostgreSQL adapter cases; and 19/19 issue-update
tests. Plain server types pass. All four actual local process/ACP
browser paths pass in 1.3 minutes. Fresh databases, no skips or retries
in these cohorts, frozen source hashes, and independent review clear.

- The preceding head `3e4e1c1c` passes [all PR CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820),
including Build and required `ci / verify` and `ci / e2e`. Both the
original Rust failure and the previously load-sensitive lineage fixture
pass with unchanged Linux concurrency. Master advanced afterward and
required this reconciliation.
- Final master composition: 448/448 focused UI tests, 186/186 adapter
tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI,
server, shared, and adapter types pass. Token gates and diff checks
pass. Independent server and UI reviews are clear.
- Stop-registration regression: both real-service cases fail against
exact `a95` source and pass with the fix. The full corrected
recovery/control suite passes 265/265. Duplicate-owner and failed-Stop
controls also pass. Plain server types pass. The readiness barrier
prevents provider startup without adding an acknowledgment to an already
terminal run.
- Final qualification strengthens terminal-field equality and repeats
both affected cases successfully on a fresh database. All four actual
local process/ACP browser paths pass again in 1.3 minutes, without skips
or retries. The final screenshot shows Cancelled, a paused subtree,
retained input, and no error toast.
- Two new actual-service regressions fail before the merge fix. They
prove that queued-comment adoption could consume a dedicated chat answer
or add unrelated input to that answer. The fixed four-case cohort
passes, including ordinary upstream continuation and adapter Stop
controls. Full recovery passes 257/257. All four actual local
process/ACP Stop browser flows pass in 1.4 minutes, without skips or
retries, on a fresh database.
- The unchanged runner artifact was qualified with 171/171 transport
tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11.
Six controlled reader tests prove the exit/drain repair. Its local
serial Rust workspace passed 546 top-level cases plus two invoked
helpers; the later passing Linux CI supplies default-concurrency
evidence.
- Prior exact-source full chat integration passes 995/995. Settings
regressions cover concurrent stale pages, 501 destinations, pending
state, rejected updates, and explicit retry. These deterministic tests
do not prove live provider behavior.
- Retained failed attempts and their causes are in the [qualification
log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md).
The first merge adapter run timed out while macOS slept for 290 seconds.
Its unchanged repeat passed with a temporary sleep guard. No assertion,
deadline, or CI gate was weakened.

Review commands include `pnpm --filter @paperclipai/server exec vitest
run src/__tests__/heartbeat-process-recovery.test.ts
src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec
playwright test --config tests/e2e/playwright.config.ts
tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh
disposable databases. See the [browser
runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md)
for provider setup and separate live acceptance steps.

## Risks

- This remains experimental. Deterministic tests and bounded live
evidence do not establish every provider feature, tenant, permission
layout, or media shape. Teams work-tenant qualification is still open.
- Failed and uncertain provider effects remain visible and can require
operator action. A transport receipt does not prove recipient
visibility.
- Native controller and runner artifacts must remain compatible.
Preserve lease ownership, terminal authority, source binding, and
quarantine during future changes.
- Access and audit rows commit together, but activity notifications
remain best-effort. This is not a new durable event outbox.
- The PR operation does not deploy a live server, replace its runner, or
change provider permissions. Remaining live qualification is documented
in the [temporary
handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md).

## Model Used

OpenAI Codex assisted with implementation, tool execution, testing, and
review. The work records `gpt-6-astra` assistance. The environment does
not report a context-window size. No private reasoning traces are
included.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:06:45 -05:00
DottaandPaperclip ebaeba40ee feat: simplify agent onboarding and configuration (#13011)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators create agents and configure their runtimes in the board
UI.
> - The old creation flow presents several choices and a large form
before an agent can start.
> - The existing onboarding controls already provide clear provider
connection steps.
> - This pull request uses those controls in a new-agent wizard and
organizes the full configuration pages.
> - Operators can connect, test, save, and assign a first task while
keeping the existing configuration tools.

## Linked Issues or Issue Description

Related: #10974. That earlier open PR also reorganizes agent
configuration. This PR follows the reviewed Storybook designs for agent
creation and the current configuration tabs.

**What existing behavior does this improve?**

Agent creation, provider connection, runtime tests, and full agent
configuration.

**Current behavior**

The creation dialog leads to a large manual configuration form. Provider
login controls differ from onboarding. Environment variables and secret
access appear in separate places.

**Proposed behavior**

Choose a name and adapter. Connect Claude or Codex through the existing
onboarding controls. Configure and test the runtime, save the agent, and
open a task dialog with that agent assigned. Use the same design on the
existing configuration tabs.

**Reason and benefit**

The first setup asks for fewer decisions. The full editor keeps
instructions, skills, runtime controls, secret access, permissions,
keys, and revisions available in clear sections.

**Breaking changes**

The board creation and configuration layouts change. The
test-environment API adds an optional, allowlisted `testCredentials`
field for one-shot probes. Database contracts stay the same. Native ACPX
tests now reject unsupported local platforms before a CLI login can mask
the runtime restriction.

## What Changed

- Added a new-agent wizard with numbered steps, adapter branding,
provider connections, editable model choices, runtime tests, and
confirmation.
- Added Codex app-server, Claude ACPX, and OpenCode runner choices.
- Stored API credentials through existing secret APIs and persisted
references in agent configuration. New setup keys are isolated from
credentials used by existing agents.
- Preserved external-agent invitations beside the wizard, including
optional messages, one-time prompts, and clipboard fallback.
- Added OpenRouter provider and secret bindings for Pi and OpenCode.
- Added adapter-specific prerequisite fields for Cursor, Gemini, Kimi,
and Hermes. Cursor Cloud keys are saved as new organization secrets.
- Fixed Cursor Cloud repository field mapping, omitted empty remote
environment values, and added useful model and repository error
messages.
- Preserved complete MCP assignments when multiple valid profiles
contain more than 250 tools in total. Generated profiles retain exact
tool selectors.
- Added service branding and deployment-aware adapter choices. Cloud
setup offers Claude, Codex, and OpenCode; local native runners require
the experimental setting.
- Made the agent list responsive at intermediate widths.
- Applied the reviewed design to the real agent configuration pages.
Kept the instruction editor, skills, and existing mutations.
- Combined secret access and environment variables under one Save and
Discard action.
- Added interactive Storybook screens for setup, configuration,
confirmation, authentication, and test results.
- Fixed Pi provider-error parsing and thinking-effort persistence.
Native ACPX validates Linux x64 on the actual local, SSH, or sandbox
target.
- Redacted the complete transient probe-credential field from HTTP error
logs, including rejected provider names.

## Verification

- Current head `df0292fe6` has a fresh Greptile 5/5 review with no
unresolved findings. All 31 executed CI checks passed, including the
aggregate verification gate and all browser E2E shards. Storybook visual
regression is skipped by its workflow; the local Storybook build passed.

- Browser tests completed real assigned tasks with direct Codex, Claude,
OpenCode, Pi, and native Codex.
- Verified external-agent invitation generation and automatic prompt
copying in the live browser.
- Pi and OpenCode used an existing OpenRouter secret. Browser checks
covered save and reload, instruction edits, skill selection,
environment-variable Save and Discard, and assigned task creation.
- Invalid Claude API credentials remained on the connection step with an
error. A live Pi/OpenRouter invalid-key probe returned a provider
failure and left the user-secret inventory unchanged (zero entries
before and after).
- Full workspace typecheck and build passed after rebasing onto current
master. After review fixes, server and UI typechecks, token gates, and
the full build passed again. Storybook built successfully.
- All 5,542 local UI tests passed. The Cursor Cloud and Pi adapter
regressions passed all 24 tests. Review regressions passed 69 server
tests and all 18 agent-list tests.
- The local full test command ran 6,971 general server tests
successfully. Editing review fixes during that long run caused nine
tests to use stale modules; fresh isolated runs passed. An unrelated
embedded-Postgres fixture hit the host shared-memory limit; its 15
affected tests passed when the fixture groups ran separately.
- Local workspace groups passed after rerunning 18 CLI tests
sequentially to avoid host database limits and parallel-load timeouts.
The local full command stopped at the general server phase, so
serialized server verification comes from the five passing CI shards.
- Browser testing at 390px confirmed that the agent action menu opens
and the page has no horizontal overflow. CI browser E2E shards passed.
- Review the `Onboarding / New agent` and `Agents / Configuration
refresh` Storybook groups. In the real app, create an agent, run its
connection test, save it, assign a task, and reload its configuration.

## Risks

- This changes the main agent setup and configuration UI. Regression
tests cover routing, persistence, secret bindings, and form actions.
- Native Claude ACPX requires Linux x64. Direct Claude works on macOS.
Remote checks execute a bounded platform probe and reject unsupported or
unverified targets.
- A native OpenCode task reached the provider context limit because of
its tool payload. Its provider connection test passed. Direct OpenCode
completed a task. This existing native execution limit is not fixed
here.
- Claude and Codex connection keys use the existing user-secret store.
Other runtime setup keys use distinct organization secrets. Existing
credentials are never rotated. Probes do not store entered keys. Failed
agent creation removes newly staged credentials.
- Cursor Cloud has not completed a live task. Its authenticated account
still needs GitHub repository access. The live run passed MCP
provisioning, remote environment validation, and explicit Auto model
selection before the repository prerequisite blocked execution.
- Generated runtime MCP profiles can exceed the public profile-edit
request limit. They still contain exact catalog selectors and preserve
permission boundaries.
- No database migration, dependency, lockfile, or workflow changes are
included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser automation. The runtime did not expose the exact
model ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-08 14:22:32 -05:00
DottaandPaperclip 1cc45086d3 feat: use the responsible person's GitHub for shared agent operations (#13005)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Several people can send instructions to the same agent and task.
> - A fixed GitHub token in the provider process can keep the first
person's access after another person's message is accepted.
> - Task ownership cannot select credentials for each accepted
instruction or preserve the identity of an operation already in
progress.
> - This pull request records ordered execution identity contexts and
resolves credentials when managed Git, gh, or GitHub tools start.
> - The benefit is automatic personal GitHub access for shared agents,
with durable continuation rules and no teammate credential fallback.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: orchestration, connection grants, database, runtime
adapters, native runners, and run details.

**Problem or motivation**

A shared agent must use the person whose instructions it has accepted. A
queued message must retain its author. A retry or approval without new
instructions must retain the originating identity. GitHub must remain
optional for ordinary work.

**Proposed solution**

Persist execution identity separately from task ownership. Give new
processes a run-scoped broker capability and token-free managed
launchers. Capture identity at operation start. Keep an explicit
dedicated-agent grant as an override. Show redacted diagnostics in run
details.

**Alternatives considered**

Per-task ownership, fixed provider tokens, and mutable repository author
configuration do not handle accepted steering or concurrent operations.
A manual account-selection action would add unnecessary setup to each
turn.

**Roadmap alignment**

This completes the existing Multiple Human Users, MCP Tool Gateway &
Apps, Secrets Manager, and Self-healing Runs capabilities. The
implementation follows the maintainer-approved plan.

Related work: Refs #12843, Refs #12907. Existing proposals #4618 and
#8945 cover per-agent or per-worktree author configuration. This change
instead follows the accepted human instruction across runtime types.
Refs #11831 for governed personal connection delegation; this change
preserves connection audience checks and does not use standing
delegation as a personal credential fallback.

## What Changed

- Add durable, ordered identity contexts and active run references.
Preserve message authors through consolidation, steering, retries,
delegation, approvals, routines, and restart.
- Add an authenticated operation-time GitHub credential broker and
local/remote managed git and gh launchers. Keep personal tokens out of
the long-lived provider process.
- Resolve GitHub gateway and server-side Git operations through the same
responsible-person or dedicated-grant selection rules.
- Make absent and unavailable GitHub credentials non-blocking at generic
startup. Clear host and prior-person credentials. Keep anonymous Git
access where supported.
- Add run-detail identity history and the dedicated-account warning.
Keep task ownership and queue-versus-steer decisions unchanged.
- Preserve personal OAuth declarations through connection edits. Retain
exact selected grants in the gateway.
- Fix continuation races found during real acceptance: verify a warm
owner before credential rotation, and wait for bounded durable runner
suspension before the next run starts.
- Make migrations replay-safe. Retain identity through agent/run
deletion, remove it with its company, and clean terminal launcher
directories before releasing execution environments. Document
coordinated release and rollback.

## Verification

- Full workspace typecheck, build, and token gates passed. The complete
local suite passed in its normal test groups: 17,120 passing tests,
including all 143 serialized server suites. After integrating the newly
merged runner API work, full local typecheck and build passed again,
along with 890 focused integration tests. All 31 checks on the
integrated revision passed, including build, browser E2E, release
registry, canary dry run, typecheck, security and all test suites.
Greptile is 5/5 with all review threads resolved.
- Current focused checks passed: 142 native executor tests, 67 runtime
lifecycle tests, 9 durable identity tests, 75 credential/routine tests,
19 low-trust/resumption tests, and the executable migration replay test.
- Authenticated browser acceptance with two Paperclip users and two
GitHub accounts on one shared native agent passed. Real commits and
pushes followed A → B accepted steering → queued A continuation in the
same saved conversation. GitHub commit author and committer identities
matched all three operations. Both runs succeeded and task ownership
stayed unchanged.
- Real GitHub MCP calls switched from A to B after accepted steering. A
delegated subtask retained its originating identity across a server
restart.
- Disabling B's GitHub connection left ordinary work successful. Managed
gh was unauthenticated and the provider had no inherited GH_TOKEN or
GITHUB_TOKEN.
- The browser displayed run-detail diagnostics and the exact
dedicated-account warning. A final controller-restart check followed by
another-person continuation retained the conversation, selected the
correct GitHub login and Git author, and removed each terminal launcher
directory.
- Company-lifetime migration and all five previously failing CI suites
passed locally (167 tests). Same-token gateway A → B → A and six
broker/launcher boundary tests passed.
- Remote callback, launcher, sandbox, and runtime contract tests passed.
Both native and legacy Codex completed actual Daytona executions on the
integrated revision ([campaign
results](https://github.com/paperclipai/paperclip/actions/runs/34155056509)).
The remote package-manager shim staging regression also passed locally.

## Risks

- Deploy the migrations, server broker, launchers, and runner artifacts
together. Existing processes finish with their original contract. New
managed processes need the broker endpoint for GitHub operations.
- Finish or stop new managed executions before rolling application code
back. Keep the additive schema and identity history during rollback.
- Scripts that require a persistent raw GH_TOKEN must use managed git,
gh, or GitHub gateway tools. Run capabilities authorize code executing
within that run to acquire its current identity; this is not
hostile-code isolation within one execution principal. Managed commands
prevent automatic credential carryover; arbitrary code deliberately
copying a credential is outside that boundary.
- Uncertain steering acknowledgement deliberately holds new credential
acquisition until reconciliation. Already-started operations retain
their captured identity.
- GitHub private access and provider outages can still fail the specific
operation that needs them. Dedicated grant failure does not fall back to
personal access.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, review, shell
execution, and browser acceptance. The exact model variant and
context-window size are not exposed in this session. Tool use included
TypeScript and Rust tests, database integration tests, GitHub CLI, and
authenticated browser control.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-07 14:32:20 -05:00
DottaandPaperclip 7dfc769f3b fix(server): honor proxy trust for forwarded host (#12832)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-04 10:02:28 -05:00
Devin Foley 2177b85eb5 fix(server): retry cloud-tenant auth sync once on a dropped DB connection (#12773)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Managed-cloud deployments authenticate tenant requests through
trusted headers. The middleware syncs the tenant's user, company, and
membership rows on the way through.
> - Pooled Postgres endpoints sometimes close an established connection
under an in-flight query (pooler recycle, compute suspend). The driver
reconnects on the next query, but the statement on the wire fails.
> - In this path a single dropped statement fails the whole request with
a 500. This happened live on 2026-09-03: the idempotent company
bootstrap insert died with `write CONNECTION_CLOSED`.
> - This pull request retries the actor resolution exactly once when the
error chain carries a postgres.js closed-connection code. The sync is
idempotent end to end, so the replay is safe.
> - The benefit is that a routine pooler blip no longer fails an
authenticated request on the entry path.

## Linked Issues or Issue Description

**What happened?**

A cloud tenant request hit the trusted-header authentication middleware
while the pooled Postgres endpoint closed the connection mid-query. The
insert failed with `write CONNECTION_CLOSED <host>:5432` wrapped in a
`Failed query: insert into "companies" …` error, and the request failed.

**Expected behavior**

The driver reconnects on the next query, and every statement in the
tenant sync is idempotent (upserts, on-conflict inserts, deletes; the
write debounce records only after the full sync succeeds). One
in-request retry should absorb the blip and serve the request.
Non-transient failures must keep failing fast.

**Steps to reproduce**

1. Run an authenticated public deployment against a pooled Postgres
endpoint.
2. Have the pooler close the connection while the middleware's tenant
sync insert is on the wire.
3. Before this change the request fails with a 500; after it the retry
serves the request.

**Deployment mode**

Authenticated public (managed cloud), external pooled PostgreSQL.

## What Changed

- `resolveCloudTenantActor` now delegates to the (unchanged) resolution
body through `retryOnTransientDbConnectionError`, which retries exactly
once on a transient closed-connection failure
- `isTransientDbConnectionError` walks the error `cause` chain (drizzle
wraps the driver error) for the postgres.js codes `CONNECTION_CLOSED`,
`CONNECTION_ENDED`, `CONNECTION_DESTROYED`; both helpers are exported
for tests
- New unit test file `cloud-tenant-transient-db-retry.test.ts`:
detection matrix (including a `23505` staying non-transient),
retry-once-then-succeed, no-retry on non-transient,
propagate-on-second-failure

## Verification

- `pnpm vitest run
src/__tests__/cloud-tenant-transient-db-retry.test.ts` — 5 passed
- `pnpm vitest run
src/__tests__/cloud-tenant-company-provisioning.test.ts` — 7 passed
against embedded Postgres, driving the real resolution path through the
new wrapper

## Risks

- Low risk. The retry is bounded to one attempt, gated on three explicit
driver codes, and wraps an operation that is already idempotent by
design. Every other failure propagates unchanged.
- A genuinely down database now fails after two attempts instead of one
— a few milliseconds of added latency on an already-failing request.

## Model Used

Claude Fable 5 (claude-fable-5) via Claude Code, extended thinking with
tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above (none found for connection-retry work in this path)
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (doc
comments; no user-facing docs affected)
- [x] I have considered and documented any risks above
2026-09-03 12:16:55 -07:00
DottaandPaperclip 9dd6526b47 fix(security): harden privileged server boundaries (#12776)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server controls secrets, host files, outbound requests, and
workspace commands
> - A red-team review found cases where restricted callers could cross
these trust boundaries
> - These cases could expose credentials or let untrusted input reach
privileged resources
> - This pull request applies least-privilege checks at each affected
server boundary
> - The benefit is safer agent execution without changing the
private-instance bootstrap contract

## Linked Issues or Issue Description

**What happened?**

Several server paths used authorization, redaction, or content-delivery
rules that were too broad. Restricted agent keys could obtain
company-level operational data. Some adapter and instruction paths could
reach server-owned network or file resources without the required owner
approval.

**Expected behavior**

Paperclip must redact credential values, enforce restricted-key scopes,
guard outbound network access, prevent same-origin script execution, and
reserve host-level file and command controls for authorized operators.

**Steps to reproduce**

1. Configure an authenticated development instance at the parent commit.
2. Exercise the affected APIs with a restricted agent key or a
non-instance-admin company user.
3. Observe that the parent commit returns privileged data or accepts a
privileged operation.
4. Repeat on this branch and observe a redacted response, a safe
download, or an HTTP 403 response.

**Paperclip version or commit**

The findings reproduce from commit `39898ab22` and are fixed by this
pull request.

**Deployment mode**

Authenticated self-hosted server and local development modes.

**Installation method**

Built from source with pnpm.

## What Changed

- Redact generic secret `value` and `token` fields recursively in
structured logs.
- Classify exact and separator-suffixed `KEY` environment names as
secrets in company exports.
- Limit restricted self-identity responses and protect company run, log,
and secret catalog APIs.
- Route HTTP adapter requests through DNS-pinned SSRF protection with
exact private-origin allowlisting.
- Download HTML, SVG, and other script-capable assets with `nosniff` and
a sandbox CSP.
- Require instance-admin access for external instruction roots and
exports that read them.
- Block agent-authenticated host command persistence across supported
workspace runtime shapes.
- Apply the central runtime-management decision before workspace command
controls.
- Keep the documented first-user instance-admin claim contract
unchanged.
- Add regression tests and server-owner configuration documentation.

## Verification

- `pnpm -r typecheck` passes.
- The Node 24 remediation suite passes with 365 tests. It skips 25
environment-gated tests.
- `pnpm build` passes under Node 24.
- `git diff --check` passes.
- The full local runner reaches known macOS-only general-server harness
failures before the serialized route lane. The Linux PR matrix is the
authoritative full-suite gate.

## Risks

- Restricted agent keys now receive HTTP 403 responses from company-wide
run, log, and secret catalog endpoints.
- Script-capable assets now download instead of rendering inline.
- External instruction roots now require instance-admin access.
- Private HTTP adapter endpoints now require an exact origin in
`PAPERCLIP_HTTP_ADAPTER_PRIVATE_ENDPOINT_ALLOWLIST`.
- Public HTTP adapter endpoints remain enabled. Redirects and metadata
or link-local targets remain blocked.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, GPT-5. The exact serving snapshot and context-window size
are not exposed. The model used tool-enabled reasoning, repository
access, code execution, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-03 14:15:32 -05:00
DottaandCodex 0798c77fde Secure Cloud canonical runtime identity (#12766)
Accept and persist Cloud-signed canonical runtime identity before activation, then route absolute self-URLs through the durable runtime identity provider.

Co-Authored-By: Codex <codex@openai.com>
2026-09-03 12:01:47 -05:00
DottaandPaperclip 8c3b8c432a Simplify app connections and enable managed Google access (#12728)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps subsystem gives humans and agents governed access to
external tools.
> - The current connection flow hides Apps behind an experimental gate
and repeats setup text.
> - Google sharing choices and generic MCP permissions do not use one
consistent opening model.
> - Self-hosted installs also need a safe default origin for managed
OAuth without a manual config file.
> - This pull request makes Apps available, simplifies connection setup,
and applies one governed permissions model.
> - The benefit is a shorter connection flow that works on a clean
self-hosted install.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the Apps connection setup flow, managed Google connection
flow, generic MCP connection flow, navigation, and runtime origin
discovery.

**Subsystem affected**

Cross-cutting. This changes `ui/`, `server/`, `packages/shared/`,
connector documentation, and browser tests.

**Current behavior**

Apps require an experimental switch. Setup pages repeat titles and
explanatory copy. Connection names require manual input. Google
credential sharing does not always offer both personal and organization
access. Generic MCP providers do not start with the same permission
choices. Managed OAuth needs a public URL setting even when the request
already has a safe HTTPS origin.

**Proposed behavior**

Apps are available by default. Setup asks only for required permissions
and sharing choices. Paperclip creates conflict-free connection names.
Google apps and generic MCP providers use the same human and agent
access model. Managed OAuth derives a validated same-origin HTTPS URL
when no explicit public URL is set.

**Reason and benefit**

A clean self-hosted install can connect a managed Google app without
hidden setup. Humans can share a service account with their
organization. The shorter flow reduces duplicated choices and setup
errors.

**Breaking changes**

The Apps experimental switch is removed. Existing connection APIs remain
compatible. New connections can receive a numeric suffix when a name
already exists.

No duplicate or related public issue was found.

## What Changed

- Removed the Apps experimental gate and the breadcrumb that leaves the
Apps section.
- Simplified all connection setup pages and moved optional provider
requirements into one small link.
- Added consistent human and agent access choices for Google apps,
Zapier, and generic MCP connections.
- Added organization sharing to Google Workspace credentials while
keeping personal access available.
- Generated connection names automatically and resolved name conflicts
with numeric suffixes.
- Derived a validated public HTTPS origin from the request for
config-free managed OAuth.
- Updated connector contracts, tests, browser coverage, and authoring
documentation.

## Verification

- `pnpm check:token-gates`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts` (273 passed)
- Targeted UI/service regression suite (308 passed)
- Six targeted Playwright connection journeys on a fresh onboarding
instance (6 passed)
- Fresh-install browser proof through Tailscale HTTPS: enrolled with
Paperclip Cloud, connected managed Google Drive, and completed a real
read operation.
- [Exact-head CI
run](https://github.com/paperclipai/paperclip/actions/runs/33669760711):
all 23 matrix jobs passed, including build, typecheck, server,
serialized, canary, and all browser shards.
- Greptile 5/5 on `0ae2a859f269984ee950d0af231a5b09a06f3dfd`, with no
unresolved review threads.

## Risks

Apps are now visible to all operators. The removed experimental flag no
longer hides unfinished app definitions. Managed Google availability
still depends on the Cloud profile rollout and active instance
enrollment. Automatic conflict handling changes only the display name of
a newly conflicting connection.

> I checked [`ROADMAP.md`](ROADMAP.md). MCP Tool Gateway and Apps are
shipped. Connected Apps is planned, and this change improves the
existing shipped connection flow.

## Model Used

OpenAI Codex, GPT-5, with reasoning, browser control, tool use, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-02 14:05:53 -05:00
DottaandDev Agent 25cf079ec5 feat(runner): add Codex-native application integration (#12591)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner package is useful only when the application can start,
observe, and recover a native Codex run safely.
> - Existing direct adapters must keep their current execution and
finalization paths.
> - The application boundary therefore needs additive persistence,
authorization, coordination, and recovery behind an explicit
experimental adapter.
> - This pull request adds that Codex-only boundary without activating
generalized providers, remote environments, or the later task/SDK
surfaces.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, database persistence, adapter utilities, server
native-runtime services, and the experimental Paperclip Runner adapter.

**Problem or motivation**

The already-landed runner package has a qualified Codex path, but the
application needs durable native-run state, guarded runtime selection,
authenticated coordination, tool security, finalization, and recovery
before the experimental adapter can be exercised safely.

**Proposed solution**

Add a Codex-only `paperclip_runner` application path behind the existing
default-off native-runner setting. Bind native state and coordination to
company/run identity, preserve persisted-run recovery, and leave every
direct adapter on its existing legacy execution path.

**Alternatives considered**

The earlier stack boundary introduced a generalized executor and
remote-environment lifecycle here. That made this PR depend on
implementations in higher PRs and changed reusable sandbox behavior
globally. Those pieces are now deferred together to #12592.

**Roadmap alignment**

ROADMAP.md does not list a conflicting native-runner integration
project. This change adds the application boundary for the existing
Runner architecture.

## What Changed

- Added native run/result/finalization/provider-trace persistence,
shared validators, and idempotent migration/replay coverage.
- Added guarded Codex-only runtime selection, authenticated PRP
coordination, recovery, finalization, and interaction services.
- Added run/company-bound tool-gateway authorization, credential
redaction, SSRF protections, and replay-safe behavior.
- Added the explicit `paperclip_runner` adapter behind the default-off
rollout setting.
- Preserved legacy answered-question wake projection and direct-adapter
execution/finalization paths.
- Hardened cancellation so only owned in-memory child processes are
signaled; persisted recycled PIDs/process groups are never trusted.
- Retained the narrow Claude ACPX isolated-context security follow-up
discovered after #12590.
- Deferred the generalized executor, provider ingress, remote lifecycle,
SDK/lab/eval work, release-process changes, and lockfile.

## Verification

- Changed-file delta against `master`: 133 files.
- GitHub Actions is the authoritative verification environment for this
PR.
- Full CI, security, and Greptile review will run on this lowest
unmerged stack PR.
- Local tests/build/typecheck were not run because this checkout is
resource constrained.
- Static diff/reference checks pass, and `pnpm-lock.yaml` is unchanged.

## Risks

- This touches central heartbeat and agent-route code, so legacy
compatibility is the primary risk.
- Runtime selection remains Codex-only and explicit; direct Codex,
Claude, OpenCode, process, HTTP, and plugin adapters remain on their
existing paths.
- Fresh native starts fail closed while the rollout flag is off;
persisted native records remain readable and recoverable.
- Cancellation, company/run binding, tool calls, status decisions, and
completion writes are guarded or replay-safe.

> For core feature work, check [ROADMAP.md](ROADMAP.md) first and
discuss it in #dev before opening the PR. Feature PRs that overlap with
planned core work may need to be redirected.

## Model Used

OpenAI Codex, GPT-5.6, with repository tools, code execution, and
parallel agent review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass — GitHub Actions is
authoritative for this resource-constrained checkout
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI and security gates are green
- [ ] Greptile is 5/5 with no open actionable findings
- [x] I will address all Greptile and reviewer comments before merge

## Stack

- Position: 3 of 5 overall; lowest of 3 currently unmerged
- Base: `master`
- Previous:
[#12590](https://github.com/paperclipai/paperclip/pull/12590), qualified
Claude ACPX runtime — merged
- Next: [#12592](https://github.com/paperclipai/paperclip/pull/12592),
generalized Codex executor, task experience, and developer SDKs

---------

Co-authored-by: Dev Agent <dev@paperclip.ing>
2026-08-31 14:38:38 -05:00
DottaandPaperclip b3343dbd64 feat(connections): add self-serve intent runtime (#12345)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a governed way to request app connections during issue
work.
> - The catalog now describes the available providers and setup methods.
> - A request must become a durable, company-scoped intent before an
operator acts on it.
> - This pull request adds that intent runtime across server, agent,
CLI, and shared contracts.
> - The benefit is a safe bridge from agent need to operator-approved
setup.

## Linked Issues or Issue Description

Refs #11965

This is stack 7 of 11. It depends on stack 6 and replaces another
reviewable part of #11965.

## What Changed

- Add connection intent types, validation, service logic, and routes.
- Add agent runtime tools and CLI support for connection requests.
- Add issue-thread interaction support for connection intents.
- Add runtime, route, adapter, and contract tests.
- Hold the final resolved-continuation row lock through asynchronous
adapter preparation until an actual process spawn, so parking or
reassignment cannot cross that boundary.
- Report Hermes Gateway's first remote run request through the shared
dispatch hook so the resolved-intent lock is released at the true
dispatch boundary.
- Revalidate the addressed user's live non-viewer membership and
connection-management authority for every intent mutation, including
OAuth completion.

## Verification

- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts`
- Result: 176 tests passed.
- `pnpm build`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/heartbeat-stale-queue-invalidation.test.ts` (32 passed;
includes non-process dispatch lock-release coverage)
- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/connection-intents-service.test.ts -t
"addressed-user mutation"` (1 passed)
- `pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/tool-access-service.test.ts -t "binds OAuth
callback completion to the initiating board session"` (1 passed)
- `pnpm --filter @paperclipai/hermes-paperclip-adapter test --
src/gateway/server/execute.test.ts` (23 passed; includes dispatch-hook
ordering and exactly-once coverage)
- `pnpm --filter @paperclipai/hermes-paperclip-adapter typecheck`

## Risks

- A malformed intent could create an unusable operator request.
- Validators and company checks reject invalid or cross-company
requests.
- The final continuation gate holds the issue row lock through adapter
preparation until process or remote dispatch; later operator changes use
the normal active-run interruption path.
- The change does not add a database migration.

> I checked `ROADMAP.md`. This stack continues the existing app
connection work from #11965 and does not duplicate another planned item.

## Model Used

OpenAI Codex, GPT-5. The runtime model ID and context window were not
exposed. The model used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have linked the public source pull request with `Refs #`
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-29 12:08:34 -05:00
DottaandPaperclip 20ccf3f476 feat(apps): add connection grants and delegated identities (#12341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - External tools need explicit identity and access boundaries.
> - Shared connection credentials cannot represent every user-scoped use
case.
> - Grants must stay company-scoped and support safe delegation.
> - This pull request adds connection grants, identity rules, and their
database contract.
> - The benefit is durable control over which identity an agent may use.

## Linked Issues or Issue Description

Refs #11965

This is stack 3 of 11. It depends on stack 2 and replaces another
reviewable part of #11965.

## What Changed

- Add company and user connection grants.
- Add delegated identity and membership rules.
- Synchronize database, shared, server, and UI contracts.
- Register the grant-member replacement route in the OpenAPI surface in
the same layer that mounts it.
- Add migration 0231 with replay-safe guards and coverage.

## Verification

- `pnpm -r typecheck`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/openapi-routes.test.ts` (5 passed)
- `pnpm --filter @paperclipai/db check:migrations`
- `pnpm build`

## Risks

- Incorrect grant selection could expose the wrong credential scope.
- The service enforces company and subject boundaries before credential
use.
- Migration 0231 is generated, ordered after 0230, and safe to replay.

> I checked `ROADMAP.md`. This stack continues the existing app
connection work from #11965 and does not duplicate another planned item.

## Model Used

OpenAI Codex, GPT-5. The runtime model ID and context window were not
exposed. The model used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked a public issue or pull request with `Refs #`
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-29 12:08:33 -05:00
DottaandPaperclip cabc9146d0 feat(apps): add secure remote MCP and PostHog setup (#12339)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give those agents governed access to external tools.
> - Remote MCP setup needs secure endpoint validation and durable
credentials.
> - PostHog needs both browser sign-in and personal API key setup paths.
> - This pull request adds the shared remote MCP foundation and the
PostHog definition.
> - The benefit is a secure and reusable base for later app connection
work.

## Linked Issues or Issue Description

Refs #11965

This is stack 1 of 11. It replaces the first reviewable part of #11965.

## What Changed

- Add guarded remote MCP setup and credential handling.
- Add PostHog OAuth and API key connection methods.
- Add focused server, shared contract, and UI coverage.
- Keep the migration replay-safe and idempotent.
- Give the late-close security regression the same 10-second CI headroom
as the adjacent real-timer handshake test.
- Synchronize fake-timer handshake tests at the exact ensure-session
boundary so real filesystem setup cannot race the fake deadline.
- Drive PTY overflow coverage only after listener registration so
scheduling cannot reorder the test fixture.

## Verification

- pnpm exec vitest run
packages/adapter-utils/src/acpx-engine/execute.test.ts
server/src/__tests__/plugin-worker-manager.test.ts (220 passed; affected
cases also passed five focused stress repetitions)
- `pnpm exec vitest run
packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never leaks a
sandbox-provided value from a late close rejection into logs or the
result"` (1 passed)
- `pnpm exec vitest run
packages/adapter-utils/src/acpx-engine/execute.test.ts -t "never
promotes a late ensureSession resolution|closes a late-resolving real
handle exactly once"` (2 passed)
- `pnpm -r typecheck`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/tool-access-service.test.ts`
- `pnpm --filter @paperclipai/db check:migrations`
- `pnpm build`

## Risks

- Remote endpoint validation can reject configurations that previously
passed without checks.
- OAuth configuration errors can block setup until the operator corrects
the provider settings.
- The migration uses guarded statements so repeated execution is safe.
- The test-only synchronization changes do not affect runtime behavior;
they remove filesystem/fake-clock and listener-registration races
observed under parallel CI load.

> I checked `ROADMAP.md`. This stack continues the existing app
connection work from #11965 and does not duplicate another planned item.

## Model Used

OpenAI Codex, GPT-5. The runtime model ID and context window were not
exposed. The model used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-29 12:08:32 -05:00
Devin Foley 76f7019bdf feat(server): derive hosted-tenant issue prefixes from the company name and follow renames (#12292)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every company has an issue prefix. It is the visible half of each
task and case identifier, and a self-hosted company derives it from the
name it was created with
> - A hosted or managed instance does not use the create-company flow.
The trusted-header auth path claims the tenant company instead
> - That path minted the prefix from a hash of the stack id, and it
wrote a placeholder description that nobody chose
> - So a hosted company showed opaque task IDs such as `PC7F2A-14`, and
a rename never changed them
> - This pull request derives the prefix from the company name on that
path too. It re-derives the prefix when the name changes on a managed
instance, and it rewrites the stored issue and case identifiers so
existing tasks follow the rename
> - It also repairs each company that an earlier build claimed. The
repair runs once, on the next authenticated request
> - The benefit is that task IDs on a hosted instance read like the ones
on a self-hosted instance, and they stay correct after a rename

## Linked Issues or Issue Description

No public issue exists. The description below follows
`.github/ISSUE_TEMPLATE/enhancement.yml`.

**What existing behavior does this improve?**

The tenant company claim in `resolveCloudTenantActor`
(`server/src/middleware/auth.ts`) and the company update in
`companyService.update` (`server/src/services/companies.ts`). Both
decide the
`issue_prefix` and the `description` of a company on a hosted or managed
instance.

**Subsystem affected**

`server/` — REST API and orchestration services. One small hint was also
added
in `ui/`.

**Current behavior**

A self-hosted company gets its issue prefix from its name. "Acme
Robotics"
becomes `ACM`, and its tasks read `ACM-14`.

A hosted or managed instance claims the company through the
trusted-header auth
path. That path wrote a different prefix: `"PC"` plus the first four hex
characters of the SHA-256 of the stack id. The same path also wrote a
placeholder description, `"Provisioned by ... for stack <stack id>."`.

The result is a task ID such as `PC7F2A-14`. It says nothing about the
company.
A later rename of the company does not change it, because nothing
re-derives
the prefix after creation.

**Proposed behavior**

The claim path derives the prefix from the company name, exactly as the
create-company flow does. It writes no description.

On a managed instance, a rename re-derives the prefix. The stored issue
and
case identifiers move with it, so `ACM-14` becomes `NOR-14` when "Acme
Robotics" becomes "Northwind Traders". A rename that keeps the same
three-letter
base keeps the current prefix, including any disambiguating suffix.

A self-hosted instance is unchanged. A rename there still keeps the
prefix the
company was created with.

Companies that an earlier build already claimed get a one-time repair on
their
next authenticated request. The repair re-derives the prefix from the
current
name, re-keys the identifiers, and clears the placeholder description.

**Reason and benefit**

A task ID is the primary handle for a task. People type it, paste it
into chat,
and read it in a URL. On a hosted instance that handle was an opaque
hash, and
it disagreed with the company name that the same user chose during
signup. The
name is the only prefix source a hosted user ever supplies, so the
prefix now
follows it.

**Breaking changes**

Yes, on hosted and managed instances only. A company rename now rewrites
the
stored issue and case identifiers. Links that carry an old identifier
stop
resolving after the rename. The company settings page states this before
the
user saves. The one-time repair applies the same rewrite once to
companies that
carry the old hash prefix. Self-hosted behavior does not change.

## What Changed

- Added `server/src/services/issue-prefix.ts`. It holds the prefix
helpers that
used to live inside the `companyService` closure:
`ISSUE_PREFIX_FALLBACK`,
  `deriveIssuePrefixBase`, `issuePrefixSuffixForAttempt`, and
  `isIssuePrefixConflict`. The companies service now imports them.
- Added `pickAvailableIssuePrefix` to that module. It reads the prefixes
in one
base family and returns the first free candidate. A standalone `INSERT`
can
retry on a unique violation, because each failed statement is its own
implicit
transaction. A caller that already holds a transaction cannot, because
the
violation aborts the whole transaction. Such a caller picks first, then
writes.
- Added `rekeyCompanyIssueIdentifiers` to that module. It rewrites the
prefix of
the stored `issues.identifier` and `cases.identifier` values of one
company in
  the caller's transaction, and it returns the two row counts.
- `companyService.update` re-derives the prefix when the name changes on
a
managed instance, re-keys both tables in the same transaction, and
writes a
  `company.updated` activity entry after the commit.
- `resolveCloudTenantActor` claims the company with a name-derived
prefix and a
null description. The claim retries with the next suffix when the prefix
is
  taken.
- `resolveCloudTenantActor` also runs a one-time repair for companies
that carry
the old hash prefix. An exact-match fence on the update lets a
concurrent
rename win. The repair is idempotent, because its guards stop matching
after
  it lands.
- The rename takes a row lock on the company before it compares anything
against
it, and it re-keys from the prefix it reads under that lock. Only patch
and
environment facts gate the lock, so no stale read can steer the
decision. Two
overlapping updates would otherwise leave a company whose prefix
disagrees with
its own identifiers, in either direction: two renames, where the second
re-keys
from a prefix the first already moved; or a rename plus a stale form
that
resubmits the original name, where the second sees an unchanged name,
skips
re-derivation, and restores the old name on top of the first rename's
prefix.
Only a managed instance takes the lock, and only for an update that
carries a
  name.
- Both helpers compare an exact identifier head instead of a LIKE
pattern. A
  stored prefix is data, so it must never be read as a pattern.
- The company settings page shows a hint under the name field on a
managed
  instance: renaming can change the task ID prefix.

## Verification

Automated tests:

```
pnpm --filter @paperclipai/server exec vitest run \
  src/services/issue-prefix.test.ts \
  src/__tests__/companies-service.test.ts \
  src/__tests__/cloud-tenant-company-provisioning.test.ts \
  src/middleware/cloud-tenant-actor.test.ts \
  src/__tests__/auth-session-route.test.ts \
  src/__tests__/cloud-routes.test.ts \
  src/__tests__/cloud-instance.test.ts \
  src/__tests__/company-branding-route.test.ts \
  src/__tests__/company-cloud-floor.test.ts \
  src/__tests__/companies-route-cross-company-authz.test.ts \
  src/__tests__/companies-route-path-guard.test.ts \
  src/__tests__/company-portability.test.ts
pnpm --filter @paperclipai/ui exec vitest run
pnpm --filter @paperclipai/ui typecheck
```

New coverage:

- `server/src/services/issue-prefix.test.ts` covers the derivation, the
suffix
  ladder, the cause-chain walk of the unique-violation detector, and
  `pickAvailableIssuePrefix` against a stubbed select.
- `server/src/__tests__/companies-service.test.ts` covers a managed
rename
against a real Postgres database: the prefix moves, both identifier
tables are
re-keyed, and the activity entry is written. It also covers a same-base
rename,
  a collision that takes the suffixed candidate, a non-name patch, and a
self-hosted rename that leaves the prefix alone. Two more tests drive
the
overlap cases: two concurrent renames of the same company, and a rename
racing
a stale form that resubmits the original name. Both assert that the
surviving
name's base matches the company prefix and that the stored identifiers
sit on
  that prefix.
- `server/src/__tests__/cloud-tenant-company-provisioning.test.ts`
covers the
claim path and the repair against a real Postgres database: a
name-derived
prefix, a null description, a suffixed prefix on collision, the full
repair,
  a second pass that changes nothing, a description-only repair, and an
  operator-written description that the repair leaves alone.
- `ui/src/pages/CompanySettingsRenameHint.test.tsx` covers the hint on a
managed
  instance and its absence on a self-hosted instance.

The `substring` cast in `rekeyCompanyIssueIdentifiers` is load-bearing
and the
database tests prove it. The driver binds the offset as text. Without
the
`::int` cast Postgres resolves the SQL-regex overload of `substring`,
and every
identifier becomes NULL.

## Risks

- **Re-keying changes existing identifiers and URLs.** This is
deliberate, and
it happens on hosted and managed instances only. After a rename, a link
that
carries an old task identifier stops resolving. The settings page warns
about
  this before the user saves.
- **Identifiers inside comment text are not rewritten.** Only the
`identifier`
columns of `issues` and `cases` move. A task ID that someone typed into
a
  comment, a description, or a document keeps the old prefix.
- **A lost prefix race inside the rename transaction surfaces as a
conflict.**
The rename picks a free prefix and then writes, because a unique
violation
inside a transaction aborts the whole transaction. Two *different*
companies
renamed onto the same base at the same moment can still collide. The
loser
sees its PATCH fail with the unique violation. The write is retryable by
the
client, and the window is a single statement wide. Two renames of the
*same*
company no longer race: the row lock serializes them, and the second one
  re-keys from what the first committed.
- **The rename holds a row lock.** A managed rename takes `SELECT ...
FOR UPDATE`
on its own company row for the rest of the transaction. It is one row,
and no
other path in the transaction locks a company row, so there is no
lock-order
cycle. A self-hosted instance and every non-rename company update never
reach
  the lock.
- **The one-time repair is best effort.** It runs inside a try/catch and
logs a
warning on failure, so it never blocks authentication. A failed pass is
retried
  on the next request, because its guards still match.
- No schema change and no migration.

## Model Used

- Provider: Anthropic (Claude)
- Model: Claude Opus, model id `claude-opus-5[1m]`
- Context window: 1M
- Reasoning mode: extended thinking
- Capabilities used: agentic tool use through Claude Code (file edits,
shell,
  test runs against an embedded Postgres database)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-27 11:34:42 -07:00
Nicky LeachandPaperclip 8f1e3cfe24 feat(observability): add opt-in Sentry error monitoring for the server and the browser (#12190)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server and the browser need clear error reports when an operator
enables external monitoring.
> - Paperclip already uses an opt-in OpenTelemetry pattern for server
traces.
> - Sentry can provide error reports for both runtime paths when the
operator sets one data source name.
> - This pull request adds one opt-in Sentry gate for the server and the
browser.
> - The benefit is faster diagnosis while the default setup sends no
Sentry data.

## Linked Issues or Issue Description

**What is improved?**

Paperclip gains optional error monitoring for server and browser
failures.

**Subsystem affected**

Cross-cutting (server, UI, and shared authentication data).

**Current behavior**

Paperclip has no built-in Sentry error capture for server failures or
browser boundary failures. Operators must inspect local logs and browser
tools.

**Proposed behavior**

When the operator sets `SENTRY_DSN`, the server and authenticated
browser use the same Sentry project. When the variable is absent, both
paths stay inactive. The server loads Sentry dynamically and fails open
when the optional package is absent.

**Reason and benefit**

Operators can inspect runtime errors in one Sentry project. The default
setup remains local and sends no monitoring data.

**Breaking changes**

None when `SENTRY_DSN` remains unset. Authenticated session responses
add the optional `sentryDsn` field.

**Additional context**

The implementation uses built-in Sentry privacy options. It disables
default HTTP context and breadcrumb integrations and keeps
`sendDefaultPii` false.

## What Changed

- Add an opt-in server Sentry gate with dynamic package loading and
fail-open behavior.
- Add the Sentry data source name to the authenticated session response.
- Add an authenticated browser Sentry gate and React error boundary
capture.
- Add tests for server, browser, route, and application error paths.
- Document activation, installation, privacy settings, capture behavior,
and operator controls.

## Verification

- Run `npx vitest run server/src/__tests__/sentry.test.ts`.
- Run `npx vitest run ui/src/lib/sentry.test.ts`.
- Run `npx vitest run server/src/__tests__/auth-routes.test.ts
server/src/__tests__/shutdown.test.ts`.
- Confirm that the full continuous integration suite passes on this pull
request.
- Leave `SENTRY_DSN` unset and confirm that the server and browser gates
stay inactive.
- Set `SENTRY_DSN` and install the optional Sentry packages before a
manual capture check.

## Risks

The operator controls the Sentry project and accepts the data risk when
the operator enables the feature. Error objects can contain messages,
stacks, or cause chains with private values. The default configuration
sends no data because the feature stays off without `SENTRY_DSN`. A
missing optional server package does not stop server boot.

## Model Used

OpenAI Codex, GPT-5, with tool use, repository inspection, GitHub CLI
operations, and code review support.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-26 11:15:45 -07:00
417336f8be fix(workspaces): attach PR preparation to existing branches (#11703)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100): short sentences, one instruction per sentence, simple
approved vocabulary, and the active voice. -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Execution workspaces isolate an agent task from the primary
checkout.
> - Pull request preparation can need a branch that already contains
completed work.
> - The workspace policy could not require an exact existing branch.
> - Workspace cleanup also treated worktree creation as branch
ownership.
> - This pull request adds an exact existing-branch policy and separate
branch ownership metadata.
> - The benefit is safe pull request preparation that preserves every
existing commit and operator-owned branch.

## Linked Issues or Issue Description

**What happened?**

A pull request preparation run could not pin its execution workspace to
an exact existing branch. Workspace reuse and cleanup could also confuse
worktree creation with branch ownership.

**Expected behavior**

The run must attach only to the requested branch in an isolated Git
worktree. It must fail if the branch is missing, busy, or inconsistent.
Cleanup must not delete a branch that Paperclip does not own.

**Steps to reproduce**

1. Create a branch that contains completed work.
2. Configure a pull request preparation task to use that branch.
3. Start the task and observe that the prior policy cannot require the
exact branch.

**Paperclip version or commit**

This behavior reproduces on the base revision before this pull request.

**Deployment mode**

Local development with isolated Git worktrees.

## What Changed

- Add `existingBranch` to the execution workspace policy and shared
validation contracts.
- Require `existingBranch` to use an isolated Git worktree and reject
conflicting branch templates.
- Attach to the exact branch without creating, renaming, resetting, or
deleting it.
- Track branch ownership separately from worktree creation and use that
ownership during cleanup.
- Return HTTP 422 for invalid existing-branch settings on every
issue-producing route.
- Add a bounded repair script for existing pull request preparation
tasks.
- Add focused policy, route, heartbeat, runtime, and ready-comment
tests.
- Document the exact-branch behavior and safety rules.

## Verification

- `pnpm exec vitest run
server/src/__tests__/execution-workspace-policy.test.ts
server/src/__tests__/heartbeat-workspace-session.test.ts
server/src/__tests__/issue-existing-branch-validation-status.test.ts
server/src/__tests__/workspace-runtime.test.ts
server/src/services/workspace-runtime-exposure.test.ts
server/src/services/workspace-runtime-ready-comment.test.ts` passed 335
tests.
- `pnpm -r typecheck` passed for all workspace projects.
- `pnpm test:run` passed 4,431 tests. Two unrelated embedded-Postgres
setup hooks timed out under aggregate load. Their isolated rerun passed
74 tests.
- `pnpm build` passed for all workspace projects.
- The two review regressions passed with 139 unrelated tests skipped.
- All latest-head CI gates passed after one unrelated timing-sensitive
test passed on rerun.
- Greptile scored the latest head 5/5 with no unresolved review threads.

## Risks

- Invalid workspace settings now return HTTP 422 instead of a generic
validation response.
- The exact branch must already exist and must not be checked out by
another worktree.
- The new policy fails closed when it cannot prove branch identity or
ownership.
- This change has no database migration.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex from the GPT-5 family assisted with this change. The
runtime did not expose its exact deployment ID or context window. The
agent used high-reasoning mode, repository tools, shell execution, and
code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 17:23:18 -05:00
DottaandPaperclip a2bf936f9a feat(workspaces): sign the workspace login handoff and gate readiness (#11671)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed worktree services run isolated Paperclip instances with
cloned databases.
> - A reachable service was reported as ready even when its database,
runtime identity, or login path was not usable.
> - The first candidate added verified database seeding and managed
repair in #11665.
> - This pull request consolidates that candidate with signed login
handoff and a complete readiness contract.
> - Post-QA fixes close five defects in repair identity, repair
responses, UI retry, seed journal handling, and seed-source trust.
> - The benefit is a workspace that either opens safely or reports one
accurate recovery action.

## Linked Issues or Issue Description

No public GitHub issue exists for this work, so the problem is described
here.

**What happened**

Managed workspace URLs could return HTTP 200 and report ready while
login failed. QA also found cases where repair used the wrong instance
identity, returned a generic error, left the UI stuck, rejected a safe
journal lag, or trusted a mutable workspace manifest.

**Expected behavior**

Opening a ready workspace signs the board user in to the correct
isolated instance. Provisioning and repair use a registered source and
report a structured recovery state.

**Actual behavior**

Entry depended on a password copied into the clone. Several failure
paths could publish stale readiness, hide the repair precondition, or
trust state that the workspace could modify.

**Additional context**

This pull request includes the commits first published in #11665. That
pull request keeps the original base head for review history. This
consolidated pull request is the merge candidate. Related open readiness
work includes #11575 and #11621.

## What Changed

- Adds a short-lived, signed, single-use login ticket. It binds the
user, workspace, instance, and runtime origin.
- Exchanges the ticket through Better Auth. It creates the session and
cookie through the supported adapter path.
- Adds protected workspace readiness fields for the database, clone
data, login handoff, seed phase, and runtime identity.
- Fails readiness closed when the guest has no company or
execution-workspace binding.
- Binds ticket issuance to the exact cloned user and active company
membership selected for the handoff.
- Verifies every current active board identity through the exact-user
handoff before publication or reuse.
- Gates managed runtime publication on the readiness contract and the
recorded worktree instance identity.
- Refreshes runtime work products from the live runtime row after a port
change.
- Adds one workspace access card with ready, degraded, repairing, and
failed states.
- Uses the runtime response identity for repair. It returns structured
repair precondition errors.
- Lets a valid source journal lag converge during provisioning.
- Binds seed and repair manifests to a source registered outside the
agent-writable worktree.
- Clears recovered UI errors so a successful retry can open the
workspace.
- Makes runtime tests register canonical sources and avoid ports owned
by live host listeners.
- Keeps Vitest on source suites when compiled `dist` trees exist.
- Isolates CLI and adapter tests from ambient AWS and runtime API
environment variables.
- Preserves a 404 response for cross-company workspace ID lookups before
runtime authorization.
- Makes concurrent single-flight coverage independent of
path-canonicalization scheduling order.

## Verification

The following checks passed on the integrated head:

```sh
pnpm -r typecheck
pnpm build
pnpm check:token-gates
pnpm --filter @paperclipai/db check:migrations
```

- The server source lane passed 420 files and 4,953 tests. Five tests
were skipped.
- The CLI lane passed 57 files and 385 tests.
- The database lane passed 26 files and 97 tests.
- The shared package passed 58 files and 506 tests.
- The adapter utility lane passed 640 tests. Four tests were skipped.
- The Claude adapter passed 220 tests. One test was skipped.
- The Codex adapter passed 323 tests.
- The OpenClaw adapter passed 13 tests.
- The OpenCode adapter passed 42 tests.
- The plugin SDK passed 45 tests.
- The workspace runtime suite passed 124 tests.
- The caller-scoped readiness and handoff suite passed 52 tests.
- The workspace provisioning shell suite passed 7 tests.
- The runtime exposure suite passed 17 tests while live host mappings
occupied fixed test ports.
- `git diff --check` passed and the worktree is clean.

The serialized route lane will run in GitHub CI with its normal shards.
No deployment or active-workspace migration was performed.

## Risks

- This is a medium-risk authentication and runtime-readiness change.
- The login ticket uses exact origin, workspace, instance, and user
binding. It has a short expiry and a one-time nonce.
- Runtime publication is stricter. A real readiness, identity, per-user
handoff, or control-plane database disagreement now blocks publication.
- This pull request supersedes #11665 as the merge candidate. Close
#11665 after this pull request merges.
- No new database migration is included. The lockfile and workflow files
are unchanged.
- Deployment and active-workspace migration are intentionally outside
this pull request.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

Claude Opus 5 (`claude-opus-5[1m]`), 1M context, extended thinking, tool
use, and code execution produced the main candidate. OpenAI GPT-5
(`gpt-5`) through Codex, with agentic reasoning, tool use, and code
execution, integrated the post-QA fixes and hardened the test gates. The
Codex context-window size was not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-19 02:37:02 -05:00
1c366a9059 fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100): short sentences, one instruction per sentence, simple
approved vocabulary, and the active voice. -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server authenticates each agent request in `actorMiddleware`
before it attributes chat comments
> - When an agent bearer token failed verification, the middleware
called `next()` with no error and the request continued without an agent
actor
> - The request then fell back to the local user actor, so the server
stored agent replies as user comments
> - The task chat UI renders user comments in blue bubbles, so agent
messages appeared as blue user bubbles
> - This pull request rejects invalid agent credentials with 401 instead
of a silent downgrade
> - The benefit is that agent messages keep agent attribution, and
broken credentials fail loudly with a clear retry message

## Linked Issues or Issue Description

**What happened?**

A user cancelled an onboarding question card. The agent posted a
follow-up reply. The reply appeared in a blue bubble, which the UI
reserves for human messages. The agent run held an expired local agent
JWT. The auth middleware could not verify the token, called `next()`
without an actor, and the request fell back to the local user identity.
The server stored the agent comment as a user comment.

**Expected behavior**

Agent messages always render as agent bubbles. A request with invalid
agent credentials must fail with 401 so the adapter can refresh
credentials and retry. It must not post content under a human identity.

**Steps to reproduce**

1. Start a local Paperclip instance.
2. Give an agent run an expired or malformed agent JWT.
3. Let the agent post an issue comment through the API bridge.
4. Before this change: the comment is stored with the local user
identity and renders as a blue bubble. After this change: the request
fails with 401 and a message that tells the caller to obtain fresh
credentials.

## What Changed

- `server/src/middleware/auth.ts`: a bearer token that fails
verification now produces a 401 `unauthorized` error instead of a silent
fall-through to the anonymous/local-user actor.
- The 401 message states the cause: expired token, unverifiable token,
empty bearer token, missing agent record, agent record in another
company, terminated agent, or agent pending approval.
- The API-key path now also rejects an agent record whose company does
not match the key.
- `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now
writes a `comment id: <id>` marker to the run log for each posted issue
comment, so misattributed comments can be traced to a run.
- `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression
test asserts that a recovered `local-board` comment with a derived agent
author renders as an agent bubble, not a user bubble.
- `server/src/__tests__/agent-auth-middleware.test.ts` and
`packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests
cover each rejection path and the log marker.

## Verification

- Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in
`server/` — 14 tests pass.
- Run `pnpm vitest run execution-target-sandbox` at the repo root — 44
tests pass.
- Run `pnpm vitest run
src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests
pass.
- Manual check: post an issue comment with an expired agent JWT; the API
returns 401 with a retry message and no comment is stored.

## Risks

- Behavioral shift: requests that previously continued as anonymous or
local-user actors after a failed agent-token verification now receive
401. Any caller that relied on the silent downgrade must refresh its
credentials. This is the intended fix, and the adapters already handle
401 with a credential refresh.
- No schema or migration changes. Low risk otherwise.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with
extended thinking and tool use (agent harness with shell, file, and git
tools).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 21:36:11 -07:00
Nicky LeachandPaperclip fdb9a4880d fix(security): route paperclipai CLI guidance through safe npx form (CWE-78) (#11400)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip provides CLI commands and guidance for operators and
agents
> - The `pnpm paperclipai` script can pass argument values through a
shell
> - Shell re-parsing can execute command substitutions inside quoted
values
> - This pull request routes guidance through inert-argv `npx
paperclipai` commands and adds regression coverage
> - The benefit is safer operator guidance across documentation and
runtime hints

## Linked Issues or Issue Description

This pull request fixes a command-injection-class defect in Paperclip
CLI guidance.

**What happened?**

The `pnpm paperclipai <sub> --flag "$VALUE"` form can re-parse argument
values through a shell. A command substitution inside a quoted value can
execute on the host.

**Expected behavior**

Paperclip guidance must pass CLI values as inert argument values.
Host-derived values must not appear in copyable commands.

**Steps to reproduce**

1. Run a Paperclip guidance command that uses the `pnpm paperclipai`
script.
2. Provide a quoted value that contains a command substitution.
3. Observe that the shell can evaluate the substitution before the CLI
starts.
4. Compare the result with the `npx paperclipai` form.

**Paperclip version or commit**

`5670984b75d109950c968542a0111ebb6967f4da`

**Deployment mode**

All deployment modes that show or use the affected CLI guidance.

**Installation method**

Built from source and installed CLI guidance.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Database mode**

Not database-related.

**Access context**

Both.

**Additional context**

The earlier merged PR
[#11343](https://github.com/paperclipai/paperclip/pull/11343) used the
unsafe `pnpm exec paperclipai` form. This fresh PR replaces that
guidance with the safe `npx paperclipai` form.

## What Changed

- Standardize documentation and runtime hints on `npx paperclipai`.
- Remove the broken `pnpm exec paperclipai` guidance.
- Use a static `<host>` placeholder in private-hostname guidance.
- Add regression tests for unsafe forms, continued lines, static hosts,
and offline guidance.

## Verification

- `git diff --check
origin/master...origin/fix/paperclipai-cli-npx-safe-invocation` passes.
- The branch adds `server/src/__tests__/cli-invocation-safety.test.ts`
and updates private-hostname tests.
- CI must run the new tests, typecheck, lint, and build checks.
- Local Vitest execution was not available because this worktree has no
installed Vitest binary.

## Risks

- The change affects operator and agent documentation text.
- The runtime hints now show `<host>` instead of a request-derived host
value.
- No database schema or migration changes exist.
- CI will detect any missed unsafe invocation or type error.

## Model Used

OpenAI GPT-5, exact model ID `gpt-5`, with tool use and code-review
assistance. The model used repository inspection, Git operations, and PR
preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] CI ran the test suites and they pass; local test execution was
unavailable in this worktree
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I addressed all Greptile and reviewer comments before requesting
merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-14 22:11:16 -07:00
Nicky LeachandPaperclip 5ca7b4c1fe fix(security): standardize paperclipai CLI guidance on safe npx path (#11343)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip provides CLI guidance to agents and operators through
documentation and runtime messages.
> - Content-bearing `pnpm paperclipai` examples send arguments through a
shell.
> - Shell evaluation can execute command substitutions in untrusted
argument content.
> - Runtime hostname guidance can also place request-derived content
inside a shell command.
> - This pull request uses `npx paperclipai` for content-bearing
guidance and uses a static hostname placeholder.
> - The benefit is safer copy-paste guidance for agents and operators.

## Linked Issues or Issue Description

**Issue type**
Incorrect information

**Where is the issue?**
CLI guidance in `doc/CLI.md`, `skills/paperclip/SKILL.md`,
documentation, and runtime-generated hints.

**What's wrong?**
Content-bearing `pnpm paperclipai` commands can pass argument text
through `/bin/sh`. Shell command substitution in an argument can execute
before the CLI receives the value.

**Suggested fix**
Use `npx paperclipai` for content-bearing commands. Use a static
`<host>` placeholder when runtime guidance displays the allowed-hostname
command.

## What Changed

- Replace content-bearing `pnpm paperclipai` examples with `npx
paperclipai` across the documentation and agent-facing guidance.
- Update runtime-generated CLI hints to use a static `<host>`
placeholder.
- Add safety notes to `doc/CLI.md` and `skills/paperclip/SKILL.md`.
- Add scans and regression tests for unsafe invocation and hostile
hostname headers.
- Keep fixed lifecycle commands and `pnpm --filter @paperclipai/*` build
commands unchanged.

## Verification

- Run `tsc --noEmit` for the changed server files.
- Run `cli-invocation-safety.test.ts`.
- Run `private-hostname-guard.test.ts`.
- Confirm that hostile hostname headers do not enter shown shell
commands.
- Confirm that the three commits contain the required Paperclip
co-author trailer.

## Risks

- This change updates documentation and diagnostic text across many
surfaces.
- Fixed lifecycle and setup commands remain unchanged.
- The tests fail if content-bearing `pnpm paperclipai` guidance returns.
- The change does not alter the CLI argument parser.

## Model Used

OpenAI Codex, GPT-5, tool use, code execution, and repository review
assistance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-13 16:43:21 -07:00
Constantineandcucurigoo 276730d63e fix(server): recognize cross-package Zod errors (#10168)
## Thinking Path

> - Paperclip validates API request bodies with Zod and converts
validation failures into client errors.
> - The global error handler recognized Zod failures with `instanceof
ZodError`.
> - Monorepo dependency layouts can provide more than one installed Zod
module instance.
> - A valid Zod error from another instance fails that identity check
and falls through as HTTP 500.
> - This pull request keeps the native path and adds a narrow structural
fallback for named Zod errors with an issues array.
> - The benefit is stable HTTP 400 validation semantics regardless of
package-instance identity.

## Linked Issues or Issue Description

Related but not duplicate: Refs #6908. That PR catches `instanceof
ZodError` inside validation middleware and returns 422; it does not
cover errors created by a second Zod module instance, which is the
reproduced failure here.

**What happened?**

An invalid `POST /api/issues/:id/work-products` payload raised a real
Zod validation error but returned HTTP 500 because the error came from a
different Zod package instance.

**Expected behavior**

All genuine Zod validation failures return HTTP 400 with validation
details, independent of module identity.

**Steps to reproduce**

1. Submit a work-product body missing the required `provider`,
`externalId`, and `url` fields.
2. Ensure the route schema is resolved from a different installed Zod
instance than the server error handler.
3. Observe HTTP 500 before this fix.
4. Observe HTTP 400 after this fix.

**Environment**

- Paperclip base: `14f20be92b86a49ff2c35495e5b0fa4d719998ef`
- Deployment: self-hosted, built from source
- Access context: board API
- Adapter scope: not adapter-specific

- [x] I searched open PRs for `ZodError`, validation errors, and
work-product validation and linked related work above.

## What Changed

- Add a narrow `readZodIssues` helper that accepts native Zod errors or
structurally valid cross-package Zod errors.
- Preserve existing HTTP 400 response shape and structured error
context.
- Add a regression for a Zod error object from another module instance.

## Verification

- `pnpm exec vitest run server/src/__tests__/error-handler.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- Full upstream CI test/build/e2e matrix passed.
- Local post-deploy smoke returned HTTP 400 for the previously failing
invalid work-product payload.

## Risks

- A deliberately thrown object named `ZodError` with an `issues` array
will be treated as a client validation failure. The effect is limited to
returning HTTP 400 instead of 500; no authorization or persistence
behavior changes.
- No schema or migration changes.

> This is a bug fix, not roadmap feature work.

## Model Used

OpenAI Codex `gpt-5.6-sol`, with tool use, code execution, repository
inspection, and independent read-only review agents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked related public work and described the bug
in-PR following the bug template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have considered documentation; no user-facing documentation
change is required
- [x] I have considered and documented risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: cucurigoo <cucurigoo@users.noreply.github.com>
2026-08-12 16:05:53 -05:00
e31951a17d feat: Claude agent setup-token login in a sandbox (#11286)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Claude agents that run in a remote sandbox need a safe in-product
login path
> - The existing host login route cannot open a pseudo-terminal inside
that sandbox
> - The login flow must protect the browser code, the login URL, and the
OAuth token at every step
> - This pull request adds the parser, the runner, a Daytona
pseudo-terminal transport, and a guarded, owner-bound session route
behind an injectable transport
> - The route stays inert in the default build and fails closed until a
sandbox provider binds the live transport
> - The benefit is a company-scoped setup-token flow with one-time
secret delivery, redaction, and fail-closed transport checks, ready for
a later staged production rollout

## Linked Issues or Issue Description

**Agent or provider**

Claude Code setup-token login for sandbox agents.

**Why this adapter is useful**

Sandbox agents need a supported way to sign in without host credentials.
An authorized owner completes the browser step and receives the token
one time.

**How the agent is invoked**

When a sandbox provider binds the injectable transport, the server
starts `claude setup-token` through a sandbox pseudo-terminal, sends the
browser code to the matched prompt, and returns the token through the
guarded session route. The default build does not bind the transport. In
that state the start route fails closed with a fixed no-secret `503`. It
does not start a process and it does not hold a sandbox lease.

**Additional context**

The transport is injectable, so each sandbox provider binds its own
pseudo-terminal. This pull request adds the Daytona transport but does
not bind it in the production server. A production wiring needs a lease
manager, a live pseudo-terminal factory, a durable token store, and its
own security review. The route keeps secrets out of logs, activity
details, errors, telemetry, and non-owner responses.

## What Changed

- Add strict parsers for the setup-token URL, the prompt, and the
success token.
- Add a login runner that drives the `claude setup-token` command
through a pseudo-terminal.
- Add the Daytona pseudo-terminal transport and the sandbox plugin
wiring.
- Add a company-scoped, owner-bound login session service with rate
limits, a reaper, cleanup, and one-time token delivery.
- Add the guarded session routes at
`/agents/:id/setup-token-login-sessions/*` behind an injectable
transport. The routes become the live login path only when a provider
binds the transport.
- Keep the start route fail-closed in the default build. It returns a
fixed no-secret `503` and it does not bind `setupTokenLogin`.
- Keep the existing host route `POST /agents/:id/claude-login` in place.
This pull request does not replace it.
- Keep confidential responses behind a fail-closed TLS transport guard
with `Cache-Control: no-store`, and extend redaction for the new fields.
- Export the parser and the runner from the Claude local server entry,
and document the new session routes in the OpenAPI spec.

## Verification

- `pnpm --filter @paperclipai/server exec vitest run setup-token-route
setup-token-session`
- `pnpm --filter @paperclipai/adapter-claude-local exec vitest run`
- `pnpm --filter @paperclipai/server run typecheck`
- Confirm that the pull request checks pass on GitHub.

## Risks

- Low user-facing risk on merge. The default build does not bind the
transport, so the production start route stays fail-closed with a `503`.
The merge does not change the production login behavior.
- When a provider later binds the transport, the flow starts a live
sandbox process and holds a short-lived in-memory secret. Cleanup must
stop the child before it releases the sandbox lease.
- The transport guard fails closed when the deployment does not provide
a trusted TLS path. A wrong proxy allowlist can block a valid request.
- The production wiring is out of scope. It needs a lease manager, a
live pseudo-terminal factory, a durable token store, and its own
security review before the server binds `setupTokenLogin`.

## Model Used

Anthropic Claude Opus 4.8 assisted the implementation. It used extended
reasoning, code execution, repository tool use, and a 200,000-token
context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (the
OpenAPI spec covers the new session routes; no user-facing documentation
needs changes)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-12 13:02:49 -07:00
Devin Foley 2c53437fc9 fix(server): authenticate cloud-proxied browsers on the live-events websocket (#11290)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The UI receives live run/issue events over a websocket at
`/api/companies/:id/events/ws`; the server authorizes upgrades with a
bearer token or a Better Auth session
> - On a cloud-managed deployment, browsers authenticate through trusted
`x-paperclip-cloud-*` headers injected by the managing front door — they
never hold a local Better Auth session, and the Express middleware lane
that understands those headers is not consulted for websocket upgrades
> - Every browser websocket upgrade behind the front door therefore
resolves no identity and is rejected 403: the live-events socket has
never connected on a managed instance, leaving permanent reconnect churn
and console failure noise while the UI silently degrades to polling
> - This pull request adds a cloud-actor lane to the upgrade
authorization, reusing the same trusted-header resolver the HTTP
middleware uses
> - The benefit is working realtime updates on managed instances, an end
to the reconnect churn, and unchanged self-hosted behavior

## Linked Issues or Issue Description

No existing issue. Description follows the bug template:

**What happened?**

On a cloud-managed instance, the browser console shows `WebSocket
connection to 'wss://…/api/companies/<id>/events/ws' failed:` repeating
indefinitely for every company, on a healthy instance. The server
rejects each upgrade with 403 because `authorizeUpgrade` in
`server/src/realtime/live-events-ws.ts` only knows bearer tokens and
Better Auth sessions, while cloud-proxied browsers authenticate via
`x-paperclip-cloud-*` trusted headers (handled only by the Express
`actorMiddleware` lane in `server/src/middleware/auth.ts`).

**Expected behavior**

A browser that authenticates through the trusted cloud headers can open
the live-events websocket for any company in its membership scope,
exactly as it can call the HTTP API for those companies.

**Steps to reproduce**

1. Run Paperclip in `authenticated` mode behind a proxy that injects the
`x-paperclip-cloud-*` headers with a valid
`PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN`.
2. Load any company page in a browser (no local Better Auth session).
3. HTTP API calls succeed; every `/events/ws` upgrade is rejected 403
and the UI retries forever.

## What Changed

- `server/src/middleware/auth.ts`: `resolveCloudTenantActor` now accepts
a minimal `CloudActorHeaderSource` (`header(name)`) instead of an
Express `Request` — `Request` satisfies it unchanged — plus
`cloudActorHeaderSourceFromHeaders` to adapt raw
`IncomingMessage.headers`.
- `server/src/realtime/live-events-ws.ts`: `authorizeUpgrade` gains an
injected `resolveCloudActor` lane, tried before the Better Auth session
fallback in `authenticated` mode. A resolved cloud actor is
authoritative: the upgrade is authorized only for a company in the
actor's membership scope (`companyIds`, the same scope the HTTP lane
grants). Absent/unresolvable cloud headers fall through to the session
path.
- `server/src/index.ts`: wires `resolveCloudActor` through
`resolveCloudTenantActor` + the header shim. The resolver self-gates:
without `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` and a matching trust token
it returns null, so self-hosted deployments never take this path.
- Tests: upgrade authorized for an in-scope company (session resolver
not consulted), rejected for an out-of-scope company, fall-through to
session auth when no cloud actor resolves; header-shim resolution from a
raw lowercased header map including `string[]` values.

## Verification

- `pnpm vitest run server/src/__tests__/live-events-ws.test.ts
server/src/middleware/cloud-tenant-actor.test.ts` — 25 tests pass.
- `pnpm typecheck` in `server/` — clean.
- Not verified live end-to-end: that requires a managed instance running
this build; the direct probe evidence (HTTP authenticated fine, every WS
upgrade 403) matches the code path exactly.

## Risks

Low risk. The new lane only activates when the deployment configures the
cloud trust token and the request presents it; both checks already
protect the HTTP lane. Authorization scope is the same `companyIds` set
the HTTP middleware computes (primary stack company plus the user's real
membership rows). The cloud resolver's user/company materialization
writes are debounced (existing behavior shared with the HTTP lane), so
websocket reconnect storms do not amplify database writes. Self-hosted
instances see no behavioral change, covered by the fall-through test.

## Model Used

- Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code
CLI with extended thinking and tool use (code search, edit, test
execution; diagnosis included live websocket handshake probes against a
managed instance).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-12 11:22:50 -07:00
b847e8b6f6 perf(server): reduce issue detail request overhead (#10414)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI-agent work
> - Opening an issue fans out into several authenticated issue-detail
reads, so repeated work on that path directly affects perceived latency
> - Those reads repeated issue and authorization lookups, returned full
private JSON even when unchanged, and performed non-critical bookkeeping
writes on the request path
> - Interaction reads also performed lifecycle writes even though `GET`
must be read-only
> - This pull request adds request-scoped reuse, private conditional
responses, read-only interaction access, and bounded write debouncing
without crossing actor, request, or company boundaries
> - The result is less database, serialization, logging, and
response-body work while preserving authorization and interaction
lifecycle invariants

## Linked Issues or Issue Description

This is the server-only latency phase. Related work is tracked
separately in #10415 (aggregate view), #10416 (warm navigation, merged
into the base), and #10463 (bundle split). This pull request
intentionally excludes those scopes.

**What happened?**

Opening an issue detail view caused avoidable server costs: repeated
issue and authorization reads within one request, full private JSON
responses when a representation was unchanged, writes during
interaction-list reads, production debug transport setup, and immediate
bookkeeping writes for cloud tenant activity and board-key usage.

**Expected behavior**

All successful JSON `GET /api/issues/:id/*` responses should support
strong private ETags and `304 Not Modified`. Repeated work may be reused
only within the current request. `GET /interactions` must not modify
stored interactions. Non-critical activity timestamps may be debounced
without weakening authentication or stale instance-admin cleanup.

**Steps to reproduce**

1. Start Paperclip in local development or self-hosted server mode.
2. Open one issue and request its detail subresources with the same
authenticated actor.
3. Repeat a successful JSON request with its `ETag` in `If-None-Match`.
4. Observe `304 Not Modified`, no interaction writes from `GET
/interactions`, and unchanged authorization boundaries.

**Deployment mode / installation**

- Local development or self-hosted server
- Built from source
- Core server behavior; not adapter-specific

## What Changed

- Added strong ETags and `Cache-Control: private, must-revalidate` to
successful JSON reads under `/api/issues/:id/*`, including
standards-compliant `If-None-Match` handling.
- Added request-scoped promise memoization for issue and authorization
lookups; no authorization result survives the request.
- Made `GET /interactions` read-only, moved supersession and
terminal-state handling to mutation paths, and prevented plugin callers
from accepting or rejecting interactions after an issue closes.
- Removed the production debug-file logger transport while preserving
development formatting.
- Debounced cloud-tenant activity and board-key `lastUsedAt`
persistence, while keeping stale instance-admin deletion unconditional
and authentication checks per request.
- Added focused tests for ETags, request isolation, authorization
lifecycle behavior, interaction invariants, logger configuration, and
retry-safe debounce behavior.

## Verification

- `pnpm exec vitest run server/src/__tests__/private-json-etag.test.ts
server/src/__tests__/issue-thread-interaction-routes.test.ts` — 2 files,
23 tests passed.
- Focused Vitest run covering request memoization, authorization,
interactions, plugin orchestration, logger, cloud tenant, board auth,
and issue services — 9 files, 264 tests passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check origin/master...HEAD` — passed.
- Scope guardrails: 21 changed files under `server/src`; no lockfile,
workflow, migration, UI, aggregate-view, or bundle-split changes.

## Risks

- Strong ETags hash each successful serialized JSON response. This adds
a small CPU cost but avoids transferring unchanged bodies.
- Debounced bookkeeping timestamps can lag by the bounded debounce
interval. They are non-critical usage metadata; authentication still
runs per request, and stale instance-admin deletion remains
unconditional.
- Legacy pending interactions on terminal issues are projected as
expired by reads and are finalized only by mutation paths. The stored
record remains unchanged on `GET` by design.
- No database schema or migration changes are included.

> This is a focused performance correction and does not duplicate a
planned core feature in `ROADMAP.md`.

## Model Used

OpenAI Codex using `gpt-5.3-codex` for the initial implementation and
`gpt-5.6-sol` for isolation, verification, and PR preparation, with
reasoning, repository tool use, code execution, and GitHub CLI access.
The runtimes did not expose authoritative context-window sizes.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Dev Agent <dev@paperclip.ing>
2026-08-11 14:39:23 -04:00
DottaandPaperclip 5858ccb981 feat: make in-app features cloud-aware (#10850)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators use the same board application in self-hosted and
Paperclip Cloud deployments.
> - A Cloud tenant contains one company, so an in-app company switch
does not change the active Cloud stack.
> - Cloud operators need the sidebar and company surfaces to use the
signed-in user's stack portfolio.
> - The server must derive Cloud identity and links from trusted
instance context instead of client input.
> - This pull request adds canonical Cloud context, a trusted stack
portfolio proxy, and Cloud-aware navigation.
> - The benefit is consistent stack switching on Cloud while self-hosted
company behavior stays unchanged.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: server REST routes and the React board UI.

**Problem or motivation**

A Cloud-managed instance contains one company. The existing company
switcher could only switch records inside that tenant. It could not move
the operator to another Cloud stack. The existing header also gave long
organization names too little width.

**Proposed solution**

Expose a canonical public Cloud context in health data. Add a trusted
server proxy for the current user's stack portfolio. Use that data in
the board UI to switch stacks with top-level navigation. Keep the
existing company behavior on self-hosted instances. Move search into the
navigation and keep long organization names inside the sidebar panel.

**Alternatives considered**

An in-app `/stacks` route was rejected because Cloud tenant hosts
reserve that path and stack selection must wake or authenticate another
tenant. Client-supplied user identity was rejected because the server
can derive the trusted Cloud actor.

**Roadmap alignment**

This change advances the Cloud deployments milestone. It keeps the
product local-first and Cloud-ready without changing the self-hosted
mental model.

## What Changed

- Added canonical Cloud instance context and public health metadata.
- Added a Cloud-only stack portfolio proxy with trusted actor forwarding
and per-user caching.
- Prevented normal company creation on Cloud-managed instances.
- Switched the sidebar and Companies page from company actions to stack
actions on Cloud.
- Added full-page stack navigation and Cloud create-stack links.
- Moved search into the sidebar navigation so the organization name
keeps more width.
- Added truncation and hover recovery for long organization and stack
names.
- Added server and UI regression coverage for Cloud and self-hosted
behavior.
- Updated the implementation specification for the Cloud contracts.

## Verification

- `node scripts/check-token-gates.mjs` passed. All three token gates are
clean.
- `pnpm --dir server exec vitest run src/__tests__/health.test.ts
src/__tests__/cloud-instance.test.ts src/__tests__/cloud-routes.test.ts
src/__tests__/company-cloud-floor.test.ts
src/__tests__/company-portability-routes.test.ts` passed: 5 files and 66
tests.
- `pnpm --dir ui exec vitest run
src/components/SidebarCompanyMenu.test.tsx` passed: 1 file and 11 tests.
- Pre-PR QA report `7da87ca7` passed all 8 acceptance criteria with real
HTTP route factories and real Chromium screenshots in Cloud and
self-hosted modes.
- Security reviews passed for the canonical Cloud context and stack
portfolio proxy.

## Risks

- Cloud stack switching depends on the configured Cloud application and
tenant portfolio URLs.
- The new health `cloud` block is public by design, but it contains only
canonical public instance metadata.
- The stack proxy fails closed on self-hosted instances and derives the
user identity from the trusted actor.
- Self-hosted navigation and company creation retain their existing
paths and behavior.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, model `gpt-5`. The run used reasoning, repository tools,
shell execution, and GitHub integration. The deployment did not expose
its context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-04 23:00:14 -05:00
DottaandPaperclip 2495e29f7f Preserve company display names in cloud tenants (#10845)
Prefer the trusted organization name, repair known machine-generated legacy names with compare-and-set safety, and preserve the audited fallback behavior required by PAP-16331.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-08-04 19:37:08 -05:00
Devin Foley 1c52f02d34 Let cloud tenant sessions reach companies they hold memberships in (#10524)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - On Paperclip Cloud, each stack authenticates its users to the tenant
app through trusted headers (`resolveCloudTenantActor`), which seed a
primary company for the stack
> - That actor was pinned to exactly one company — the seeded primary —
regardless of any other companies the user actually holds a membership
in
> - Companies created later (via the import flow, or company creation)
write real membership rows for the user, but the pinned actor ignored
them, so those companies showed up in listings yet returned "User does
not have access to this company" when opened
> - This pull request unions the pinned primary with the user's own
active membership rows, exactly as a locally authenticated session
already does
> - The benefit is that a Cloud user can reach every company they belong
to — most visibly, a company they just imported

## Linked Issues or Issue Description

- Refs #10507 (Import/Export — imported companies were unreachable on
Cloud stacks). No open issue; bug described above (companies visible in
listing but unreachable; expected: reachable when the user holds an
active membership).

## What Changed

- Extracted the session path's own active-membership query into
`loadActiveUserCompanyMemberships(db, userId)` (single-sourced; the
session path now calls it too).
- `resolveCloudTenantActor` unions its result with the pinned primary:
`companyIds = [primary, ...others]`, memberships likewise, primary
first. Strictly per-user; a membership-read failure degrades to
primary-only (mirrors the existing fail-closed owner-elevation pattern).
No change to owner instance-admin elevation, grant seeding, the stale
instance-admin purge, or trusted-header validation.
- Grants are seeded at membership creation across all flows (company
create, invite/join, import), not per request — so no extra seeding was
added here.

## Verification

- `@paperclipai/server` typecheck clean.
- `cloud-tenant-actor.test.ts` (+ union / other-user-excluded /
inactive-excluded / no-rows-identical cases),
`auth-session-route.test.ts` (route-level: trusted headers reach a
unioned company through `assertCompanyAccess`), plus agent-auth,
authz-company-access, cross-company-authz, portability-routes — 83 tests
green.

## Risks

- Low and tightly scoped: only widens a Cloud actor's reachable
companies to those it already holds active memberships in; users without
extra memberships, other users' rows, and owner elevation are all
unaffected. Read failure fails closed to primary-only.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), Claude Code CLI,
extended thinking + tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-30 16:36:50 -07:00
Devin FoleyandPaperclip c274f10abc feat(server): computed owner instance-admin elevation for cloud-managed instances, behind platform floors (#10343)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Cloud-managed instances authenticate tenant users through a
trusted-header path (`resolveCloudTenantActor`) that deliberately never
grants `instance_admin`, so every tenant user is company-scoped
> - On a dedicated (single-owner) managed instance that leaves the
paying owner unable to administer their own instance: instance settings,
the environments admin surface, and the custom sandbox image flow are
all instance-admin gated (the environments UI can't even show the
provider/image of the platform sandbox because the restricted read view
blanks `config` entirely)
> - Re-granting the old blanket `instance_user_roles` row would repeat
the mistake the shared-pool hardening fixed: DB rows go stale, resurrect
via restores, and elevate through every auth path
> - This pull request elevates only the stack `owner`, computed per
request at the trusted-header boundary behind a new managed-tier feature
flag, and ships that elevation together with code floors on the
platform-owned surfaces an instance admin must not control on a managed
instance
> - The benefit is that dedicated-stack owners can administer their own
instance while platform credentials, execution policy, backups, and
runtime code-install stay platform-owned, and self-hosted behavior is
unchanged

## Linked Issues or Issue Description

No public GitHub issue exists for this change; the underlying issue is
described here following the feature-request template.

### Problem or motivation

- On cloud-managed instances, tenant users resolved from trusted headers
are always company-scoped. For dedicated instances with a single paying
owner, the owner cannot reach any instance-admin surface of their own
instance (instance settings, environments administration, custom image
setup), and the restricted environment read view hides even structural
fields like the sandbox provider and image.
- The previous hardening intentionally removed blanket elevation (and
purges stale `instance_user_roles` rows on every trusted-header
authentication). That protection must not regress for shared
multi-tenant pools.

### Proposed solution

- Owner-only, computed, flag-gated elevation plus code floors on
platform-owned surfaces, in one PR so the elevation can never ship
without the floors.

### Alternatives considered

- Re-inserting an `instance_user_roles` row for owners (the
pre-hardening model): rejected — DB rows go stale, survive restores, and
elevate through every auth path; #7525 removed exactly this.
- Widening only the environments read view without any elevation:
rejected — it fixes one screen but still leaves a dedicated-instance
owner unable to administer instance settings or custom images.
- Elevating additional stack roles (`member`/`admin`/`support`):
rejected — only the owner has an ownership claim over the whole
instance; other roles stay company-scoped.

### Roadmap alignment

- Extends the shipped "Cloud deployments" roadmap work (multi-tenant
isolation, company-scoped cloud tenants, managed-instance bootstrap)
without overlapping planned core items, and leaves self-hosted behavior
unchanged.

## What Changed

- **New feature key** `enableOwnerInstanceAdmin` (`packages/shared`):
boolean flag in `instanceExperimentalSettingsSchema`, catalog tier
`managed`, `cloudDefault: true`, `selfHostedDefault: false`. Inert on
self-hosted instances — the elevation path only exists behind the cloud
tenant trust token.
- **Computed elevation** (`server/src/middleware/auth.ts`):
`resolveCloudTenantActor` now returns `isInstanceAdmin: true` only when
the trusted-header stack role is `owner` **and** the flag is enabled.
The flag is resolved through the instance-settings service so the
managed-config overlay applies (the control plane can disable elevation
fleet-wide without touching tenant databases; a DB row edit or restore
cannot resurrect it). Resolution fails closed on settings read errors.
The `instance_user_roles` never-insert and the per-request stale-row
purge are byte-identical. `member`/`admin`/`support` stack roles stay
company-scoped.
- **Authorization guard split**
(`server/src/services/authorization.ts`): the blanket-allow now trusts
the actor's *computed* `isInstanceAdmin` flag (only the attested
resolver can set it for `cloud_tenant` actors) while keeping the
`instance_user_roles` DB lookup excluded for `cloud_tenant` — a stale or
hand-inserted role row still elevates nothing.
- **Floor F1 — platform environment credentials**
(`server/src/routes/environments.ts`): on cloud-managed instances,
platform-provisioned environment rows (`managedByPaperclip` marker, plus
the legacy managed-Kubernetes marker) use a single floored view for
every reader on all environment routes (list, get, create, update,
delete responses): `envVars` are never echoed and credential-shaped
`config` keys (reusing the managed-config
`SECRET_LIKE_CONFIG_KEY_PATTERN`) are dropped — for **all** actors
including instance admins — while structural config (provider, image,
template, region, …) and the managed markers stay visible. This also
fixes the environments UI for managed sandboxes, which previously lost
the provider/image entirely in the restricted view. The floor also
covers writes: `PATCH /environments/:id` and `DELETE /environments/:id`
on a platform-provisioned row are rejected (403,
`environment_platform_managed`) for every actor including instance
admins, and the guard binds to the persisted row's markers so a patch
cannot strip the managed marker to lift the floor. The one recovery path
is a metadata-only PATCH that solely clears the marker keys
(null/false), for rows stamped through the old unrestricted API before
the markers became reserved — and it never applies to a row whose slot
markers are live platform state: the single local row
(`environments_local_driver_idx`), which `ensureLocalEnvironment` adopts
and stamps on cloud-managed instances from every caller (company
creation, the heartbeat, run orchestration), and the single marked
sandbox row (`environments_managed_sandbox_idx`) while a managed-sandbox
bootstrap path is configured (managed-config `environments` section or
`PAPERCLIP_EXECUTION_MODE=kubernetes`) and the provisioner therefore
adopts and refreshes it on every boot. Clearing a live slot row's
markers would let the next write reclassify it as tenant-managed and
bypass the floor; conversely, when no sandbox provisioning path is
configured the platform holds no claim on any sandbox row, so a platform
marker there is stale by definition and the recovery patch applies.
Every marker outside a live slot is clearable, so no legacy row is ever
locked permanently. Custom-image setup and probes on the platform
sandbox stay available to instance admins — those are the owner-facing
flows this elevation exists for. The marker keys themselves are
reserved: client create/update payloads that set `managedByPaperclip` or
`managedKubernetesSandbox` are rejected (422,
`environment_platform_marker_reserved`) on cloud-managed instances, so a
tenant row can never be stamped platform-provisioned through the API and
self-locked behind the write floor (the provisioner writes markers at
the service layer, not through these routes). Tenant-created
environments are otherwise unaffected.
- **Floor F2 — executionMode**
(`server/src/routes/instance-settings.ts`): on cloud-managed instances,
`PATCH /instance/settings/general` rejects writes that would change
`executionMode` (403, `execution_mode_platform_managed`). Same-value
echoes pass so settings forms that submit the full general-settings
object keep working. The boot-time execution-policy bootstrap path is
untouched (it calls the service directly).
- **Floor F3 — manual database backups**
(`server/src/routes/instance-database-backups.ts`): floored off on
cloud-managed instances (403, `database_backups_platform_managed`);
backups are platform-owned there, and the result would also echo a
server-side filesystem path.
- **Floor F4 — adapter code install** (`server/src/routes/adapters.ts`):
`POST /adapters/install` and `POST /adapters/:type/reinstall` are
floored off on cloud-managed instances (403,
`adapter_install_platform_managed`). Adapter packages execute in the
server process, so a runtime install would let an instance admin read
the platform trust anchors out of the process environment. This mirrors
the existing bundled-only plugin install floor; adapter code on managed
instances comes bundled with the platform image.

## Instance-admin surface audit

Before widening who can hold `isInstanceAdmin`, every
instance-admin-gated surface in `server/src` was enumerated and reviewed
for whether its response or side effects could echo process environment
values or platform credentials (tenant trust token, JWT signing keys,
database connection strings, provider API keys): 29 distinct gate
definitions covering ~90+ call sites, in four groups — sole
instance-admin gates (12), instance-admin-or-company-permission gates
(10), response-shaping/scope-widening sites (6), and the central
`allow_instance_admin` short-circuit in the authorization service (58
`decide()` call sites).

Findings and dispositions:
- **Environment read/write responses** exposed platform sandbox
`envVars`/credential-shaped config to instance admins → closed by floor
F1.
- **Manual backup trigger** echoed a server filesystem path and triggers
a platform-owned operation → closed by floor F3.
- **Adapter install/reinstall** loads externally fetched code into the
server process (indirect, complete env exposure) → closed by floor F4.
The sibling plugin-install path already had a bundled-only floor on
managed instances and needed no change.
- **Token-minting surfaces** (gateway tokens, custom-image
terminal/connection tokens) mint credentials scoped to the instance's
own resources, not platform trust anchors → acceptable for an
owner-admin of a dedicated instance; unchanged.
- All remaining gated surfaces return ordinary instance-scoped business
data; none echo `process.env` or platform secrets directly. OAuth client
secrets are referenced by env-var *name* only; SSH private keys are
stored as secret refs before persistence and are not echoed.

Operational note for managed platforms: this model assumes the process
environment of a managed instance holds only that instance's own
credentials. Platform operators should keep provider credentials
per-instance (never fleet-shared) since an instance admin ultimately
controls in-process code on their own instance.

## Verification

- `pnpm vitest run server/src/middleware/cloud-tenant-actor.test.ts` —
resolver matrix: owner × flag on/off, flag via managed overlay
(on-over-DB-off and off-over-DB-on), member/admin/support × flag on,
no-token self-hosted, fail-closed settings read, purge still runs and no
role row is ever inserted (14 tests).
- `pnpm vitest run server/src/__tests__/authorization-service.test.ts` —
computed flag elevates a `cloud_tenant` actor; a stale
`instance_user_roles` row still never does; `session` actors unchanged
(full suite, embedded Postgres).
- `pnpm vitest run server/src/__tests__/environment-routes.test.ts` —
F1: no secret echo to admins on get/list, structural config visible to
restricted readers, platform-row PATCH/DELETE rejected for admins
(including a marker-stripping patch), marker-clear recovery allowed for
stale legacy rows and for a marked sandbox row when no provisioning path
is configured, but refused on the managed local row and on the sandbox
slot row under a managed-config `environments` entry or the forced
kubernetes execution mode, client marker-stamping creates/patches
rejected, tenant rows still readable and writable, self-hosted
read+write regression (60 tests).
- `pnpm vitest run server/src/__tests__/environment-service.test.ts` —
`ensureLocalEnvironment` adopts a pre-existing local row on
cloud-managed instances (marker stamped, other metadata preserved,
idempotent — no rewrite on re-ensure) and leaves self-hosted rows
untouched (22 tests, embedded Postgres).
- `pnpm vitest run server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/instance-database-backups-routes.test.ts` — F2
change-vs-echo matrix incl. self-hosted regression; F3 floor for both
admin shapes (32 tests).
- `pnpm vitest run server/src/__tests__/adapter-routes-authz.test.ts` —
F4 floor; self-hosted install/reinstall behavior unchanged (existing
cases).
- `pnpm vitest run server/src/__tests__/first-admin-claim.test.ts
server/src/__tests__/bootstrap-claim-routes.test.ts
server/src/__tests__/managed-config.test.ts
server/src/__tests__/health.test.ts
server/src/__tests__/instance-settings-managed-overlay.test.ts
server/src/services/managed-environments.test.ts
server/src/services/execution-policy-bootstrap.test.ts` — first-admin
bootstrap gate and managed-config behavior unchanged (91 tests).
- `pnpm vitest run packages/shared/src/feature-catalog.test.ts` —
catalog/schema sync tests cover the new key (selfHostedDefault must
equal the schema default).
- `pnpm run typecheck` — all 31 workspace projects clean.

## Risks

- Self-hosted behavior is unchanged: every floor binds to
`isCloudManagedInstance()` (tenant trust token present), the new flag
defaults off with no elevation path, and regression tests pin the
self-hosted branches.
- The elevation is fail-closed and stateless: turning the flag off
(managed overlay or DB) de-elevates on the next request; there is no
role row to clean up and restores cannot resurrect elevation.
- On a cloud-managed instance a pre-existing unmarked local row is
adopted (stamped `managedByPaperclip`) by the next ensure and becomes
platform-owned — the intended managed-product semantic: the platform
owns the single local slot. Self-hosted instances are untouched.
- F1 widens restricted readers' view of platform-provisioned rows from
fully blanked `config`/`metadata` to structural-only `config` plus
markers. Platform-delivered config is guaranteed secret-free by the
managed-config contract (secret-shaped keys fail startup), and the floor
re-drops secret-shaped keys defensively.
- One extra instance-settings read per trusted-header request for
owner-role actors (the resolver already performs several queries per
request).

## Model Used

Claude Fable 5 (Anthropic) — model id `claude-fable-5`, extended
thinking enabled, agentic tool use via Claude Code; read-only explore
subagents on the same model were used for the surface audit sweep.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-27 18:58:31 -07:00
DottaandPaperclip cac3c0fa1a feat(connections): add runtime subjects and grants (#9982)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Connections is the subsystem that lets operators connect external
apps and govern which subjects may use those credentials
> - #9958 established the v3 schema foundation and #9981 adds the
AppDefinition catalog layer
> - The runtime still needs subject-aware authorization state, scoped
key handling, and API/OpenAPI routes so connected apps can actually be
granted and used safely
> - This pull request adds the runtime grants/authorization behavior on
top of the catalog branch, while keeping unrelated dependency and
workflow sync commits out of the stack
> - The benefit is a reviewable runtime layer that can land after the
catalog PR, then unblock the wizard and orchestrator cutover work

## Linked Issues or Issue Description

Refs #9958 and #9981.
Refs #9981.

No public GitHub issue exists for this branch. This is the runtime layer
for the Connections v3 stack and is rebased onto `master` after #9981
landed.

## What Changed

- Adds the connection user authorization state migration and schema
wiring.
- Adds shared runtime subject/grant types and validators.
- Adds runtime grant and scoped key behavior in the tool-access service.
- Adds runtime route coverage and registers the routes in OpenAPI.
- Replays only the Connections runtime commits on top of the catalog
branch, dropping unrelated sync/dependency history from the prior closed
runtime PR.

## Verification

- `pnpm run preflight:workspace-links`
- `pnpm exec vitest run
packages/shared/src/validators/tool-access.test.ts
server/src/__tests__/tool-access-service.test.ts`

## Risks

- Medium: runtime grant enforcement is security-sensitive and must fail
closed for unknown key scopes.
- Migration ordering depends on the schema and catalog layers already
merged through #9958 and #9981.
- This PR is rebased and retargeted to `master` with runtime-only
commits.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-5 Codex coding agent with repository tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-21 16:17:02 -05:00
NoSoloSoftandAndrew Aymeloglu 156830006b fix: redact HTTP cookies from server logs (#7977)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators rely on Paperclip server logs for maintenance, incident
triage, and support handoffs.
> - The HTTP logger persisted request metadata and only redacted
authorization headers.
> - Request cookies and set-cookie headers can contain active session
material and should not be written to durable logs.
> - This pull request keeps the fix intentionally narrow: centralize the
HTTP log redaction path list and include cookie-bearing headers.
> - The benefit is lower credential/session leakage risk from routine
server.log collection or sharing.

## Linked Issues or Issue Description

No GitHub issue exists for this exact local finding. Inline bug report:

- Type: security/privacy bug.
- Affected area: server HTTP logging middleware.
- Observed problem: local Paperclip maintenance found raw cookies
present in server.log.
- Expected behavior: durable HTTP logs redact authorization and
cookie-bearing request/response headers.
- Impact: anyone with access to copied/exported logs could see
session-bearing cookie values.
- Related/open PRs found during dedup search: #7242, #7306, #7346. This
PR is the minimal local fix branch created from the verified local
maintenance patch; those PRs may be better upstream candidates if
maintainers prefer their broader coverage.

## What Changed

- Added `HTTP_LOG_REDACT_PATHS` for HTTP logger redaction paths.
- Kept existing `req.headers.authorization` redaction.
- Added redaction for `req.headers.cookie`, request `set-cookie`, and
response `set-cookie` paths.
- Added focused tests asserting the required redaction paths are present
and that pino-http output redacts live request/response header secrets.

## Verification

- `pnpm exec vitest run server/src/__tests__/http-log-redaction.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- Pre-commit TruffleHog scan: 0 verified/unverified secrets.
- PR CI observed passing so far for policy, Typecheck + Release
Registry, Build, e2e, Socket, Snyk, security-review, and
serialized/workspace suites; remaining jobs may still be running.

## Risks

- Low runtime risk: this only expands pino redaction paths.
- Possible coverage risk: broader redaction helpers in related PRs may
cover more serialized variants beyond the pino-http request/response
header pipeline tested here.
- No migrations, schema changes, or UI changes.

## Model Used

- OpenAI Codex via Hermes Agent, model gpt-5.5, tool-using coding/ops
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Andrew Aymeloglu <aaymeloglu@gmail.com>
2026-07-20 17:17:08 -05:00
Devin Foley 2f42a4968d Treat cloud-managed instances as bootstrapped in the health gate (#9912)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip instances can be self-hosted, or provisioned and managed
by a cloud control plane that authenticates users through trusted
headers validated against `PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN`
(`resolveCloudTenantActor`)
> - In `authenticated` deployment mode, the health route reports
`bootstrapStatus: bootstrap_pending` until at least one `instance_admin`
exists, and the UI locks everyone out at the "waiting on its first
admin" claim screen until then — correct for self-hosted instances,
where a human operator must claim the instance
> - But the cloud-tenant trust middleware, by deliberate security
hardening, never grants `instance_admin` and actively purges legacy
grants — so a cloud-managed instance can never leave
`bootstrap_pending`: the gate demands a role the middleware forbids
> - Every control-plane-provisioned instance is therefore permanently
locked at the claim screen even though its users and memberships exist
> - This pull request makes the gate cloud-aware: when the tenant server
token is configured, the instance is considered bootstrapped, because
the control plane owns identity and there is no operator claim step
> - The benefit is that cloud-managed instances become usable while
self-hosted behavior stays byte-for-byte identical, now pinned by a
previously missing regression test

## Linked Issues or Issue Description

Refs #2927 (introduced the browser-native first-admin bootstrap flow
this gate feeds). No existing public issue for the deadlock; inline
description per the bug report template:

- **What happened?**: an instance configured with
`PAPERCLIP_DEPLOYMENT_MODE=authenticated` and
`PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` reports `bootstrapStatus:
bootstrap_pending` forever. All users — including ones created via the
trusted-header path with owner-level company membership — are locked out
at the "This Paperclip is waiting on its first admin" screen.
- **Expected behavior**: a control-plane-managed instance has no
first-admin claim step; users arriving with control-plane identity
should reach the app.
- **Steps to reproduce**:
1. Run the server with `PAPERCLIP_DEPLOYMENT_MODE=authenticated` and a
`PAPERCLIP_CLOUD_TENANT_SERVER_TOKEN` set
2. Create users only through trusted cloud headers (the middleware
upserts them but never grants `instance_admin`, and purges any legacy
grants)
3. `GET /api/health` → `bootstrapStatus` stays `bootstrap_pending`; the
UI shows the claim screen for every visitor, and no supported path
exists to create the `instance_admin` the gate requires
- **Paperclip version or commit**: reproducible on `master` as of
2026-07-20; present since the cloud-tenant `instance_admin` purge
hardening landed.

## What Changed

- `server/src/middleware/auth.ts`: new exported
`isCloudManagedInstance()` predicate beside the trust middleware that
defines the tenant-token contract.
- `server/src/routes/health.ts`: the authenticated-mode first-admin gate
is skipped when the instance is cloud-managed; `bootstrapStatus` reports
`ready`.
- `server/src/__tests__/health.test.ts`: two new tests — authenticated
without the token → `bootstrap_pending` (previously untested regression
baseline), and with the token → `ready` despite zero instance admins.

## Verification

- `pnpm vitest run src/__tests__/health.test.ts` in `server/` — 13/13
- `pnpm vitest run src/middleware/cloud-tenant-actor.test.ts` — 6/6
- Manual: with the env vars from the repro steps set, `GET /api/health`
now returns `bootstrapStatus: "ready"`; without the token, behavior is
unchanged

## Risks

- None for self-hosted deployments: without the env var the gate is the
prior behavior, now pinned by the new regression test.
- For cloud-managed instances the claim screen and
`bootstrapInviteActive` flow no longer appear — intended; browser-based
claim was already disabled in that configuration.

## Model Used

- Claude (Anthropic) — model id `claude-fable-5`, via the Claude Code
CLI harness with tool use (shell, file edits, test execution). Diagnosis
and change agent-assisted, human-directed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (n/a —
behavior documented in code comments and pinned by tests)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-07-20 15:59:08 -04:00
89ce36d7af feat(skills): open-by-default company skill policy and core UX (#9564)
## Thinking Path

> - Paperclip uses company skills to make agent capabilities reusable
across an organization.
> - Skill operations currently mix capability availability with
permission checks, which creates avoidable setup friction and
inconsistent denial handling.
> - The policy contract needs to remain open by default while allowing
company-scoped restrictions for governed deployments.
> - Core owns the canonical policy actions, persistence, evaluation, API
behavior, safe import boundaries, and generic denial/read-only UI.
> - Enterprise policy-editor implementation belongs in the separate
`paperclip-ee` repository and is intentionally excluded from this PR.

### Problem or motivation

Company skill operations can encounter permission dead ends even when no
explicit restriction has been configured, and import-source
classification can drift between policy evaluation and execution.

### Proposed solution

Define eight canonical skill policy actions, default all actions to
allowed, persist company-scoped restrictions, expose policy evaluation
APIs, normalize import sources at the boundary, and update Skill Studio
to present actionable restriction states without embedding Enterprise
Edition implementation in the core repository.

### Alternatives considered

Keeping capability checks distributed across routes and UI surfaces was
rejected because it duplicates policy logic and makes denial behavior
inconsistent. Shipping the Enterprise policy editor in this repository
was rejected because `paperclip-ee` is a separate repository and must
receive its own PR.

### Roadmap alignment

Extends the completed **Skills Manager** roadmap area by adding coherent
governance and removing workflow dead ends.

### Additional context

The core API contract remains suitable for a separate Enterprise Edition
editor, but this PR contains no `paperclip-ee` package or EE-specific UI
integration code.

## What Changed

- Added the company skill policy contract to product and implementation
documentation, including the open-by-default rule, eight canonical
actions, decision shape, and core/EE ownership boundary.
- Added the company-scoped policy schema, migration `0170`, shared
validators, policy service, REST routes, OpenAPI coverage, and focused
server tests.
- Hardened import policy enforcement by normalizing import sources and
keeping source classification consistent between policy evaluation and
execution.
- Updated core Skill Studio behavior to remove generic permission dead
ends and show actionable policy/platform denial states only when an
operation is actually denied.
- Removed the `plugin-paperclip-ee` package, Docker wiring, EE
discovery/deep-link helpers, and EE-specific UI tests/stories from this
PR so that implementation can be submitted separately to the EE
repository.
- Preserved open-by-default behavior when no explicit company
restriction exists.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/skill-studio/SkillPolicySurfaces.test.tsx
src/lib/skill-policy-denial.test.ts` — 20/20 passed.
- `pnpm --filter @paperclipai/ui exec tsc --noEmit` — passed.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/worktree-config.test.ts` — 12/12 passed.
- `pnpm check:token-gates` — passed with all gates clean.
- `git diff --check` — passed.
- `git diff --name-only origin/master | rg
'paperclip-ee|ee-skill-policy'` — no matches.

## Risks

- Migration `0170` introduces company policy persistence; rollout
depends on the migration applying before policy routes are exercised.
- Open-by-default is an intentional behavioral policy: deployments
expecting implicit denials must configure explicit restrictions.
- Import normalization is security-sensitive and should retain focused
review.
- The separate EE editor must stay contract-compatible with the core
policy API as policy actions evolve.

## Model Used

- OpenAI Codex CLI, runtime model identifier and context-window size not
exposed by this execution environment; reasoning, repository tool use,
shell execution, and code review capabilities enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available to this runtime)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked or
described the result above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green on the latest head
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Evyatar Bluzer <bluzername@users.noreply.github.com>
2026-07-15 11:42:40 -05:00
DottaandPaperclip cfa5e0704e feat(mcp) [split 3/8]: add tool access policy core (#9558)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Governed MCP access spans contracts, runtime enforcement, adapters,
UI surfaces, and operator verification
> - The parity reference PR #9534 is too large for effective automated
or human review
> - The feature therefore needs a linear stack whose individual diffs
stay below the 100-file review limit
> - This pull request is split 3/8 and focuses on tool-access policy and
authorization core
> - The benefit is a standalone, testable review boundary while
preserving byte-for-byte parity at the top of the stack

## Linked Issues or Issue Description

- Related parity reference: #9534
- Problem: Authorization, OAuth binding, secret projection, content
guards, and policy evaluation need a security-reviewable server
boundary.
- Proposed solution: Adds tool-access services/routes/tests plus the
runtime service dependencies directly imported by the core, without
registering the routes in the application.
- Alternatives considered: keeping #9534 as one 403-file review, or
rewriting the feature to manufacture seams; both were rejected in favor
of path extraction plus compile-driven boundary moves.
- Roadmap alignment: this advances the existing governed MCP/tool-access
work already represented by #9534; it does not introduce a separate
roadmap initiative.
- Stack position: base branch is `pap10341-split/02-schema-shared`.
- Merge policy: merge bottom-up, in order, only after the complete
eight-PR stack has been reviewed and the top-of-stack parity gate
remains empty.
- Requested review: SecurityEngineer for authz, OAuth, secrets, and
content guards; Greptile on every PR.

## What Changed

- Adds tool-access services/routes/tests plus the runtime service
dependencies directly imported by the core, without registering the
routes in the application.
- Keeps this PR below 100 changed files and independently typecheckable.
- Preserves the final tree from #9534 when combined with the other seven
stack levels.

## Verification

- `pnpm typecheck`
- Focused server Vitest run — 4 files, 143 tests passed

## Risks

- Authorization bugs could permit cross-company or over-broad tool
access; the PR remains inert until PR 4 wiring and requires dedicated
security review.
- Stack risk: merging out of order can expose incomplete layers;
mitigate by following the documented bottom-up merge policy.
- Parity risk: later edits to an intermediate branch can drift from
#9534; mitigate by re-running the empty top-of-stack diff before merge.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context
window; medium reasoning with repository, shell, Git, GitHub CLI, and
code-execution tools enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] Internal references are omitted except the execution-plan link
explicitly required for this coordinated split stack
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Stack Coordination

- Internal execution plan:
[PAP-13874](/PAP/issues/PAP-13874#document-plan)
- Parity reference: #9534
- Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563
- Merge bottom-up only after full-stack review and an empty parity diff
at #9563.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-14 14:22:39 -05:00
DottaandClaude Fable 5 23f34491e2 Fix apiCompression corrupting and dropping Better Auth responses for gzip clients (#9381)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its server fronts every API route — including Better Auth sign-in —
with Express middleware, and #9190 added an `apiCompression` middleware
that gzips JSON responses over 1KB
> - That middleware buffers `res.write()` chunks with `String(chunk)`,
but Better Auth (via better-call) streams `Uint8Array` chunks and
commits headers with `writeHead()` before streaming
> - `String(Uint8Array)` serializes the body to comma-separated decimal
bytes (~3.4x inflation), and once the inflated body crossed the 1KB
threshold, `setHeader()` threw `ERR_HTTP_HEADERS_SENT` and the catch
handler destroyed the socket
> - Every real browser sends `Accept-Encoding: gzip`, so sign-in
returned zero bytes (`net::ERR_EMPTY_RESPONSE` / "Failed to fetch"),
while curl without `Accept-Encoding` worked — making the bug easy to
misdiagnose as a client or network issue
> - This pull request makes the middleware byte-safe for `Uint8Array`
chunks, passes through responses whose headers are already committed,
and falls back to the uncompressed body instead of destroying the
connection when compression fails
> - The benefit is that browser sign-in (and any other streamed
binary-chunk response) works again for gzip-accepting clients, with
regression tests locking in all three behaviors

## Linked Issues or Issue Description

Refs #9190 (introduced the `apiCompression` middleware).

No public GitHub issue exists; bug description:

- **What happened:** Sign-in from any real browser failed with
`net::ERR_EMPTY_RESPONSE` / "Failed to fetch". The server logged
`ERR_HTTP_HEADERS_SENT` from the compression middleware and destroyed
the response socket, so zero bytes reached the client.
- **Expected:** `/api/auth/*` responses are delivered intact regardless
of the client's `Accept-Encoding`.
- **Steps to reproduce:** Run the server with API compression active,
open the web UI in a browser (which sends `Accept-Encoding: gzip`), and
attempt email/password sign-in. The auth response body exceeds ~300
bytes, so after the ~3.4x stringification inflation it crosses the
1024-byte compression threshold and the response is destroyed. `curl`
without `Accept-Encoding` succeeds against the same server.
- **Scope:** Any route that streams `Uint8Array` chunks and/or commits
headers via `writeHead()` before writing — in practice all Better Auth
routes served through better-call.

## What Changed

- `server/src/middleware/api-compression.ts`:
- Buffer `res.write()` chunks with a `toBodyBuffer()` helper that
converts `Uint8Array`/`ArrayBuffer` views via `Buffer.from()` instead of
`String()`, so binary chunks are preserved byte-for-byte.
- Pass responses through untouched once headers are already sent
(`writeHead()`-style streaming), since compression headers can no longer
be set at that point.
- On any compression failure, write the original uncompressed body
instead of calling `res.destroy()`, so clients get a valid (just
uncompressed) response rather than a dropped connection.
- `server/src/__tests__/api-compression.test.ts`: three new regression
tests — small `writeHead`+`Uint8Array` responses are delivered
byte-for-byte, large ones no longer drop the connection, and
`Uint8Array` JSON bodies gzip without corruption (includes
`/api/auth-bridge` and `/api/uint8-json` test routes mirroring
better-call's streaming pattern).

## Verification

- `cd server && pnpm vitest run src/__tests__/api-compression.test.ts` —
10/10 passing (7 pre-existing + 3 new regression tests).
- Manual: with the fix, browser sign-in against a dev instance succeeds
for gzip-accepting clients; before the fix the same request returned
`net::ERR_EMPTY_RESPONSE`.

## Risks

- Low risk. The middleware still compresses large text/JSON responses
exactly as before; the changes only affect paths that previously
produced corrupted or destroyed responses.
- Behavioral shift: responses whose headers were already committed are
now delivered uncompressed instead of being (incorrectly) buffered —
this is strictly less surprising than the previous corrupted output.
- Failure-path shift: a compression error now yields an uncompressed 200
response instead of a dropped connection.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking
enabled, running via Claude Code / Paperclip agent harness with tool use
(shell, file edit, test execution).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 16:04:29 -05:00