Commit Graph
98 Commits
Author SHA1 Message Date
DottaandPaperclip 71af2fbc3b fix(slack): teach agents how people connect their accounts (#15576)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack connections let people talk to an agent with their own
Paperclip permissions.
> - Each bot has a saved command that starts account linking.
> - The Access page explains this flow, but Slack agent turn context
omitted it.
> - An agent could guess the command or confuse channel membership with
Paperclip access.
> - This pull request gives agents the saved command and account
confirmation instructions.
> - People can ask the agent how to join without changing access
controls.

## Linked Issues or Issue Description

**What happened?**

Slack agent context explained messages, tools, and questions. It did not
explain how a teammate can connect an account. The saved slash command
can differ from the current agent name.

**Expected behavior**

Give the saved bot command, such as `/research_ops connect`. The
teammate runs it themselves. They sign in through a private confirmation
link. A new company member must receive admin approval before account
confirmation. The connection manager can copy instructions from Access →
Invite people.

**Steps to reproduce**

1. Configure a Slack bot with a custom slash command.
2. Rename its assigned agent.
3. Inspect a fresh or resumed Slack task prompt. Before this change, it
contains no account invitation instructions or saved connect command.

**Paperclip version or commit**

Base: `d6df12cef69fcaf2d2fe66a393168931d5b8b4e7`.

**Deployment mode**

Applies to local and hosted Slack chat connections. Deterministic tests
used a local isolated database. The new model probe has not run against
a live Slack bot.

Related work: Refs #15413 and #13638. This fixes agent guidance for
their existing account-linking flow. It adds no new membership system or
invitation endpoint.

## What Changed

- Read only the saved public slash command from the company-scoped
conversation endpoint. Supply it only to the assigned agent for Slack
turns with an active or verifying connection.
- Validate the command with the shared Slack configuration schema. Refer
to Access → Invite people when it is missing or invalid. Never guess
from the current agent name.
- Explain personal account confirmation, link expiry, and company
membership approval on fresh and resumed turns. Distinguish channel
invitations from Paperclip access.
- Add deterministic guidance regressions, a manual invitation model
probe, and setup documentation.

## Verification

- Passed: 63 tests in `heartbeat-context-summary.test.ts` and
`heartbeat-chat-task-link.test.ts`.
- Passed: four real-heartbeat regressions in
`heartbeat-slack-invitation.test.ts`. They check the saved command after
an agent rename, a persisted resumed session, full and compact prompts,
missing-command fallback, inactive endpoints, and a different assigned
agent. They also reject caller-supplied command fields and exclude other
setup metadata.
- Passed: the isolated `chat-channels.integration.test.ts` case
`discovers a Slack connect identity without starting work or granting
access`. It covers the private link, duplicate connect requests, and
nonmember access requests without a membership grant.
- Passed: `pnpm --filter @paperclipai/server typecheck` and `pnpm
--filter @paperclipai/server build`.
- Passed: `pnpm test:slack-connector --list` and `git diff --check`.
- Passed: repository `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Typecheck required local IPC access for the
migration check.
- Passed: final-commit CI, with 54 successful checks and two optional
Storybook checks skipped. CI includes all server, chat, workspace,
serialized, Runner, and browser test shards, typecheck, build, and
canary dry run.
- The full local `pnpm test:run` was started. It was stopped after full
CI passed; it had no final local summary. The focused invitation checks
passed locally. Do not count the interrupted local run as a full-suite
pass.
- Greptile gave the exact final commit
`f37e52bde621ef7f5b2bb345074d53345f8e4ee9` a 5/5 score. Both review
findings were fixed and their threads resolved.
- The new `invite-person` model probe is manual. These deterministic
results do not establish a live model or Slack acceptance pass.

## Risks

- Model guidance cannot prove that a person joined. The existing
account-linking and membership checks remain authoritative.
- Legacy rows without a saved command use the Access page fallback.
Invalid command text is excluded from the prompt.
- The query selects only the public command. It does not expose
registration secrets, tokens, or personal confirmation links.
- No schema changes, new provider requests, permission grants, or
telemetry changes.

## Model Used

- OpenAI GPT-6 through Codex. The exact backend variant and context
window are not exposed in this session. Used reasoning, repository
inspection, code editing, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 10:28:14 -05:00
DottaandPaperclip f47614046d fix(slack): upload agent avatars directly during Cloud setup (#15566)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack setup creates a dedicated bot for an agent.
> - The setup should upload that agent's avatar.
> - The current upload asks Slack to fetch an image from the board
origin.
> - Cloud requires a tenant session at that origin, so Slack receives
HTTP 401.
> - This change renders the PNG on the server and uploads the file
directly.

## Linked Issues or Issue Description

Refs #15413.

**What happened?**

Automatic Slack setup left the app and bot with default icons on Cloud
staging. An unauthenticated request to the exact avatar URL returned
HTTP 401 with `tenant_session_required`. Local setup did not expose this
Cloud ingress requirement.

**Expected behavior**

New Slack bots receive the assigned agent's 512-pixel avatar with the
Paperclip dark background.

**Steps to reproduce**

1. Create a Slack app through automatic setup on a Cloud tenant.
2. Complete installation.
3. Inspect the bot avatar in Slack. The previous URL-based upload cannot
fetch the image without a tenant session.

## What Changed

- Render the assigned agent's preset PNG with the existing bounded
worker pool.
- Send PNG bytes as multipart `file` data to `apps.icon.set` instead of
passing a board URL.
- Keep the temporary token in the Authorization header. Let fetch set
the multipart boundary.
- Recheck management permission and credential-lease ownership after
rendering. Close the worker pool during chat service shutdown.
- Cover actual PNG dimensions, uploaded bytes, failure recovery, and
secret-safe responses. Update deployment documentation.

## Verification

- Passed: 84 focused tests across automatic Slack registration and
on-demand agent avatars.
- Passed: full repository build.
- Passed: full repository typecheck. All 54 current-head GitHub checks
passed, including the complete test matrix, all browser shards, build,
typecheck, canary, and security checks. Greptile completed on
`79e79f6fc` with 5/5 and no actionable findings or open review threads.
- The Cloud fetch failure was reproduced without browser credentials. No
Cloud access rule was changed.
- A real Slack upload with this new path still requires deployment and a
fresh automatic setup. Existing apps retain the manual avatar-upload
fallback.

## Risks

- The renderer can time out or Slack can reject the upload. Both
failures preserve the saved app and leave installation usable.
- The renderer adds a bounded, lazy worker pool to Slack registration.
Shutdown closes it.
- No migration, bot permissions, credential retention, or Cloud
authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser inspection. The runtime does not expose a more
specific authoring model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:12:18 -05:00
DottaandPaperclip d66acb7ac1 feat: automate Slack bot app setup and installation (#15413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors give each agent a customer-owned bot and task-backed
conversations.
> - Manual Slack setup requires app creation and copying durable
credentials.
> - Operators need a shorter setup that an assisting agent can use
safely.
> - This pull request creates the app through Slack's Manifest API and
installs it through OAuth.
> - Durable registration state supports recovery without creating
another app.
> - A four-screen wizard, automatic avatar upload, and OAuth account
linking reduce setup work.
> - Connector settings and per-turn tool guidance support daily use
after installation.

## Linked Issues or Issue Description

**Subsystem affected**

Native Slack bot setup, company secret storage, chat connector
management, and agent tool guidance.

**Problem or motivation**

New Slack bots require manual app creation and copying a signing secret
and bot token. Interrupted setup can create duplicate apps. The setup
and management screens contain unnecessary controls. Agents also need
guidance for native questions, files, thread replies, and governed Slack
actions.

**Proposed solution**

Use a temporary app-configuration access token to create a
customer-owned app. Save durable secrets in the vault. Bind OAuth to the
initiating actor, company, endpoint, registration revision, scopes, and
configured origins. Preserve manual and existing-app recovery. Link the
installing user's account, send a welcome DM, and advance from saved
server evidence. Keep request-URL recovery instructions available if
automatic connection detection waits.

**Alternatives considered**

The Slack CLI adds installation requirements. Socket Mode changes
transport. A shared Paperclip-owned app changes app ownership. These
alternatives are outside this change.

**Roadmap alignment**

This extends existing chat connectors and secrets capabilities. Related
public work: #14037 and #13954 cover Slack MCP prerequisites and user
OAuth. No duplicate bot-registration PR was found.

## What Changed

- Share one reviewed manifest builder between automatic registration and
manual setup.
- Add replay-safe migration 0318 and company-bound registration state
with vault references and uncertain-creation recovery.
- Add registration, installation, callback, and resume APIs with
short-lived, single-use OAuth state.
- Save installation credentials before downstream checks and preserve
bot identity constraints.
- Reduce automatic setup to four screens. Keep advanced app details,
manual recovery, and existing-app setup.
- Upload the agent avatar with the Paperclip dark background. Link the
OAuth installer's account and send setup DMs.
- Show agent and connector-owner avatars. Simplify settings, access, and
conversation screens.
- Discover joined Slack channels and enable them by default. Start a
task from a bare mention and admit same-thread follow-ups.
- Refresh Slack tool guidance each turn. Add native-form, file,
approval, and delivery regressions plus manual model probe definitions
and sanitized acceptance records.
- Update deployment/database docs, OpenAPI, redaction, removal cleanup,
production Storybook stories, and provider browser tests.
- Merge current master and move the registration migration after its
latest migration without rewriting published commits.

The completed Slack success view intentionally has a single centered
**Done** action and no **Save & exit**, as explicitly requested by the
product owner. `DESIGN.md` records this exception; unfinished setup
steps retain the aligned wizard footer.

## Verification

- Passed after the master merge: repository typecheck, full build,
Storybook build, design-token gates, module-boundary gates, and
migration generation.
- Passed: all 352 focused Slack deterministic tests and all 14 affected
provider browser tests. Browser tests use controlled provider fixtures
and a separate throwaway instance.
- Passed on current head `c5d01e0e2`: the complete GitHub test matrix
(general server, chat, all workspaces, serialized server, and Runner),
all eight browser shards, typecheck/release registry, build, canary dry
run, security checks, and policy gates. There are 52 passing checks and
no pending or failing checks.
- Greptile completed on the exact current head with 5/5 and no
actionable findings or open review threads.
- Local repair verification passed 93 focused tests, including same-app
reinstall after revocation and rejection of consent started before
revocation, the AgentMail browser journey, and repository typecheck.
Local build and Storybook build also passed. The redundant local
full-suite rerun was stopped after the complete current-head CI matrix
passed.
- Real Slack setup and agent replies were exercised in the authorized
isolated test drive during the setup iteration.
- The ten additional model probes were attempted with legacy
`codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were
partly verified. Native runtime is not qualified. See
`server/src/services/connectors/slack/evals/2026-10-08-acceptance.md`
for evidence and limits.
- Passing model probes cover native forms, downloaded file bytes, bare
mentions with thread replies, explicit posts/reactions, and saved
approval denial.
- The controlled uncertain-write probe found wrong delivery-check IDs.
The canvas fallback attempt used an invented tool name. Search
pagination/native search, a private-source denied-tool receipt, and
distinct board/webhook origins remain unqualified.

Reviewer path: enable Chat connectors, start Slack chat setup, select an
agent, enter an app-configuration access token, and approve Slack
installation. Send a message to the bot and confirm that setup advances
to success. Inspect settings and allowed channels. See
`doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery.

## Risks

- Slack app creation has no provider idempotency guarantee. A timeout
after dispatch stays uncertain until the operator checks Slack.
- OAuth needs a stable public HTTPS board origin. Webhook ingress may
use a separate configured HTTPS origin. Workspace policy can delay
installation.
- Migration 0318 can replay safely on instances that applied the earlier
development migration.
- OAuth installation now links the installer to the initiating Paperclip
user. Identity checks and company access rules still apply.
- Joined channels now enable bot responses by default. Linked-user
authorization and per-action approval rules still apply.
- Model behavior has the documented delivery-check and canvas fallback
failures. A passing CI run does not establish that every model probe
passed.
- Removing the connection does not delete the customer's Slack app. No
new first-party telemetry is added.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose a more
specific authoring model ID or context-window size. The live bot probes
used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:33:31 -05:00
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
Devin FoleyandPaperclip 892b0b3606 fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:20:37 -07:00
f77fcbf4bf feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Research agents need current web information.
> - The Apps catalog connects agents to remote MCP tools through the
normal access rules.
> - Telem.AI supplies web search and page reading through one API key.
> - This PR adds its catalog entry, optional search settings, artwork,
and setup guide.
> - The connection keeps an operator's saved header policy when they
reconnect.

## Linked Issues or Issue Description

Refs #15302. Related catalog work: #13881.

This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the
connector and its review fixes. All seven original commits are
preserved. GitHub denied the attempt to push to the contributor's fork,
so this branch retains the repair commit and merges current master.
Master now includes the same test fix.

For a squash merge, keep the original author in the final commit
message:

```text
Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
```

A search found no separate public Telem issue or competing Telem PR. The
existing Apps path matches the roadmap.

**Agent or provider**

Telem.AI provides web search and page reading through a hosted MCP
server.

**Why this adapter is useful**

Agents can use multiple search providers through one governed
connection. Operators can set the search tier, auto routing, and
provider lists.

**How the agent is invoked**

The remote MCP server uses Streamable HTTP at
`https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer`
header. See the [official MCP
guide](https://docs.telem.ai/integrations/mcp/).

## What Changed

- Add the Telem.AI definition, research entry, permission review, and
generated registry entry.
- Add four optional settings. Unset settings send no request header.
- Add official light and dark artwork, source records, and a setup
guide.
- Forward company, issue, agent, run, project, and correlation IDs by
default. Preserve a saved policy, including disabled forwarding, on
reconnect.
- Add catalog and connection tests.

## Verification

Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`.
Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`.

- Resolved five shared catalog conflicts after the Superagent connection
merged.
- Keep both providers in the research ledger, generated registry,
generator, branding manifest, and connection guide index.
- Correct the combined catalog totals: 52 self-serve candidates, 55
research entries, and 68 Apps entries.
- The published Git tree exactly matches the tested local resolution.
- Catalog and Apps UI suites: **295 tests pass** after the catalog count
fixes.
- Connection service suite: **387 tests pass** in the full run. Its only
failure was the old catalog count. That test passes on a focused rerun
after the fix. This gives **388 passing service tests** across the two
runs.
- Total focused coverage: **683 passing tests**. The first runs exposed
four fixed-count assertions that needed the combined totals.
- Shared package build and plugin SDK compile pass. Token gates and
whitespace checks pass.
- Generation with `--definitions-only` reproduces the Telem definition
and registry. The unrelated AgentMail and Linear drift remains excluded.
- Local UI typecheck ended with exit 137 at the container memory limit.
Full local typecheck, test, and build are not claimed. Earlier runs also
recorded missing Cargo and Node development headers.
- GitHub reports a clean merge state against master `228f0e280`.
- All 54 checks are complete: **52 passed and two Storybook checks
skipped**. No check failed or remains pending.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406)
passes on this head. This includes typecheck, build, tests, browser
shards, Runner checks, and Canary Dry Run.
-
[Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537)
is **5/5 on this head**. There are no review threads, open P2s,
recommendations, or follow-ups.
-
[Superagent](https://github.com/paperclipai/paperclip/runs/112548142930)
passes.
- Final recovery checks confirm all seven original commits and current
master remain in history. Token gates and whitespace checks pass.
- The final recovery run makes no source change. It verifies the
published repair and retains the local check limits below.
-
[Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943)
passes with no failures. Its only informational note asks the merger to
keep the author trailer above.
- All seven original contribution commits remain in history. The diff
against master contains the same 14 Telem files. It adds no dependency,
lockfile, schema, or workflow change.
- The managed GitHub CLI capability was missing in this run. The
installed GitHub connection applied the base files, merged master, then
restored the tested combined catalog. No history was rewritten.

The previous head `2d32a0094` passed all remote gates and had Greptile
5/5. Those results do not verify this new head.

The original PR reports live setup, discovery, settings headers, gateway
calls, and context-header forwarding. This repair does not repeat those
account-bound checks. The permission record still marks maintainer live
qualification as outstanding.

## Risks

- Telem.AI receives the six context IDs by default. The saved header
policy controls forwarding. Search use is billed to the account that
owns the key.
- All agents on a connection share its search settings.
- The merge uses master's route-test setup unchanged. The
company-boundary assertions remain intact.
- No schema, dependency, or workflow change is included.
- Live provider evidence is attributed to the original contributor.
Maintainer live qualification remains outside this CI repair.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`),
1M-token context, through Claude Code with shell, editing, and test
tools, as disclosed in #15302.
- CI repair and review: OpenAI `gpt-6-astra`, through Codex with
reasoning, shell, editing, and GitHub tools. The runtime does not expose
the context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Yifei Ai <aiwen324@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:33:18 -07:00
Devin Foley 8cbd21b3e7 feat(apps): add Superagent connection (#15394)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents reach external services through the Apps catalog. Each
catalog entry is a reviewed `AppDefinition` that connects a provider's
hosted MCP server to Paperclip's shared vault, grants, policies,
gateway, and audit trail.
> - Superagent (`superagent.sh`) is a security platform. Its hosted MCP
server lets agents read and triage security findings, start red-team
reports, run Contributor Trust and dependency update jobs, and score web
pages, email, files, skills, MCP repositories, and packages before an
agent trusts them.
> - Superagent is not in the catalog. Operators must use the generic
"Connect your own MCP server" flow. That flow has no branding, no key
guidance, and no warning about billable or destructive tools.
> - The connector playbook supports this provider with existing
definition fields. The server accepts an organization API key as a
bearer header. It publishes no OAuth authorization-server metadata, so
browser sign-in is not possible.
> - This pull request adds the Superagent definition, its artwork, the
research and permission-review ledger rows, documentation, deterministic
tests, and one reviewed risk-classification rule.
> - The benefit is a branded, governed Superagent connection with clear
key guidance and a warning about tools that cost credits or delete data.

## Linked Issues or Issue Description

**Problem or motivation**

Teams that use Superagent for PR security, red teaming, and agent
guardrails want their Paperclip agents to read findings, return
structured reports, and score content before they use it. Superagent is
not in the Apps catalog. Operators must paste the MCP URL and an
`Authorization` header into the generic remote-MCP flow. That flow gives
no branding and no provider guidance. It also does not tell the operator
that the key reaches the whole organization, or that some tools consume
credits or permanently delete findings.

**Proposed solution**

Add a catalog-only Superagent connection that follows the connector
playbook. It has one method: a customer organization API key
(`sk_live_...`), sent as an `Authorization: Bearer` header to
`https://www.superagent.sh/mcp`. The field helper text explains that
Superagent keys are not scoped. The method warning tells operators to
set billable and destructive actions to Ask first before agents run
unattended. All discovered tools stay governed by the normal per-action
policies.

**Alternatives considered**

A browser sign-in method was not added. The server's protected-resource
metadata names `https://superagent.sh` as its authorization server, but
that origin publishes no `oauth-authorization-server` or
`openid-configuration` document, so Paperclip cannot discover OAuth
endpoints. A plugin was not needed because the connection needs no
custom UI, tables, workers, or webhooks. Relying on the generic risk
classifier was not enough. Several Superagent mutations
(`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names
that it reads as reads, so a narrow reviewed Superagent rule was added
instead.

**Roadmap alignment**

This extends the existing self-serve remote-MCP connection catalog. It
does not overlap planned core work.

## What Changed

- Added the `superagent` row to
`packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk
tier S4).
- Added the `superagent` provider to
`scripts/ingest-app-definitions.mjs` (category, key placement and
placeholder, console links, guidance, description). Regenerated
`packages/shared/src/app-definitions/superagent.json` and the generated
registry.
- Added the `superagent/mcp-api-key` permission review to
`doc/connections/tool-method-permission-reviews.json`, with
key-permission text and evidence links.
- Added Superagent's official mark
(`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the
official `superagent-ai` GitHub organization) and the brand manifest
entry.
- Added gallery copy for the Superagent card.
- Added a reviewed Superagent rule to `classifyRisk` in
`server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools,
and tools that Superagent marks read-only, are reads. `delete_*` and
`revoke_agent_client` are destructive. All other tools are writes, so
billable and rule-changing tools can be set to Ask first.
- Put the Ask-first advice in the API-key helper text, because the key
form shows helper text and not method warnings.
- Added `doc/connections/SUPERAGENT.md` (transport and auth, why there
is no OAuth, administrator setup, capabilities and policy, manifest,
brand provenance, validation hook). Linked it from the connections
README and the permission audit.
- Tests: definition shape, store visibility and artwork, URL
recognition, the bearer header on discovery with the key kept out of
connection config, read/write/destructive classification of fixture
tools, the Superagent risk rule (including `triage_finding` and
`restore_agent_builtin_rule`), the visible Ask-first advice and API-key
gating of the connect form, and the pinned catalog counts.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
packages/shared/src/app-definitions-url.test.ts
ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed.
- `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts`: 385 passed.
- `node scripts/check-app-brand-assets.mjs` and `node --test
scripts/app-brand-validation.test.mjs`: passed.
- `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter
@paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui
typecheck`: clean.
- Manual: in a local instance, open Apps → Browse and confirm the
Superagent card and icon. Open `/apps/connect?source=superagent`.
Confirm the single API-key method, and confirm that Connect enables only
after a key is entered.
- Live metadata probe on 2026-10-06: an unauthenticated `initialize` on
`https://www.superagent.sh/mcp` returns 401 with
`resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`.
That document returns 200. No authorization-server metadata exists at
the named issuer.

## Risks

- Low risk to existing providers. The change is additive catalog data
plus tests. The generated registry only gains one import. The new risk
rule runs only for Superagent connections.
- A Superagent key reaches its whole organization. Some tools consume
credits (`create_*_report`, `triage_finding`) or delete data permanently
(`delete_finding`). Every action starts Allowed under the current
product default. The key helper text tells operators to set these
actions to Ask first.
- The permission-review ledger records live proof as not run. No
Superagent account was used. The lifecycle checklist in
`doc/connections/SUPERAGENT.md` needs a documented pass before the entry
is fully qualified.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with
extended thinking and tool use (shell, file editing, web fetch, browser
checks).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-10-06 15:57:16 -07:00
DottaandPaperclip 2ca0d26a99 fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 15:36:01 -05:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
88ff98b83d feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path

> - Paperclip agents need governed access to Enterpret's hosted MCP
tools for customer feedback.
> - The official Enterpret MCP is read-only. Enterpret Agent’s beta
write MCP is a separate service and is outside this PR.
> - This PR supports both organization auth tokens and personal browser
OAuth for the official MCP.
> - Enterpret committed to fixing OAuth scope behavior and reducing its
token-revocation cache window. The provider fixes will follow after
connector deployment and do not gate this release. Paperclip functional
validation remains required. New tools stay quarantined on refresh.
> - This PR combines the original connector and Vivek's token-path work
on current master, with live self-hosted QA recorded in the validation
guide.

## Linked Issues or Issue Description

No public issue covers this connector. This PR incorporates the work
from #14737. The independent OAuth scope-provenance fix is #14059; it is
not required for the organization-token path. #13890 is a merged
connector precedent. #13692 provides the connector-authoring skill.

**Problem or motivation**

The generic MCP setup can reach Enterpret, but it has no reviewed
Enterpret entry, method-specific setup, artwork, or policy defaults. The
original connector branch had also fallen behind master.

**Proposed solution**

Publish the Enterpret catalog entry with an organization auth token as
the selectable method. Support browser OAuth through instance-local DCR
on the same read-only endpoint. Default `run_graph_query` to Ask first
and quarantine tools newly found on refresh.

## What Changed

- Added the Enterpret app definition, generated registry, official icon,
Browse card, setup copy, tests, and Storybook states.
- Combined #14737 and resolved the current-master conflicts, including
the connector permission audit.
- Preserved catalog quarantine on token reconnect and updated expiry
guidance to follow the token's dashboard expiry. The QA token displayed
three months, contradicting the former six-month copy.
- Enabled personal OAuth for the official read-only MCP. Removed the
write-capability inference and OAuth draft-only setup restriction. The
beta Agent endpoint is excluded.
- Preserved tool quarantine during token reconnect, with a regression
test for existing quarantined and newly discovered tools.
- Added OAuth scope and revocation-cache disclosures and an interactive
Storybook retry check.
- Merged master `d9f600043` into the existing PR branch without
rewriting history. Resolved four conflicts, regenerated the combined
registry, and retained the model-provider assertions with a catalog
count of 66.
- Recorded live QA and its limits in
[doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md).

## Verification

Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513
focused connector, discovery, OAuth, policy, shared-definition, and
UI-handoff tests passed across 11 suites. Full repository typecheck,
build, Storybook build, token gates, module boundaries, Node policy,
no-git-push policy, and typecheck-build-gaps passed. Release-registry
tests passed 129/129 with modern Bash. The four failures with macOS Bash
3.2 also reproduced on untouched master.

The full test suite passed in CI on this exact head, including every
general-test and serialized-server shard and all eight E2E shards. The
duplicate local `pnpm test:run` was stopped after these remote test
lanes passed; it did not complete locally. The combined registry was
regenerated. Regeneration also reproduces existing Linear and AgentMail
definition drift on untouched master; those unrelated changes were
excluded.

On an isolated self-hosted Paperclip instance, an organization token
connected and discovered eight actions. `get_organization_details`
returned the intended organization before and after reconnect. Catalog
refresh, invalid-token failure and valid-token recovery, activity
logging, local disable, and local removal were exercised. No feedback
quotes or graph query were requested. Paperclip revoked both QA
connection secrets and archived the QA connections.

The token-path agent access preview correctly showed an action Off, but
the board Test endpoint still invoked it as the board user. This shared
Test-path mismatch is documented; it is not evidence that a real agent
can bypass policy. An actual token-path agent-session denial, a real
agent process, expiry recovery, server/VPS, and Cloud remain unverified.

Enterpret's dashboard confirmed the QA token was revoked, but the same
token immediately reconnected and completed a safe read. Vivek explained
this as a 24-hour validity cache and committed to reducing the window.
The new maximum delay and deployment remain unverified. The earlier
OAuth check reported `mcp:write` and `email` for an `mcp:read` request.
This is a scope mismatch; it did not prove access to write tools. On
2026-10-01 the account holder relayed Vivek’s clarification that the
official MCP is read-only and the beta Agent write MCP is separate.

## Risks

The organization token can read the organization's customer feedback,
and provider-side revocation did not immediately stop access. Operators
should check each token's displayed expiry and remove Paperclip's local
connection when retiring access. The account holder confirmed the
provider fixes are fast follows after deployment. Release can proceed
with the current scope reporting and up-to-24-hour revocation delay
documented, once fresh OAuth authorization/refresh, real-agent
allow/deny, CI, and review pass. Retest the reduced revocation window
after Enterpret deploys it. Scope labels alone must not be used to claim
write access.

Greptile is 5/5 on head `6c73d3084` with no actionable findings. All
current-head CI checks pass, including the canary dry run. Live
agent-session denial and fresh OAuth refresh remain explicit deployment
QA items. The board Test panel does not prove agent authorization.
Provider scope-reporting and revocation-cache fixes are accepted fast
follows after deployment.

## Model Used

Original connector and validation record: Claude Opus 5 through Claude
Code. Integration, current-master reconciliation, and token-path QA:
OpenAI Codex, GPT-6 family (exact host variant not exposed), using
repository tools and an isolated runtime. Contributor commits retain
their authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal issue
ID
- [x] I have run relevant local tests and checks
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 with no actionable follow-ups

If squash-merging, preserve contributor credit in the squash message:

Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com>

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Vivek Kaushal <vivek@enterpret.com>
2026-10-06 14:35:51 -05:00
DottaandPaperclip 9f7057e122 feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use a harness, a model, and a credential to run tasks.
> - Connections already store credentials and control who can use them.
> - Custom providers also need an endpoint and a supported API format.
> - A per-agent endpoint would duplicate credentials and access rules.
> - This pull request stores routing on the connection and projects it
into the harness.
> - Isolated credentials and protocol checks keep the selected
connection authoritative.

## Linked Issues or Issue Description

Refs #37, #13083, #14104, #14565, #12692.

#14016 is a reference only. This PR has its own schema, vault
persistence, routing validation, runtime projection, and tests. None of
the five commits in #14016 is an ancestor of this branch. We do not
depend on or plan to merge it. #14967 addresses task-pinned account
pools. #14422 addresses another provider integration.

This is the first of two linked PRs. Merge this connection change before
#15341, which refines agent setup and adds the qualification harness.
The split keeps each review below 100 changed files. Provider catalog
entries have usable setup forms in this PR. Local browser subscription
sign-in is included.

## What Changed

- Store non-secret routing metadata on AI connections. Vault provider
API keys, including Bedrock bearer API keys. Reject general AWS access
keys.
- Enforce company, owner, human audience, agent access, connection
status, and protocol checks before resolving credentials. Keep reconnect
destinations immutable and retain connection identity during key
rotation.
- Project OpenRouter and compatible custom endpoints into Codex, Claude,
OpenCode, and local Hermes. Carry these settings through both legacy and
native runner transports. Clear conflicting host credentials and redact
keys from diagnostics.
- Preserve older OpenRouter accounts and native personal defaults. Add
Google API-key accounts and migration `0306` for the two
provider-default constraints.
- Run local Claude and Codex subscription sign-in behind the existing
browser sign-in card. Use private attempt homes and owner-bound
completion instead of a copied terminal command.
- Seed isolated Gemini authentication and preserve OpenCode workspace
permissions. Keep the selected connection authoritative. The independent
Gemini and Grok workflow fixes are in #15341.
- Keep native OpenCode custom gateway keys in a runner-owned
selected-model proxy; the harness config contains only a session-scoped
capability. Honor runtime outgoing proxy and certificate settings.
Preserve streamed responses and revoke the proxy on close or startup
failure.
- Allow ordinary members to connect native personal accounts before an
agent exists.
- Repair routed accounts from task cards using the saved provider
destination, protocol, model aliases, and connection identity.
- Add provider catalog definitions, model discovery, pinned logos, and
complete native and routed setup forms. Allow a personal routed
connection before a new agent exists. Keep endpoint authentication keys
out of Hermes terminal children.
- Recover cancelled or restarted browser sign-in with a clear restart
action. Support no-auth endpoints without a vault credential. Add
isolation and recovery regressions and runtime documentation.

## Verification

- Updated with `origin/master` at `22a3ea341`. Migration `0306` follows
the new master migration and passes migration and snapshot checks.
- The integrated connection regressions passed 152 tests and 50 native
OpenCode driver tests, including key-free child-shell configuration
reads, authenticated/no-auth forwarding, streaming, model/path
restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass.
Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34
executable also completed a turn through the proxy against a local
synthetic provider; the reusable key was absent from its config. A
second real-executable smoke passed with an HTTPS CONNECT proxy and
runtime-specific synthetic certificate trust. Certificate-file and
certificate-directory regressions pass.
- Task-card repair passed 48 tests, including OpenRouter, Bedrock, and
custom gateway reconnect cases. UI typecheck and token gates passed.
- The prior core regression set passed 133 tests across new-agent setup,
provider forms, browser sign-in, routing projection, and connection
authorization. Token gates and UI typecheck passed.
- Full workspace typecheck and production build passed again after the
latest integration and credential-proxy fix. The merged deterministic
runner E2E suite passed 1,400 Vitest tests and 128 Node tests.
- Full workspace typecheck passed on the prior linked combined
implementation. Production build, Storybook build, 1,316 browser-harness
Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the
latest master integration and regenerated migration. This exact head
passed 54 remote checks with four expected skips and Greptile 5/5; no
review threads remain open. An unchanged server fixture had a random
six-character issue-prefix collision on its first attempt. All 245 tests
passed locally and the single CI retry passed.
- A provider-free terminal check used the cited supported Hermes source
and dummy keys. Gateway and OpenRouter terminal children could not read
the selected key.
- The broad local Vitest attempt passed 15,442 tests but was not green.
It had an embedded-Postgres startup failure, an HTTP logger timeout, an
origin socket error, and a browser cancellation wait timeout. The
cancellation wait was corrected. The relevant connection tests and the
full origin test file passed separately. Latest-head CI must pass before
merge.
- Prior credential-backed acceptance exercised task creation, tool use,
artifact delivery, completion, and context-dependent follow-up. Claude
legacy and native runners passed Bedrock with `us-east-1` and
`us.anthropic.claude-sonnet-4-6`.
- Historical local qualification retained 43 passing API/gateway cells
out of 46. Those attempts span earlier builds. They do not qualify this
exact commit or staging. All subscription combinations and staging
remain unqualified.
- Verify native subscription and API-key setup. Connect a regular
provider catalog row. Verify an incompatible harness and a changed
reconnect URL are rejected. Use #15341 for the complete browser
campaign.

## Risks

- Migration `0306` changes two check constraints. It preserves rows and
is safe to reapply. It takes normal constraint-change locks.
- Credential projection touches several harnesses. CLI upgrades can
change provider configuration and session behavior.
- The native OpenCode proxy adds a loopback hop, pins requests to the
selected model, limits request bodies to 16 MiB, rejects redirects, and
expires at session close. It prevents reusable keys in the child
configuration; it is not an OS isolation boundary against a process
debugger running as the same user.
- Custom endpoints must be reachable from the agent environment. Saving
a connection does not prove connectivity. Bedrock keys require rotation
before expiry.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Provider overloads and an unresolved follow-up timeout also
affect live Gemini qualification. We have not patched the installed CLI
or marked those cases as passing.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS
identity, arbitrary auth headers, and custom routing for other harnesses
are excluded.
- These PRs do not establish production or staging qualification for
every provider and login method.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:29:03 -05:00
DottaandPaperclip 0e0b63e5a5 feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - AI Connections separate account access from models and harnesses.
> - A pool must act as one connection while retaining each task’s
account.
> - Core must enforce member access and preserve session and recovery
rules.
> - A plugin supplies rotation policy without receiving credentials.
> - This change adds durable routing and native connector setup and
management.

## Linked Issues or Issue Description

**Subsystem affected**

AI Connections, Connectors, plugins, run dispatch, and session
compatibility.

**Problem or motivation**

Operators need to rotate new tasks across saved accounts while each task
keeps its account and session. Pool setup must fit the existing
connector catalog and account workflow.

**Proposed solution**

Add an experimental router binding, a capability-gated plugin hook, and
transactional task pins. Plugins declare native pooled connectors
through `aiConnectionRouter`. Core hosts the existing-account picker,
ordering step, and account settings. Related usage contract: #14936.
Companion private plugin:
https://github.com/paperclipai/paperclip-cloud/pull/643.

**Roadmap alignment**

This extends Apps and AI Connections. Core supplies generic enforcement
and native connector UI; the private plugin owns rotation and quota
policy. The prior duplicate search found no matching router
implementation.

## What Changed

- Add a router binding without changing existing concrete bindings. Keep
the instance flag and new pools disabled by default. Require manual
operator configuration. Show no routing toggle in Experimental settings
on either open-source or Cloud installs, even after routing is enabled.
- Persist company-scoped pools, one shared cursor per pool, and pins
keyed by company, pool, agent, and task. Commit pins and cursor advances
together with revision checks and bounded retries. Persist run-ID
affinity before allocation.
- Pass only authorized metadata and normalized usage to plugins. Core
retains credential handling, member access checks, runtime
qualification, and recovery evidence. Probe outside locks with a shared
15-second budget and freshness cache.
- Resolve routing before credential preparation and backend selection.
Preserve pins through turns, session resets, removed members, and quota
waits. Retain admitted recovery after disable or uninstall.
- Separate credential session epochs from token generations. Verified
refresh preserves the epoch; reconnect and manual replacement change it.
Include the credential slot ID in session and usage-cache identity, so
reconnecting an indexed legacy account invalidates its old session even
when both epochs are zero.
- Validate pool member installations before accepting saved-agent
bindings and recheck compatibility when the harness changes. Install
only authorized members in the new-agent transaction and record their
IDs in local activity. Pool membership cannot install a restricted
shared connection.
- Preserve pool bindings when agents hire teammates through either
creation API or native caller runtime inheritance. Block stale manager
credential references; retain explicit child authentication precedence
and reject incompatible inherited pools.
- Add native connector registration through plugin metadata. Reuse the
Connectors catalog, setup header, account header, sidebar, dialogs, and
usage display. Setup selects and orders saved connections. Advanced
settings hold usage rules and member runtime defaults. New-account setup
opens in another tab.
- Use revision-checked pool archival from the Connectors catalog and
account page. Keep task pins, cursors, recovery evidence, and underlying
connections. Reject ordinary connection updates or removals that bypass
pool revisions.
- Add pool selectors, composer models, override notes, quota status, run
details, activity records, and local run-log records. Keep
session-adoption copy minimal.
- Show **Used by** below the pool connections. List current company
agents with shared avatars and profile links. Include paused agents;
exclude terminated agents and agents using another pool.
- Add Core stories for the generic connector workflow and runtime
surfaces. Cloud stories reuse these production routes and tokens through
a preview-only alias.

## Verification

- Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace
`pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
locally.
- All 422 focused connector/settings/shared-contract/migration tests and
all 156 database-backed AI connection, hiring, reconnect, and
durable-routing cases pass (69 hiring cases rerun after the final
auth-precedence fix). The merged shared contract retains connection
instructions and pool metadata. The pool migration is generated at
sequence 0299 after the latest upstream migrations; this PR makes no
lockfile changes.
- All four full-app Playwright tests pass on the final head after a cold
restart and migration, against the installed private plugin and isolated
database, with no route or pool-API mocks. They cover hidden routing
controls after manual opt-in, native pool creation, ordering, membership
edits, rename, paused defaults, enabling/save/refresh persistence, stale
edits, cancellation/removal, preserved underlying accounts, unavailable
routers, and Used by avatars and profile links. Exact command:
`PAPERCLIP_CONNECTION_POOL_E2E=1
AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f
AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test
--config tests/ai-connections-app/playwright.config.ts
connection-pools.spec.ts`.
- [Native setup, ordering, and management
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278)
address the review follow-up. [Earlier selector, quota, and run-detail
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316)
show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui
storybook`, then **Connectors / Pool host** or **AI Connections /
Connection pools**. Cloud owns its host-backed plugin stories; both
repositories’ Operator Setup Required story assertions pass.
- Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed
both exact sessions after restart, preserved pinned accounts through
explicit reset and controlled quota deferral/recovery, and committed
only two allocations across fourteen runs. A later UI-created task test
again rotated OpenAI then Anthropic and resumed OpenAI through
follow-up/restart/quota recovery. That later Anthropic execution was
blocked by its saved OAuth token expiring (provider 401). No live usage
probes ran.
- The full local `pnpm test:run` was attempted earlier and did not
complete because of macOS embedded PostgreSQL bootstrap/shared-memory
failures and the 40,000-file Git fixture timeout. The focused database
suites above now pass; full-suite verification is provided by the split
CI lanes. The preceding CI run had one runtime readiness timeout; it
passes locally both alone and inside the larger runtime suite. That
larger local suite also encountered an embedded PostgreSQL setup failure
and two macOS temporary-path alias assertions; those two assertions pass
with canonical TMPDIR=/private/tmp. All final-head CI checks are
terminal green, including full general/serialized server suites, Runner
checks, browser E2E shards, canary verification, build, and typecheck.
Greptile is 5/5 on that exact head with no unresolved threads.

## Risks

- The migration adds routing tables and a credential epoch column.
Install the private plugin only with the compatible Core contract.
- Routing and each pool require opt-in. Production distribution and
fleet defaults remain unchanged.
- Unknown usage stays eligible. Known pinned exhaustion waits; revoked
access requires operator repair.
- Legacy adapters require compatible members. Runner model and effort
overrides remain limited by qualified backend support.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes:`
/ `Refs:` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites;
full-suite limitations are reported above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 20:44:37 -05:00
DottaandPaperclip 4857799a88 feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes.

Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:16:00 -05:00
DottaandPaperclip b43073d11f feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - Aggregator gateways can expose accounts that users already connected
upstream.
> - The Apps catalog did not show those accounts or their current
provider status.
> - Separate cards and setup tasks also made account ownership unclear.
> - This pull request discovers upstream accounts and groups them under
one app card.
> - Users can find connected apps while each provider keeps control of
its accounts.

## Linked Issues or Issue Description

**Subsystem affected**

Connections across the database, shared contracts, server, and board UI.

**Problem or motivation**

Users cannot see which apps are connected through a saved aggregator
gateway. Native and upstream accounts need one app card. Discovery must
preserve company, user, gateway, and credential boundaries.

**Proposed solution**

Sync account metadata from Composio, Arcade, and supported Executor
gateways. Keep upstream account management in each provider. Use source
chips and search to browse the catalog. Preserve native setup and the
gateway's existing access policy.

**Alternatives considered**

Creating a local executable connection for each upstream account would
duplicate authorization state. Using an agent task for routine Composio
setup would add an unnecessary step. The board now calls the saved
gateway directly for that setup.

**Roadmap alignment**

This extends the shipped Connected Apps and MCP Tool Gateway features in
ROADMAP.md. The duplicate search found no open PR for managed account
discovery.

Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open
PR #12906 covers adjacent toolkit routing work.

## What Changed

- Add provider-neutral discovery, sync, and refresh APIs. Preserve the
Composio API paths.
- Cache observations by company, saved gateway, viewing user, and
credential version. Retain stale observations after failed or incomplete
scans.
- Add optional Arcade account sync credentials in the vault. Discover
Executor accounts through its supported inventory interface.
- Group native and upstream accounts in one app card. Imported account
menus open their provider. Gateway menus own refresh and sync setup.
- Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50
catalog entries per page. Keep connected accounts above discovery. Keep
explicit provider searches scoped.
- Simplify Composio app setup and refresh its connected app list on the
gateway Permissions page.
- Add a compact agent access card and task creation defaults for
connection setup. Preserve explicit blocks and approval policies.
- Add two replay-safe migrations, service and UI tests, Storybook
journeys, and acceptance stories.

## Verification

- Passed the repository typecheck, full build, token gates, and
migration ordering check.
- Passed the focused provider adapter, connection interaction, and
catalog tests after rebasing onto master.
- Passed all nine database sync and migration replay tests using a
disposable database on the test-drive PostgreSQL cluster. Removed that
database after the run.
- Verified Arcade cursor pagination against its official Go SDK and
passed all eight adapter tests, including short and incomplete pages.
- Passed all 45 interaction tests after making the exact requested tools
and their Allowed/Ask first permissions visible before granting access.
Verified the compact card in Storybook.
- Passed the complete UI suite on the final code: 683 files and 7,432
tests, including the corrected Composio destination assertions. Passed
130 focused tests for the UUID, management-link, and health-status
corrections.
- Passed 22 Composio setup/sync tests, 23 connection-intent service
tests, and the connection migration test in separate disposable
databases. Database startup alone was substituted; the suites exercised
their real SQL and services.
- Passed all 10 OpenAPI route checks and the full-stack
connection-intent browser test, including scoped consent, agent
continuation, and task completion.
- The local full runner encountered embedded PostgreSQL startup failures
on this loaded macOS host. The earlier in-flight run also held the
pre-fix Arcade transform; a fresh run of the final provider suite
passes. The final-head CI is queued during GitHub’s active Actions
incident: https://www.githubstatus.com/. The previous run also lost
several runners simultaneously; its real catalog assertion failures are
fixed and the fresh complete UI suite passes.
- Tested the real test-drive server in the embedded browser with a live
Composio gateway. Detected Airtable and Circleback. Verified refresh
progress, account rows, source chips, search scope, and 50-entry
pagination.
- Arcade and Executor coverage uses provider fixtures. Live credentials
were unavailable.
- Storybook builds successfully and includes grouped native/provider
accounts, stale and unavailable discovery, optional Arcade setup, and
mobile states. The acceptance document records the simulated and live
coverage separately.

- Greptile reviewed final commit
`217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review
threads are resolved, security scans pass, and the PR has no merge
conflicts. The outstanding remote checks are `ci / Select trusted
runner` and `review`, queued by GitHub. They need to complete before
merge.

## Risks

- Provider response changes can break inventory discovery. Failed scans
retain observations and show stale status.
- Composio scans only the supported catalog and can take time. Large
inventories run in the background with progress and a bounded lease.
- Arcade requires a project API key and user ID when the gateway cannot
supply them. This key is used only for discovery.
- Executor discovery depends on the server's exposed inventory tools.
Unsupported servers report unavailable discovery.
- Cached account rows do not grant access or create executable
connections. Gateway policies still govern tool use. Account deletion
and per-app authorization remain upstream.
- The migrations add tables and one nullable column. Replay preserves
existing rows and company-scoped foreign keys.

## Model Used

OpenAI Codex, based on GPT-6. The session does not expose a more
specific serving model ID or context limit. Used reasoning, repository
tools, code execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 14:57:04 -05:00
Devin FoleyandPaperclip 5b8b2b38ca feat(apps): add Neon connection (#14980)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents reach external services through the Apps catalog. Each
catalog entry is a reviewed `AppDefinition` that wires a provider's
hosted MCP server into Paperclip's shared vault, grants, policies,
gateway, and audit trail.
> - Neon is a widely used serverless Postgres provider with an official
hosted MCP server, but it is not in the catalog. Teams that run their
databases on Neon must use the generic "connect your own MCP server"
path, which has no branding, no guidance, and no project or read-only
controls.
> - The connector playbook requires a catalog entry for a provider like
this: the hosted server supports dynamic client registration and bearer
API keys, and the common definition fields can express every option
Paperclip can serialize.
> - This pull request adds the Neon definition, its official artwork,
the research and permission-review ledger rows, documentation, and
deterministic tests, without any provider-specific runtime code.
> - The benefit is a one-click, governed Neon connection with optional
project pinning and read-only mode, and a documented path to live
qualification.

## Linked Issues or Issue Description

**Problem or motivation**

Neon is a common Postgres host for the applications agents work on, but
Paperclip's Apps catalog has no Neon entry. Operators who want agents to
inspect schemas, run SQL, or manage branches must paste the MCP URL into
the generic remote-MCP flow, which gives no branding, no provider
guidance, no project boundary, and no read-only switch.

**Proposed solution**

Add a catalog-only Neon connection built from the connector playbook:
browser sign-in through Neon's dynamic client registration with the
reviewed `read` and `write` scopes, or a customer API key sent as an
Authorization bearer header. Both methods expose Neon's documented
`projectId` pin and `readonly` switch as optional Advanced fields. Every
discovered tool stays governed by the normal per-action policies.

**Alternatives considered**

A plugin was not needed because no custom UI, tables, workers, or
webhooks are involved. A separate read-only method was not added because
the playbook treats read-only switches as advanced fields rather than
methods. Neon's repeatable `category` query filter was left out because
tenant fields serialize lists as one comma-joined value, so it cannot be
sent correctly without new runtime code; per-action policies cover
catalog narrowing instead.

**Roadmap alignment**

This extends the existing self-serve remote-MCP connection catalog and
does not overlap planned core work.

## What Changed

- Added the `neon` provider to `scripts/ingest-app-definitions.mjs`
(category, API-key placement, methods, tenant fields, guidance,
warnings) and regenerated
`packages/shared/src/app-definitions/neon.json` plus the generated
registry.
- Added the Neon row to the self-serve MCP research ledger with
`dcr_or_api_key` auth and risk tier S4.
- Added permission reviews for `neon/mcp-oauth` (explicit scopes `read`,
`write`, taken from Neon's live authorization-server metadata) and
`neon/mcp-api-key` (provider key), with evidence links.
- Added Neon's official tile icon (`ui/public/brands/apps/neon.png`,
copied byte-for-byte from the icon linked by neon.com) and the brand
manifest entry.
- Added prosumer gallery copy for the Neon card.
- Added `doc/connections/NEON.md` (service involvement, endpoints,
administrator setup, capabilities and policy, manifest, brand
provenance, validation hook) and linked it from the connections README
and the permission audit.
- Tests: Neon definition shape, store visibility and artwork, URL
recognition, reviewed scopes with scope-widening rejection, URL
projection of the project pin and read-only flag, invalid project ID
rejection, the connect form's API-key gating, and the pinned catalog
counts.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
packages/shared/src/app-definitions-url.test.ts` — 34 passed.
- `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts` — 367 passed.
- `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — all passed.
- `node scripts/check-app-brand-assets.mjs` and `node --test
scripts/app-brand-validation.test.mjs` — passed.
- `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter
@paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui
typecheck`, `pnpm check:token-gates` — clean.
- Manual: in a local instance, open Apps → Browse, confirm the Neon card
and icon, open `/apps/connect?source=neon`, confirm both methods, the
Advanced project pin and read-only toggle, and that Connect enables
after an API key is entered. The operator completed a live connection
against a Neon account on this build.
- Live metadata probed on 2026-10-02: both `.well-known` documents at
`mcp.neon.tech` return the recorded endpoints and scopes; an
unauthenticated `initialize` returns 401 with `resource_metadata`.

## Risks

- Low risk to existing providers: the change is additive catalog data
plus tests. The generated registry only gains one import.
- Neon's hosted server grants broad project and database management. The
definition carries two warnings, recommends a development project, and
keeps every write under the normal action policies; the read-only switch
is enforced by Neon's server, not locally.
- The permission-review ledger records live proof for both methods as
not run; the full lifecycle checklist in `doc/connections/NEON.md` still
needs a documented pass before the entry is considered fully qualified.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended
thinking and tool use (shell, file editing, browser verification).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:04:52 -07:00
DottaandPaperclip 839cac1343 fix: request supported offline access for generic MCP OAuth (#14950)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use external MCP tools through the governed gateway.
> - Remote MCP connections can use OAuth access tokens that expire.
> - Some providers issue refresh tokens only after an offline-access
request and consent.
> - Resource scopes hid that identity-provider capability in the generic
connection flow.
> - This pull request requests supported offline access and tests expiry
through a local MCP server.
> - The benefit is continued tool access without another sign-in when
the provider permits refresh.

## Linked Issues or Issue Description

Related PR: #13447 addresses the same OAuth symptom together with
managed Codex configuration. This PR focuses on generic MCP OAuth. It
also covers consent, explicit scope overrides, legacy reconnects, exact
scope persistence, and real HTTP expiry tests.

**What happened?**

A generic MCP resource can advertise only its tool scopes. Its OAuth
server can separately advertise `offline_access`. Paperclip selected the
resource scopes and omitted the offline-access request. A provider could
then issue an access token without a refresh token. Tool access stopped
after the access token expired.

**Expected behavior**

Paperclip adds advertised offline access to the selected tool scopes
when the OAuth server does not exclude refresh tokens. It requests
consent and stores the scopes sent in the authorization request.
Existing connections can discover this capability when the user
reconnects. Providers without this capability keep their existing scope
behavior.

**Steps to reproduce**

1. Run `node scripts/mcp-fixtures/servers/oauth-refresh-fixture.mjs`
from the repository root.
2. Add its MCP URL as a generic connection on a local Paperclip
instance.
3. Approve the test consent page and call `read_status`.
4. Let the two-minute access token expire and call the tool again.
5. Before this fix, the connection needs another sign-in. With this fix,
the call refreshes the token and succeeds.

**Paperclip version or commit**

The integration regression reproduced the missing-refresh-token failure
on the parent of this PR's fix. The same test passes with the fix.

**Deployment mode**

Local development. Automated tests use a loopback HTTP MCP/OAuth server
and a disposable PostgreSQL database.

## What Changed

- Track offline-access capability separately from MCP tool scopes.
- Add supported offline access and consent for generic connections.
- Preserve the actual requested scopes through callback completion and
reconnect.
- Discover the capability for older connections with cached OAuth
endpoints.
- Add a reusable MCP/OAuth fixture with PKCE, token expiry, resource
binding, and refresh-token rotation.
- Test shared and personal gateway calls through two refresh rotations.
Cover scope selection, unsupported refresh, and legacy reconnects.
- Document the behavior and local test commands.

## Verification

- The two real HTTP expiry tests failed before the fix with
`oauth_refresh_missing` after the first token expired.
- Tests passed on the current head: 76 generic MCP regressions, all 365
tool-access service tests, and 4 fixture controls.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed. Server
TypeScript checks also passed after the review fixes.
- All remote checks passed on
`d22909bb34e9d54478c0002077c498ffe105932d`. Greptile gave 5/5 with both
previous findings resolved. The complete local `pnpm test:run` is still
running.
- Run `node --test
scripts/mcp-fixtures/servers/oauth-refresh-fixture.test.mjs` for the
standalone provider controls.
- Run `pnpm exec vitest run
server/src/__tests__/generic-mcp-connection.test.ts` for the Paperclip
integration tests.

## Risks

- Users can see a consent prompt when a generic provider supports
offline access.
- The provider can still decline to issue a refresh token. Access works
until expiry, then the user must reconnect.
- A provider that advertises offline access but rejects the scope
produces an OAuth error. This PR does not add an automatic retry without
that scope.
- Existing grants without refresh tokens need another sign-in. The fix
does not change them in place.
- Curated Apps keep their reviewed scope and authorization-parameter
allowlists. No database migration is required.
- This simulation verifies the suspected failure. The reported internal
MCP server has not been tested.

## Model Used

- OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The runtime did not expose a more specific model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 14:13:59 -05:00
DottaandPaperclip 7d59de6113 feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections store the AI accounts used by legacy and native runners.
> - Subscription accounts can reach session, weekly, model, or paid
usage limits.
> - Operators need to read these limits for a specific stored account
before making a routing decision.
> - This pull request adds an on-demand usage probe to the connection
service and account detail.
> - The result preserves provider limits, reset times, paid usage, and
unknown values for later consumers.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, the connection service and API, and the account detail
UI.

**Problem or motivation**

Managed AI accounts lack a common operation to read their current usage
limits. A local harness probe can read a different login from the
account selected for an agent.

**Proposed solution**

Add `aiConnectionService.probeUsage()` and a board-only connection usage
endpoint. Probe the selected credential grant on request. Support Codex,
Claude, and Grok subscriptions, plus OpenRouter API key limits.

**Alternatives considered**

Harness-specific automatic polling would couple the read to execution
and can read ambient credentials. This change uses the managed
connection credential and leaves scheduling and admission decisions to
later work.

**Roadmap alignment**

This extends the existing Personal & Shared AI Accounts capability. It
adds no routing or quota enforcement. Related: Refs #14459 for managed
OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing
and budget work. This operation reads one requested account across all
three subscription providers.

## What Changed

- Add typed usage snapshots and a probe capability flag to managed AI
connections.
- Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model
scopes, provider admission, reset periods, and paid allowances separate.
Preserve unknown values.
- Enforce company membership, credential audience, grant identity, and
connection lifecycle before reading the stored secret.
- Add a board-only `GET
/api/companies/:companyId/ai-connections/:connectionId/usage` endpoint
with `no-store` responses.
- Add manual **Check usage** and **Refresh** actions to account details.
Show compact usage bars, resets, admission and overage status; remove
repeated descriptions and account-default copy. Clear previous results
during a new request or error.
- Add Storybook previews using the production account components for all
four providers, initial checks, loading, and permission errors.
- Add provider, authorization, runner selection, API, and UI coverage.
Document provider sources and live qualification.

## Verification

- Initial provider, authorization, selection, API, and UI validation
passed (96 focused tests): `pnpm exec vitest run
server/src/services/ai-connection-usage.test.ts
server/src/__tests__/ai-connections.test.ts
ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx
server/src/__tests__/openapi-routes.test.ts`.
- `pnpm -r typecheck` passes for the initial implementation. After
simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck,
token gates, and Storybook build pass. The initial feature module
boundary check also passed.
- Real Codex, Claude, and Grok credentials were saved to encrypted
disposable connections. The actual usage HTTP route returned 200 with
`status: ok`. Legacy and native runner selection checks passed. The
tests started no model turn and exchanged no refresh token. The
disposable databases and vaults were removed.
- Live Claude responses added structured scoped limits. Live Grok
responses omitted included-plan usage. Tests now cover both shapes and
preserve the Grok omission as unknown.
- The full workspace build passes. A full local test run hit a heartbeat
feedback timeout. That case passes in isolation. The duplicate local run
was stopped after all remote checks passed. The Slack ordering and
OpenCode transport CI flakes also pass in isolation and on the CI rerun.


- Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54
active checks pass. Two Storybook checks are intentionally skipped by
the workflow. Greptile is 5/5 with no unresolved review findings; the
branch is mergeable.

## Risks

- Subscription usage endpoints can change. Credentials can lack
usage-read permission. The probe returns explicit errors without fresh
limits in these cases.
- A successful probe can contain partial data. Missing utilization or
admission remains unknown. An enabled paid-usage switch does not prove a
funded balance.
- This change adds no migration. It does not change runner admission or
automatic provider selection. Provider requests use fixed endpoints,
disabled redirects, bounded response sizes, and a 15-second deadline.

## Model Used

OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and
HTTP tools. The session does not expose the exact runtime model variant
or context window size. Real provider credentials were used only for the
authorized live checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 11:47:14 -05:00
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00
DottaandPaperclip e00d10d5d5 fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path

> - Paperclip manages AI agents and controls the credentials used for
their work.
> - Managed AI connections resolve each responsible user's provider
default.
> - Agent settings created another account but kept the old default
selected.
> - A rejected provider test left the old account marked as connected.
> - Claude ACP reported a typed login failure as a generic
terminal-access error.
> - This pull request repairs the selected account or selects the new
login explicitly.
> - Agents can save and run with the repaired credential, and failed
logins request sign-in.

## Linked Issues or Issue Description

- Fixes #14831.
- Refs #13867. Environment failures remain separate from
credential-health failures.

## What Changed

- Add an agent-settings action to reconnect an unavailable personal
default in place. Keep its connection, grant, default, and agent access.
- State that a new account becomes the user's provider default. Select
its returned grant before changing the agent binding. Keep the actual
sign-in method.
- Show default-update errors and allow retry without another provider
login.
- Show the agent-access choice. Connection managers start with
company-wide access for their own tasks. Other members start with access
for the current agent.
- Use the server's connection-manager permission in the shared list
response. This includes members with a custom management grant.
- Mark credentials as needing attention after an explicit login
rejection in Test or Save. This includes API-key 401 and 403 responses.
Network, quota, and server failures keep the credential health
unchanged.
- Reuse the credential-generation check so an old failure cannot
invalidate a newer reconnect.
- Route Claude's typed provider `access` failure to the existing
login-recovery flow. Replace its generic terminal-access fallback with a
sign-in message.
- Add regression tests and update the AI Connections documentation.

## Verification

- Red: the UI tests failed on the missing reconnect action, unused
returned grant, missing access choice, and lost default-update error.
The server tests failed because rejected credentials stayed connected.
The real ACP fixture returned `acpx_turn_failed` for typed login
failures.
- Green: 156 tests passed across the AI connection, hiring, agent field,
and New Agent suites. All 37 environment-route tests passed. The Claude
ACP authentication fixtures also passed.
- `pnpm check:token-gates` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- The full local `pnpm test:run` passed 707 files and 14,503 tests, then
exited with an agent-conversation timeout and embedded PostgreSQL
startup failures in unchanged suites. The isolated conversation and
migration tests passed on rerun. Later local test groups did not run
after this failure.
- [All CI gates
passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356)
on commit `38513dfe2`. This includes the full test matrix, browser
tests, typecheck, build, Runner checks, and canary dry run.
- Greptile reviewed commit `38513dfe2` and returned 5/5 with no open
findings.
- The regression tests use a real embedded database and a real ACP
fixture process. Live provider sign-in requires a valid account and was
not run.

## Risks

- Connecting a new account from agent settings changes the user's
provider default. The dialog states this before sign-in.
- The displayed access choice can allow all company agents to use the
account for its owner's tasks. Reconnect keeps the existing access.
Server permissions still control installs.
- Claude's typed `access` category maps to the provider's
`auth_required` signal. Tool and workspace request failures retain their
existing classification.
- No database migration or provider credential format changes are
required.

## Model Used

- OpenAI GPT-6 through Codex. The exact served model identifier and
context window are not exposed in this session. Capabilities used:
reasoning, repository tools, code editing, and command execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:57:52 -05:00
DottaandPaperclip 6d654f63d1 feat(apps): make MCP action test results readable (#14859)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connected Apps let an operator control which MCP actions an agent
can use.
> - The Permissions page lets the operator run a real action as an
agent.
> - The Test dialog displayed the nested MCP response as escaped JSON.
> - A useful result was hard to read, even when the action worked.
> - This pull request renders known MCP content as a readable preview
and keeps the raw response available.
> - The benefit is faster validation without losing the data needed to
diagnose a failure.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The per-action Test dialog on a connection's Permissions page.

**Subsystem affected**

ui/ — React board UI.

**Current behavior**

The dialog shows the gateway response as an escaped JSON blob. Text
content that contains JSON stays inside a string. The obsolete
connection Test page also keeps a separate set of stories.

**Proposed behavior**

The dialog uses structured MCP content when present. It parses JSON text
blocks when possible. It shows compact tables, cards, fields, or plain
text. It keeps the full raw response behind a control and opens that
view for errors or unknown block shapes. Stories exercise the
Permissions page dialog, and the obsolete Test page and stories are
removed.

**Reason and benefit**

An operator can inspect a successful action result at a glance and still
inspect the exact gateway response when a call fails or looks wrong.

**Breaking changes**

No API or stored data changes. The Test dialog presentation changes. The
raw response stays available.

**Additional context**

I tested a read-only Notion search through the real Permissions page.
The dialog showed three result cards and the raw response control
worked. Storybook uses invented example data.

No directly matching public issue or open PR was found in the GitHub
search.

## What Changed

- Render structured MCP output and JSON text content in the action Test
dialog.
- Show wide rows as cards, keep short rows as tables, and retain the raw
response for diagnosis.
- Remove the obsolete connection Test page and its stories.
- Add focused dialog tests and Permissions page Storybook cases for
success, errors, mixed blocks, and malformed blocks.
- Document the Test dialog result behavior in the connection playbook.
- Keep agent mention icons visible when the Lucide icon node is
unavailable in server rendering, which repaired a repeatable CI failure.

## Verification

- `pnpm -r typecheck` — passed.
- `pnpm exec vitest run --project @paperclipai/ui` — passed (7,111
tests).
- `pnpm exec vitest run
ui/src/pages/apps/app-detail/ActionTestDialog.test.tsx` — passed (11
tests).
- `pnpm exec vitest run --project @paperclipai/ui
ui/src/components/MarkdownBody.test.tsx` — passed (53 tests).
- `pnpm test:run` — started, then stopped after the review fixes changed
the head; the full sharded suite passed in CI.
- `pnpm build` — passed.
- `pnpm check:token-gates` — passed.
- Use a connected MCP app. Open Permissions, select a read action, and
run Test. Inspect the preview and the raw response control.

## Risks

- MCP tools can return provider-specific block shapes. Unknown blocks
open the raw response so the operator can inspect the exact result.
- Row and field previews limit visible data. The raw response preserves
the complete result.

> This is a targeted improvement to the existing Connected Apps item in
`ROADMAP.md`.

## Model Used

OpenAI Codex, GPT-6. The session used tool access, code execution, and
browser validation. The exact deployment ID and context window were not
exposed to the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 11:52:26 -05:00
DottaandPaperclip 4ac374103f fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.

## Linked Issues or Issue Description

Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.

**What happened?**

Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.

**Expected behavior**

Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.

**Steps to reproduce**

1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.

**Paperclip version or commit**

Reproduced from b54b2dc35c. Rebased onto
master at `0829d94af` after the single-screen setup change in #14811.

**Deployment mode**

Local development from source. The managed path also supports enrolled
self-hosted instances.

## What Changed

- Add the `asana.mcp` managed profile and the default Sign in with Asana
method.
- Use Asana's reviewed v2 protected-resource metadata before cached
endpoints.
- Require a custom MCP app's client secret and retain saved credentials
during setup or reconnect. Repair only the known Asana v1
issuer/resource binding, retaining company and callback checks.
- Expose a boolean for the acting user's saved client secret. The form
offers secret reuse only when that user has an active grant with the
required reference.
- Let users select their own app from the enrollment and
shared-app-unavailable screens, or from Advanced on the single-screen
setup page.
- Canonicalize HTTP loopback callbacks even when the request includes an
Origin header.
- Extend signed broker claims and provider URL validation for Asana.
Require refresh credentials on managed authorization.
- Document setup, distribution, and shared-app rollout requirements.
- Resolve permission-profile name collisions when finishing another
account. The live staging test found this after renaming the first Asana
connection; OAuth succeeded but profile finalization failed.

## Verification

- Final live staging proof used app commit
`c6053157c4e42ac017727117b754ae77fa5c45fa` and the real Cloud broker at
`767b63835170f664542afd0df99a76615e204b62`. In the embedded browser,
default shared sign-in required no client credentials, returned through
the central Cloud callback to the tenant, and discovered 39 actions. Get
me succeeded through the gateway as the selected QA agent (2.1 seconds).
The custom-app connection also returned a real result on this final
build (0.9 seconds).
- Retried the shared draft that failed during the first staging test. It
completed after the profile-name fix, retained the selected agent, and
kept the existing custom connection intact. Two database regressions
reproduced the collision before the fix and passed afterward. The
updated transaction rollback test also passes.
- Shared reconnect returned to the same staging connection with 39
actions. Earlier staging checks verified the custom-app fallback when
the shared profile was unavailable, saved-secret reuse on reconnect, and
Off blocking the action test. Allowed was restored after that check.
- Local live-provider checks also repaired a saved Asana v1
issuer/resource binding without reentering the secret and verified that
numeric-loopback setup uses the displayed localhost callback. Expiring
the local managed access-token timestamp triggered a real Asana refresh
and a successful Get me call. These early local broker tests used
enrollment/authentication and storage fixtures; the final staging proof
used deployed Cloud identity and persistent storage.
- The new production app is registered and configured, but production
sign-in has not been deployed or verified. Live provider revocation was
not run because the existing staging test app is shared with other
connections.

- After rebasing onto the single-screen setup flow, full `pnpm -r
typecheck`, `pnpm build`, and `pnpm check:token-gates` pass. Focused
verification passes 365 service and 36 broker-client tests. Broader
checks pass all 833 shared-package tests and all 530 connector-page
tests. The shared suite uses `TMPDIR=/private/tmp` to avoid macOS
temporary-directory symlinks in its canonical-path tests. The UI tests
verify the shared-app default and switching to a custom app with its
required client secret.
- Embedded-browser smoke on the current rebased build verified the
shared sign-in default, Advanced → custom app (client ID and secret
required), and switching back to Paperclip. Both existing Asana
connections remained connected after restart. No new provider
authorization was performed during this smoke.
- The full local `pnpm test:run` was interrupted when the execution
session restarted. Before interruption, it reported one runtime-slot
restart test failure. That test passed on an isolated retry after
clearing two unused PostgreSQL shared-memory segments. The full local
suite did not complete; CI must pass on the current head before merge.
- All 52 CI and security checks pass on
`9318fd4e9b616cdc3de12f40cdb9bd32d865af4c` (CI run `36882064080`),
including all eight browser shards, nine serialized-server shards,
build, typecheck, and canary dry run. Two optional Storybook checks were
skipped. Greptile review 4 reports 5/5 on this exact commit, with all
review threads resolved. Its updated summary identifies the current SHA;
this comment-triggered review did not publish a separate GitHub check
run.
- Provider revocation is unit-tested in the companion broker. Live
provider revocation was not run because the existing test app is shared
with other connections.

## Risks

- The shared Paperclip Asana MCP app has been registered with its
production callback and Any workspace distribution. Its secret is
provisioned in the production secret store, and the runtime client ID
and secret reference are configured. The production profile is enabled
in the saved deployment configuration. The companion broker has merged
and passed staging deployment; production sign-in still requires a
production deployment and live verification. Custom setup remains
available.
- Asana MCP uses the provider's fixed `default` grant. Paperclip action
policies limit agent tool use; they do not narrow provider consent.
- The callback correction affects HTTP loopback OAuth flows. Public
HTTPS callbacks retain their existing behavior.
- Reviewed discovery URLs now override stale cached endpoints. Tests
cover the Asana v1-to-v2 repair.
- No schema migration. Connection removal retains the existing
local-only revocation behavior.

## Model Used

OpenAI GPT-6 through Codex, with code execution, browser testing, and
GitHub tooling. The exact model variant and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:24:44 -05:00
467125fafb feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen

## Linked Issues or Issue Description

No public issue exists. This is the description, from the enhancement
template.

**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.

**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.

**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.

**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.

**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.

**Breaking changes**
None. No schema or API change. Existing connections keep their settings.

## What Changed

- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.

## Verification

- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
  - `action-permissions.test.ts` checks the group update.
  - `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".

## Risks

- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:13:47 -07:00
DottaandPaperclip c8f874311c fix(ui): hide Google connectors only on the Connections page (#14774)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connections page lists the apps and saved accounts that agents
can use.
> - Google Workspace verification is still pending.
> - Google entries must be temporarily hidden from this page without
removing their implementations.
> - This PR filters the final page rows, including saved Google
accounts, after the page resolves their provider.
> - Definitions, direct setup routes, OAuth profiles, credentials, and
runtime access stay intact.
> - Review instances can keep the prior UI by staying on their pinned
app release.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Temporary provider visibility on the Connections landing page.

**Current behavior**

The page can show Google Workspace catalog entries and saved accounts
while verification is pending.

**Proposed behavior**

Hide all nine Google Workspace rows on this page. Keep every other
connector and all Google integration code unchanged. Use an existing
release pin for review instances instead of a hostname exception in the
app.

**Reason and benefit**

Pause public discovery without disabling existing runtime tools or
removing the implementation needed for verification and later
re-enablement.

**Breaking changes**

Google accounts are no longer visible on this landing page. Direct setup
and management routes remain available. This is not an access-control
restriction.

Related completed work: #13551 used catalog-level visibility. This
change is deliberately limited to the landing page and also covers saved
account rows. #14740 reduced Google scopes; this change leaves those
scopes unchanged. No duplicate open PR or matching open issue was found.

## What Changed

- Derive the Google app slugs from the existing Workspace profile
registry.
- Filter the combined catalog and saved-account rows only inside
`Browse`.
- Cover all nine Google entries, active/draft/disabled accounts, legacy
connection metadata, mixed-provider rows, and independently identified
non-Google connectors in regression tests.
- Document the display-only hold, pinned review builds, and how to
restore visibility after approval.

## Verification

- Passed: `pnpm exec vitest run ui/src/pages/apps/Browse.test.tsx
ui/src/pages/apps/AppsConnect.test.tsx` (199 tests, including the latest
master changes).
- Passed: `pnpm check:token-gates`.
- Passed: `pnpm build`.
- Passed: `pnpm -r typecheck` and `pnpm build` after merging the latest
master. An earlier overlapping run hit a local runner codesign race;
sequential checks passed.
- Passed again after the final custom-provider fix: `pnpm --filter
@paperclipai/ui typecheck` and `pnpm --filter @paperclipai/ui build`.
- The full local `pnpm test:run` was started, then stopped after the
full remote CI suite passed to avoid continuing duplicate long-running
work on the developer machine. It is not claimed as a completed local
pass.
- All 54 latest-head CI checks passed. Two non-applicable Storybook jobs
were skipped. One serialized server job lost its self-hosted runner
connection; its single retry passed.
- Greptile: 5/5 on `aeda167bf4494feed6ee0de2585960511fb02918`, with no
unresolved review threads.
- Confirmed in the existing review instance that all nine Google entries
still appear after its current release was pinned. No new app release
was deployed to that instance.
- Reviewer steps: open Connections on this branch with Google catalog
entries and saved Google accounts. None should appear. Non-Google
connectors must remain. Direct Google setup routes must still load.

## Risks

- Existing Google accounts cannot be found on this page during the hold.
Their data and runtime access remain unchanged.
- This is a UI-only filter, not an authorization gate. Direct routes and
API access still work by design.
- Review instances must not receive this UI build until the hold is
removed. Their existing release pin excludes fleet app upgrades; an
explicit targeted upgrade must still be avoided.
- No migrations, backend changes, broker changes, or credential changes.

## Model Used

OpenAI Codex (GPT-5-based coding agent), with reasoning, tool use, code
execution, and browser inspection. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:22:25 -05:00
DottaandPaperclip f7e36ba3e2 fix: isolate repository-free low-trust tasks in private directories (#14766)
## Thinking Path

> - Paperclip manages work by agents within company boundaries.
> - Email tasks can run under the low-trust review preset.
> - These tasks must use an isolated workspace and a sandbox.
> - The default workspace strategy assumed that the project had a Git
repository.
> - A project without a configured workspace failed before the agent
could start.
> - This change gives each such task a private directory and keeps the
sandbox requirement.

## Linked Issues or Issue Description

**What happened?**
An inbound email assigned to a low-trust agent failed with
`git_worktree_base_not_git_checkout` when its boundary project had no
configured workspace. Setup had accepted the project and sandbox.

**Expected behavior**
The agent can process email without a repository. Its workspace stays
isolated from other tasks and the shared agent home.

**Steps to reproduce**
1. Select a low-trust agent with an active sandbox and a project
boundary.
2. Leave the project without a configured workspace.
3. Receive an email through AgentMail.
4. Observe that startup fails before provider work starts.

**Paperclip version or commit**
Reproduced against `5edf55d73`.

**Deployment mode**
Hosted staging with sandbox execution.

Related: #13256 added email tasks. #13636 fixed default isolation for
projects without workspaces; the explicit isolation used by low-trust
tasks still needed this path.

## What Changed

- Select private task directories for low-trust sandbox tasks with no
configured workspace or explicit workspace strategy.
- Keep each directory scoped to its company and task. Retain files
across turns and reassignment and reject symlink paths and mismatched
workspace reuse.
- Preserve Git validation for configured workspaces and explicit
strategies, plus the existing authorization and remote gates for
referenced projects.
- Add a startup regression and directory isolation tests. Document the
supported repository-free path.

## Verification

- The startup regression failed before the fix with the same Git
validation error.
- 260 targeted email, workspace policy, heartbeat, referenced-project
and directory tests pass.
- Full `pnpm -r typecheck` and `pnpm build` pass on the latest commit.
- All CI checks, including the complete sharded test suite and canary
dry run, pass on `b4ccd9802b09b2e95499df72d48b4a3906b8c328`.
- The final commit also passes the same server shard locally: 60 files,
1,024 passed / 6 skipped tests. The earlier all-groups local run was
interrupted during follow-up edits; complete-suite verification comes
from CI on the final commit.
- Deployed the reviewed commit to staging and independently verified the
full serving SHA. Two real Codex runs in Daytona succeeded and finalized
the same private company/task workspace. The first wrote a 35-byte
marker; the second read the existing file without modifying it and
returned the independently verified SHA-256
`ce3bbeb44d07ca6822826d3a5945752a38d30b356d10829f3159a191e5aa92a6`.
- Live runtime caveat: Codex reported a nested `bwrap` loopback
permission error and used its configured escalated execution inside
Daytona. The outer Daytona sandbox remained active for both runs.
- The startup regression uses a real database, production trust checks,
workspace persistence, sandbox lease acquisition and realization, and a
fake provider. It checks reassignment and allows only an authorized
referenced project.
- The transfer regression runs production archive/sync-back/merge code
against distinct filesystem roots: create output in one sandbox, restore
it, then read and update it in a fresh sandbox. The provider I/O is
emulated; live staging verification is separate.

## Risks

- The new default applies only to low-trust sandbox tasks without
workspace configuration. Standard agents and explicit Git strategies
keep their existing behavior.
- Task directories retain work across turns and consume instance
storage. The change does not migrate or copy existing shared files.
- No database migration or credential changes are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact runtime model revision and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:06:07 -05:00
Devin FoleyandPaperclip 3bbb8d0f69 Add bounded AgentMail failure diagnostics (#14768)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - AgentMail connections create inboxes and handle email tasks.
> - A failed provider request currently records only its HTTP status.
> - The same 403 can mean a permission denial, a resource limit, or
another provider restriction.
> - This pull request records a fixed operation name and a documented,
allowlisted error code.
> - Operators can distinguish these failures without exposing provider
payloads or changing retry behavior.

## Linked Issues or Issue Description

**What happened?**

An AgentMail inbox creation failure reports only `AgentMail request
failed (403)`. The response body is deliberately excluded because it can
contain private mail or credentials. That also discards the provider
code needed to identify the cause.

**Expected behavior**

Keep the HTTP failure visible with a fixed operation name and a safe
provider code. Never copy arbitrary error text, resource identifiers,
suggested fixes, or URLs into diagnostics.

**Steps to reproduce**

1. Make an inbox creation request through `agentmailApi` with a fake
provider returning HTTP 403 and `code: "missing_permission"`.
2. Observe that the old error lacks the operation and provider code.
3. With this change, verify the error includes `operation=create_inbox,
code=missing_permission`, preserves status 403, and excludes all other
response fields.

Related work: #13256 introduced the AgentMail connection. The provider
documents stable codes in its [error
reference](https://docs.agentmail.to/errors).

## What Changed

- Add a fixed method/route-to-operation map and an allowlist of
documented provider codes.
- Read at most 8 KiB for diagnostics, with a one-second deadline. Cancel
unread bodies and preserve the HTTP error if reading or parsing fails.
- Keep the existing error prefix, status, retry delay, and failure
handling.
- Add regression coverage and document the diagnostic limits.

## Verification

- `pnpm exec vitest run server/src/__tests__/agentmail-api.test.ts` — 44
tests passed.
- `pnpm build` — passed.
- `pnpm -r typecheck` — passed before the review correction. Final `pnpm
--filter @paperclipai/server exec tsc --noEmit` also passed.
- Full local `pnpm test:run` did not finish successfully; three
company-skills-service failures were observed outside the changed
module. The final-head CI server suites passed. The remaining
workspaces-b CI retry covers an unrelated HTTP/2 port collision.
- The diff passed a scan for configured secrets, private deployment
references, and non-fixture email addresses.

## Risks

- A failed request can now wait up to one extra second while reading its
diagnostic code.
- New, missing, malformed, or oversized provider codes report `unknown`.
A future provider code needs an explicit allowlist update.
- This is a diagnostics change. It does not establish or repair the
cause of an existing provider denial.
- No schema, credential policy, or retry behavior changes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 at 20853e31abacf53a32f1a63467ee208a719bbf00; the
locked-body finding is fixed and its thread resolved
- [x] I will address all Greptile and reviewer comments before
requesting merge


Final verification (September 30): all final-head GitHub checks pass at
`20853e31abacf53a32f1a63467ee208a719bbf00`, including the targeted
workspaces-b rerun after the unrelated EADDRINUSE failure. Greptile
scored 5/5 on this head and no review threads remain unresolved. The
branch is mergeable. The local full-suite run did not yield a passing
completion; CI completed successfully across all suites. This public PR
remains open for maintainer merge.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 15:29:15 -07:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip ad55d0a281 fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use Apps through a gateway that checks identity, company
access, and action policies.
> - Personal pasted credentials can point to company secrets. Setup can
show success while the gateway rejects every call.
> - Several OAuth methods also omit the scopes needed for their
supported write actions.
> - This pull request gives setup, health checks, and invocation the
same credential rules. Owners repair existing connections by
reconnecting.
> - New connections request reviewed permissions for their supported
actions. Read-only choices remain available under Advanced.
> - Agents can use the connections people give them, while existing
consent, identity boundaries, and action restrictions remain enforced.

## Linked Issues or Issue Description

Refs #14009 and #14008. This addresses the personal-credential defect.
The separate GitHub organization-identity selection defect is outside
this change.

Related work: #13942 fixed part of new personal-key setup. #14200
independently fixes legacy personal reconnect and protects managed-agent
profile credentials during removal. This PR covers that ownership
invariant across key and secret-URL setup, reconnect, health, discovery,
and invocation, and keeps owner reconnect as the repair path. #14059
tracks requested versus provider-asserted OAuth scopes; it remains
separate work. I searched open PRs and issues for Zapier, Airtable
scopes, connector writes, and personal credential failures.

**What happened?**

A Zapier secret URL saved through personal setup can become a company
secret referenced by a user grant. Health checks bypass the gateway's
ownership check, so the connection appears healthy but calls fail with
`grant_credential_invalid`. Custom-header paths can also receive a
duplicate `credentials.` prefix. Omitted OAuth scopes make write access
depend on provider defaults.

**Expected behavior**

Personal invocation credentials belong to the selected user. Setup,
health, and actual calls enforce the same rule. New connections request
documented permissions for supported read and write actions. Existing
tokens gain no permissions without provider consent.

**Steps to reproduce**

1. Connect Zapier or a generic secret URL with the personal identity.
2. Allow an agent to use the connection and complete setup.
3. Invoke a tool through a run-scoped gateway. The legacy layout fails
ownership validation despite successful setup.

**Paperclip version or commit**

The implementation started from
`44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master
at `94e8dec56`.

**Deployment mode**

Built from source. Regression tests use isolated PostgreSQL fixtures and
controlled MCP transports.

## What Changed

- Share credential writing, ownership validation, and canonical paths
across initial setup, resume, reconnect, rotation, health, discovery,
and gateway calls. Keep OAuth client-registration secrets separate from
invocation credentials.
- Existing personal connections with company-scoped credentials require
owner reconnect with a fresh key or secret URL. Reconnect creates a
correctly owned value and updates the existing grant and declarations.
There is no automatic ownership backfill or new startup hook.
- Preserve PostgreSQL timestamp precision when reconnect checks whether
a grant changed. Previously, converting the timestamp to a JavaScript
Date could reject reconnect with a false concurrent-change error.
- Protect credentials used by other grants, connections, bindings,
managed-agent profiles, routine triggers, or secret proposals from
connection removal.
- Review all 117 tool methods, including 84 OAuth methods. Record
explicit scopes or documented provider-default exceptions with official
evidence. Add Airtable's seven scopes, Hugging Face repository/job
scopes, and other documented MCP permissions.
- Prefer available write-capable methods. Put explicit read-only choices
under Advanced. Explain pasted-key permissions and offer reconnect for
missing OAuth consent. Preserve existing grants, policies, Google
availability gates, and curated scope allowlists.
- Reconnect generic secret URLs and custom headers using their stored
credential fields. Refresh the catalog after setup, correct reconnect
feedback and error guidance, and let Cancel exit invalid setup while
Save & exit retains draft-saving behavior.
- Apply ownership checks to the new GitHub repository/skill connection
picker. Align the permission audit with the Google scope reductions
merged on master.
- Add run-scoped gateway, ownership, owner-reconnect, OAuth URL,
insufficient-scope, UI, and catalog-wide regression coverage. Update the
connector playbook and permission audit.

## Verification

Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all
CI/status gates (55 completed check runs, no failures or pending checks)
and has a completed Greptile review at **5/5 with no outstanding
findings**. GitHub reports the PR as mergeable/CLEAN.

- **Embedded browser:** used the actual server and built UI from this
worktree, a fresh isolated database, and local HTTP MCP fixtures.
Completed personal bearer-key, secret-URL, and custom-header setup;
reproduced the legacy ownership failure; reconnected through the owner’s
form; and completed writes afterward. Read-back was verified for
bearer-key and secret-URL connections. Public organization-wide setup
appeared immediately in Browse without reload. Zapier URL
validation/Cancel and Google’s enrollment gate were also exercised.
- **Persistence and invocation:** verified user ownership, canonical
`credentials.authorization` / `remote.url` / `headers.X-Api-Key`
declarations, and unchanged connection/grant identity. The old company
secrets retain their ownership. Separate HTTP calls through an actual
run-scoped gateway session completed a write and read-back.
- **Backend coverage:** the final gateway suite passes all 82 cases,
including catalog Zapier and generic inline reconnect. It checks
company/user isolation, canonical declarations, same-endpoint URL
validation, fresh credentials, retained restrictions, and real gateway
read/write execution using fixture transport. A timestamp with
PostgreSQL microseconds covers the former false reconnect conflict.
- **Local checks:** 368 catalog, gateway, repository, and UI tests
passed before the final extra Zapier case; 49 GitHub skill access tests
also passed. All three Apps browser regressions pass, including
reconnect through the actual form and catalog visibility without reload.
Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final
patch, and token gates passed. Full tool-access service runs hit varying
15-second Google fixture timeouts; both affected cases and the updated
reconnect assertion pass in isolation (3 tests). The complete test
matrix passes in CI on this head.
- **Verification limits:** no live provider account was available for
Zapier/Airtable/OAuth consent or account-bound write proof. Public
metadata and local fixtures do not establish provider consent. The
original development database clone failed on a pre-existing missing
`tool_connections_transport_check` constraint; browser acceptance used a
fresh isolated database created by the normal CLI onboarding flow.

## Risks

- Existing broken personal connections stay unusable until their owner
reconnects. Health, discovery, and invocation return an actionable
ownership error; startup does not rewrite credential ownership.
- Scope changes affect new authorization requests. Providers may still
require resource selection, account roles, paid plans, or app
verification. Existing consent and action restrictions remain unchanged.
- Shared credentials are retained rather than reassigned or revoked.
Provider-default exceptions and unavailable live checks are documented
in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`.
- No new endpoint, database table, lockfile change, or CI workflow
change is included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell
execution, web research, and browser tools. The exact deployment model
ID and context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:32:34 -05:00
DottaandPaperclip 25c422ba7e fix(apps): request minimal Google service scopes (#14740)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Google connections give agents service-specific tools through
governed credentials.
> - Each connection has a reviewed OAuth scope set.
> - Docs, Sheets, and Slides request Drive permissions in addition to
their own service scopes.
> - Google documents these permissions as alternatives, not combined
requirements.
> - This pull request removes those extra permissions and redundant
Calendar write free/busy access.
> - Users grant fewer permissions without adding tools or changing
connection access policy.

## Linked Issues or Issue Description

Refs #13820. Related #14739 changes other connector permissions; it does
not reduce these Google profiles.

Companion broker PR:
https://github.com/paperclipai/paperclip-cloud/pull/615. Ship the
matching changes together after fresh-grant validation.

**What happened?**

Seven Google profiles request redundant scopes. Docs, Sheets, and Slides
request Drive scopes. Calendar write requests free/busy even though
calendar.events authorizes its availability tool.

**Expected behavior**

Each profile requests only the scopes required for its reviewed tools.
Managed and customer-owned OAuth methods use the same set.

**Steps to reproduce**

Inspect the Google profile registry and the four app definitions on the
base commit. Compare their scope sets with Google's MCP authorization
alternatives linked in the updated documentation.

**Paperclip version or commit**

Base: 94e8dec56b.

**Deployment mode**

Hosted and self-hosted Google Workspace connections.

## What Changed

- Docs, Sheets, and Slides read profiles request only their service
read-only scope. Write profiles request only their service write scope.
- Calendar write keeps calendar-list read-only and event access.
Calendar read keeps free/busy.
- Apply these sets to managed and customer-owned OAuth methods and their
exact-scope tests.
- Document scope rationale, old-grant behavior, and the coordinated
broker rollout.
- Preserve the 21-scope integration union. Drive and Workspace Search
still need their own Drive permissions.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts`: 25
passed.
- `pnpm -r typecheck`: passed with repository-pinned pnpm 9.15.4 after
refreshing locked dependencies.
- `pnpm build`: passed with repository-pinned pnpm 9.15.4.
- `pnpm test:run`: attempted on the machine's Node 26 runtime, then
interrupted after unrelated suite/collection failures. This is NOT a
passing full-local-suite result. CI uses Node 24.
- Node 24 focused diagnostic rerun: `workspace-runtime-exposure.test.ts`
and `paperclip-control-plane-port.test.ts`: 43 passed, 3 skipped. The
first suite hit a local preview-readiness timeout in CI; the failed
shard passed its single retry without code changes.
- Node 24 rerun of the three local assertion-failure suites
(`chat-discord-adapter-patch`, `ai-connections`, and
`heartbeat-active-run-output-watchdog`): 133 passed. These files were
not changed.
- Node 24 rerun of the changed manifest contract suite: 25 passed.
- Latest-head GitHub checks are green at
`da85dcbc89c44d88f3cabd5ada142c922fa51a3c`: 54 passed, 2 intentionally
skipped; no failed or pending checks. Includes typecheck, build,
general/serialized tests, all 8 browser E2E shards, Runner verification,
and canary packaging. Greptile 5/5, no review threads.
- Cross-repository comparison: all 16 app and broker profiles match
exactly.
- `git diff --check`: passed.
- Google definitions are reviewed JSON source inputs preserved by the
ingestion script. The generated TypeScript registry imports them; no
generic provider regeneration is needed.
- Fresh minimal-grant provider testing is outstanding. No new video was
recorded, and existing broader credentials are not proof of minimal
authorization.

## Risks

The Cloud broker must ship the matching exact scope sets. Mixed versions
fail closed. Validate fresh grants in staging before production.
Existing provider tokens are not retroactively narrowed; affected
managed connections must reconnect if their grants retain extra scopes
or omit explicit scope evidence at refresh. Do not revoke the shared
Google project to migrate one connection. This change does not grant
access, add tools, or alter the Google Chat unread-filter block.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, code execution,
and browser tools. The exact runtime model ID and context-window size
were not exposed to the agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:10:42 -05:00
DottaandPaperclip a36cbffa9e fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path

> - Paperclip manages agents and the services they need for work.
> - Agents use connection search to discover a setup path before they
request access.
> - Tool-only filtering hid channel and AI methods from this search.
> - Requiring every query word to match rejected normal task
descriptions.
> - A small aggregator index also omitted supported apps such as
Circleback.
> - This pull request broadens retrieval and returns purpose-specific
setup guidance.
> - Agents can choose a relevant result while existing access and
provider-choice checks still apply.

## Linked Issues or Issue Description

**What happened?**

A search such as “AgentMail create an email address and manage an agent
mailbox” returned no usable result. “Help me find tools for circle back”
also missed Composio's supported Circleback toolkit. Queries longer than
200 characters failed validation.

**Expected behavior**

Return useful native and verified aggregator matches from
natural-language queries. Include channel/email methods when Chat
connectors is enabled. Identify each method's purpose and the correct
setup path.

**Steps to reproduce**

Use the queries above with `connections_search` from an active task.
Enable Chat connectors for the AgentMail case. The regression suite
reproduces these misses before the change.

Related routing work: #13941. This change does not change the runner
failure path or add channel setup to tool-only connection cards.

## What Changed

- Rank name and capability matches. Accept extra words, split names,
small spelling errors, and queries up to 4,000 characters.
- Include tool, channel/email, and AI methods. Return company-prefix
setup links for channel and AI flows.
- Add a dated snapshot of 1,583 official Composio toolkit names and a
refresh script. Merge duplicate MCP variants for search and link each
support claim to official evidence.
- Find authorized indexed aggregator namespaces within longer queries.
Return multiple app matches when the agent needs to choose.
- Prefer exact app names over fuzzy matches for other apps; retain
existing AI readiness.
- Preserve native preference, company and identity boundaries,
administrative denials, and saved provider consent.
- Add relevance and database regressions, extend native tool-authority
coverage, and document search behavior.

## Verification

- Red: 15 new assertions failed against the previous implementation; the
existing baseline passed. Added red-green regressions for Motion versus
fuzzy Notion and existing AI access during review. A further regression
covers mixed ready/unconfigured AI results and their per-result setup
guidance.
- Green: all 103 tests in the eight focused shared, database,
runtime-tool, fixture, and route suites pass on the latest commit.
- `pnpm -r typecheck` and `pnpm build` passed.
- Latest-commit CI passed: 54 successful checks and two skipped checks,
including the full test matrix, browser E2E, typecheck, build, and
canary dry run.
- The long local `pnpm test:run` attempt began before the review fixes
and retained transformed pre-fix search code; it also hit an unrelated
timing failure. Fresh serial reruns of the affected search suites and
three timeout cases passed all 124 tests. Parallel local route shards
hit two additional database setup timeouts; both suites passed all 17
tests on a fresh serial rerun. The complete corresponding CI suites also
passed. Duplicate broad local runs were stopped after CI completed. The
local UI suite independently passed all 7,007 tests.
- Greptile: 5/5 on `cc6a0180d`; all review findings resolved.
- Browser verification passed in a disposable local instance through
real process-agent search requests: AgentMail opened its setup flow with
the requester selected; the saved Circleback choice produced the
Composio setup card; a paragraph-length Notion query produced its setup
card. No provider credentials or external accounts were created.
- The browser test caught an invalid UUID-based setup URL. The fix uses
the company prefix and has a regression assertion.
- A 3,971-character catalog query found Circleback first in a local 10
ms spot check after sharing query preparation across the catalog scan.
This is a single measurement, not a performance guarantee.

## Risks

- Broader retrieval can return extra candidates. Named services rank
first; agents must select the relevant method.
- The public support snapshot can age. It proves catalog support, not
account authorization or the availability of every requested action.
- Channel and AI methods use existing setup links. The tool connection
card still accepts tool methods only.
- No schema, migration, credential, or runner lifecycle changes.

## Model Used

OpenAI Codex (GPT-6). The session does not expose a more specific model
identifier or context-window size. Used reasoning, repository search,
code execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:20:50 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 2f6fa3b6dc fix: recover provider authentication inside tasks (#14629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a working model provider connection to run a task.
> - A provider can reject a stored credential after the task starts.
> - The failed run must ask the responsible user to repair that
connection.
> - This pull request adds that request directly to the task and reuses
Connections sign-in.
> - The user can choose an API key or subscription, then continue the
task with a fresh session.

## Linked Issues or Issue Description

**What happened?**

A run that ended with `acpx_auth_required` or another known provider
authentication error did not immediately offer an inline way to connect
the provider. A repair form could also lock the user to the failed
account's sign-in method.

**Expected behavior**

Show a provider connection card in the task as soon as the
authentication failure is saved. Allow the responsible user to connect
or repair the provider with any supported sign-in method. Keep the
connection name automatic and resume the task after successful setup.

**Steps to reproduce**

1. Run a task with a supported provider and an expired or invalid
credential.
2. Let the run fail with a provider authentication error.
3. Open the task and attempt to repair the connection.

Related work: Refs #13724 and #13726. This change adds the inline task
repair flow and method choice.

## What Changed

- Classify provider authentication failures and create one connection
request for the current task. A persisted blocked classification
suppresses automatic retries only after the repair card is created;
unsupported providers retain their existing recovery path.
- Mark only the attributed, unchanged credential as needing sign-in.
Preserve credentials that were refreshed after the failed run started.
- Reuse the provider sign-in controls inside the task. Allow API key and
subscription choices for Claude, Codex, and Grok. Keep names hidden and
generate a default from the user, provider, and method.
- Keep the existing account when reconnecting with the same method.
Create and select another account when the method changes. Validate
updates to explicit agent bindings through the normal agent save path.
- Require explicit adoption for legacy agent authentication. Validate in
the agent environment, then commit the binding, connection install,
audit, and card completion in one transaction. Keep failed setup and
account selection visible and retryable.
- Add regression tests and update the specification and Connections
documentation.

## Verification

- Fresh local verification: 199 tests passed across the inline form,
provider method selector, default naming, authentication and recovery
classifiers, run liveness, OpenAPI routes, database adoption/rollback,
and Cursor execution suites. The adoption database suite also passed
against disposable Docker PostgreSQL.
- Full repository `pnpm build` and `pnpm -r typecheck` passed on the
latest commit. Token gates are clean.
- Embedded browser: opened real task cards from seeded authentication
failures; switched Claude from API key to subscription and back;
switched Codex from subscription to API key; confirmed the name field
stays hidden. Provider sign-in was not completed with real credentials.
- The broad local `pnpm test:run` started before review fixes and was
interrupted after the working tree changed; it is not counted as a
passing full run. Fresh focused tests passed. CI supplies the full test
and browser suite results for the current commit.
- CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54
checks passed and two Storybook checks were skipped by their path rules.
The workspace preview job passed on one rerun after a local-server
startup timeout; its rerun passed 835 tests.
- Greptile is 5/5 on the same commit with no actionable findings and no
unresolved review threads.

## Risks

- Incorrect authentication classification could prompt for a connection
unnecessarily. Tests exclude tool authorization, quota, and unrelated
runtime failures.
- A method change selects the new personal provider default, which also
applies to other agents that use that user's default. Explicit account
bindings use the existing permission and runtime validation path.
- Credential invalidation must not race with refresh or reconnect. The
code compares the saved credential generation and grant update time
under locks.
- No database migration or new credential storage format is required.

## Model Used

OpenAI GPT-6 through Codex. The exact model ID and context window size
were not exposed in this session. Capabilities used: reasoning,
repository editing, shell commands, database tests, and embedded-browser
interaction.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests listed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 06:39:37 -05:00
DottaandPaperclip 2de43fc909 fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Each task has one assignee. Explicit assignment and review requests
select who should act.
> - An agent mention started another agent on a task it did not own.
Native attachment staging then rejected that run.
> - Allowing that run through startup could also let two agents work on
the same task.
> - Mentions should identify relevant context. They should not start
work or forward comments to other tasks.
> - A personal app installed on a shared agent must also wait until tool
use to resolve the current user's grant.
> - This pull request removes mention dispatch and keeps missing
personal app credentials from blocking startup.

## Linked Issues or Issue Description

**What happened?**

A native agent mentioned on another agent's task failed with
`paperclip_runner_attachment_staging_not_authorized`. The source task
could already be complete. A nearby optional-app warning was a separate
problem: personal app tools were excluded when their shared health state
required attention.

**Expected behavior**

An agent mention is context only. It does not wake the agent, take
ownership, or copy a comment onto another task. Normal feedback still
reaches the assignee. Assignment and explicit review requests still
dispatch work. An unavailable personal app does not block startup or
produce a startup warning. Tool use requests the current user's
authorization and never uses another user's grant.

**Steps to reproduce**

1. Assign a task to agent A. Post a comment that mentions agent B,
including a comment that closes A's task or references B's child task.
2. Confirm the comment retains its agent link and B receives no run or
deferred wake. A can still receive normal feedback.
3. Install an active personal MCP connection on B. Give only Alice a
grant and leave shared health at `error`.
4. Explicitly assign work to B for another user. Confirm it can finish
without using the app.
5. Ask B to use the app. Confirm its tool call shows an inline
connection request for the current user.

Related work: Refs #11144. This change uses the existing execution-time
personal grant resolution.

## What Changed

- Remove mention dispatch from standalone comments and issue updates.
Remove implicit forwarding of parent comments to a mentioned worker's
child task.
- Ignore new requests with the legacy mention wake reason before
creating a run or deferred request. Preserve already accepted queue
entries, which can combine assignments and feedback with a later
mention.
- Remove the native mention admission, staging, and finalization
exceptions from this PR. Native task ownership checks remain intact.
- Keep active, installed personal app tools available despite shared
health errors. Remove optional-app startup warnings. Tool execution
retains the current user's grant and policy checks.
- Update agent instructions and product/API docs. Refresh generated
capability source anchors.

## Verification

- Red: comment-route regressions reproduced extra agent wakes and child
comment forwarding. A separate regression proved that cancelling by the
last coalesced reason could drop an accepted assignment.
- Green: the targeted route, wake queue, heartbeat, workspace,
responsible-user, MCP discovery, and HTTP gateway suites passed. The
final queue and heartbeat rerun passed 104 tests, the restored queue
adapter passed 56, and both comment-route suites passed 135. These
include accepted assignment preservation, rejection of new mention
requests, and normal assignee feedback.
- `pnpm -r typecheck` and `pnpm build` passed locally. The full local
`pnpm test:run` attempt was interrupted for review/CI fixes, so it is
not claimed as a completed local pass. It exposed a cleanup timing race
in the concurrent-mention assertion, now fixed and verified across 10
repetitions. CI also exposed an obsolete test waiting for the removed
mention lookup; it was reproduced and fixed, then both comment suites
passed. Final full-suite verification is through CI.
- Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks
passed, 2 Storybook checks intentionally skipped; no pending or failing
checks. Full CI includes general and serialized suites, all 8 browser
shards, runner verification, typecheck, build, and canary dry run.
Greptile is 5/5 on this exact commit, with no unresolved findings.
- One unchanged Cursor adapter test hit its 10-second CI timeout. All 5
tests in that file passed locally; one retry of its CI shard passed all
674 tests (3 skipped). The aggregate verification gate then passed. No
code or timeout was changed for that retry.
- Live browser check: inserted a structured mention with the picker on a
human-owned task. The saved link remained visible. Database checks found
zero new runs and zero wake requests.
- Live Codex runner check: explicitly assigned that task with the
unavailable personal app attached. The run succeeded and committed
completion without using the app or creating a connection card.
- Live browser follow-up: asked the assignee to call PostHog and
mentioned another enabled agent as context. Only the assignee ran. It
succeeded and displayed the existing inline connection card. Only
Alice's grant existed; the run belonged to a different user.
- The HTTP regression covers tool discovery with no provider calls or
connection cards, first use returning the current user's authorization
request, and successful retry after that user's grant exists.
- App checks use an isolated local fixture and a fake MCP provider. They
do not use production app credentials.

## Risks

- Intentional behavior change: workflows that used mentions to wake
agents must use assignment, a bounded child task, or an explicit review
request.
- Already accepted queue entries retain their prior rules. An old entry
can combine assignment or feedback with a later mention; its last reason
cannot safely identify mention-only work. New mention requests create no
run or deferred wake.
- Personal apps with a shared health error remain discoverable. Actual
tool use still requires the responsible user's grant and existing policy
gates.
- No database migration or public API schema change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact serving model ID and
context-window size are not exposed in this session.
- Live native-run verification used `gpt-6-astra` through the Codex
provider.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:49:01 -05:00
Devin FoleyandPaperclip 1e3a148f46 fix(connections): retain safe broker rejection diagnostics (#14098)
Preserve fixed broker rejection reason codes within strict size/time limits while retaining public error codes and status. Unknown bodies remain generic; no raw response or credential material enters the error.

Validated by 33 consumer tests, a synthetic producer HTTP contract fixture, root typecheck/build, and full CI. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-26 00:47:47 -07:00
DottaandPaperclip 8781f06a87 feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connections let those agents use external services with explicit
access rules.
> - Zapier, Arcade, Composio Connect, and Executor already have setup
and runtime support.
> - Their experimental switch still blocks discovery and setup by
default.
> - This pull request removes those gates and the Settings toggle.
> - Users can connect these providers without enabling an experiment.

## Linked Issues or Issue Description

Refs #13755. Refs #13941.

**What existing behavior does this improve?**
Apps browsing, inline setup, and agent connection search for the four
MCP aggregators.

**Current behavior**
An instance must enable the MCP aggregators experiment before users or
agents can start setup.

**Proposed behavior**
All four providers are available by default on local and managed
instances. Old stored and managed values still parse but cannot disable
them.

## What Changed

- Remove the aggregator gates from Apps, inline setup, server setup, and
agent search.
- Remove the Settings toggle and its UI hook.
- Retain the old setting key only for upgrade compatibility. Normalize
it to true and ignore managed overrides, as Apps already does.
- Replace opt-in fixtures with default-on coverage. Test old false
values, all four setup flows, provider choice, and the removed toggle.
- Update current connector guidance and remove the opt-in from the
runner acceptance fixture.

## Verification

- 306 focused tests passed across eight files: shared remote MCP
contracts; server remote MCP lifecycle, aggregator fallback, settings
normalization, and managed overlay; UI Apps browsing, setup, and
experimental settings.
- Server and UI TypeScript checks passed.
- UI token gates and `git diff --check` passed.
- The full local suite was not run, per the maintainer's instruction.
All 54 CI checks passed; two checks were skipped. One unrelated
workspace-preview readiness timeout passed on one failed-shard retry.
- The setup fixtures use simulated MCP responses. This change does not
claim new live provider acceptance.

## Risks

- Existing instances now show all four providers, even if the old flag
was false. This is intentional.
- External provider choice, credentials, company isolation, agent
grants, and tool policies still apply. Showing a connector does not
authorize an external account.
- No data migration is required. The compatibility key keeps old managed
configuration documents valid.
- Historical Zapier live acceptance remains incomplete in the existing
evidence report. The maintainer explicitly requested the default-on
rollout for all four existing providers; the report records that scoped
exception.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository
tools, and test execution. The context window size is not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 17:31:32 -05:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00
Devin FoleyandPaperclip 0f8750627f fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path

> - Paperclip runs agents and recovers failed tasks.
> - Codex ACP can report usage exhaustion as a typed terminal failure.
> - The shared engine removes provider text before it stores the result.
> - Codex did not use the existing terminal classifier hook, so a quota
failure became a generic turn failure.
> - This change classifies explicit usage exhaustion before that text is
removed.
> - Recovery can then use quota backoff and the supported reset clock.

## Linked Issues or Issue Description

**What happened?**

A Codex ACP `limit` failure with explicit usage-exhaustion text produced
`acpx_turn_failed`. Recovery could schedule the ordinary short retries
because the quota classification and reset time were lost.

**What did you expect to happen?**

Keep the failure visible and use the existing provider-quota wait.
Preserve a supported reset timestamp without storing provider text.

**Steps to reproduce**

Return a terminal ACP session failure with category `limit` and title
`You've hit your usage limit for GPT-5. Switch to another model now, or
try again at 4:30 PM (America/Chicago).` The real child-process
regression tests exercise both pinned ACPX versions in persistent and
oneshot modes.

**Paperclip version**

Base commit: `32573876d4`.

Related work: #13651 and #13831 provide the shared hook and Claude
classification. #11854 includes quota handling as part of optional
credential rotation, but reads the already-sanitized result; this patch
handles the typed terminal boundary without adding rotation. #13549
reads recovery text after this boundary and cannot recover discarded
provider text. #9011 concerns the Codex CLI backoff. This change leaves
those other mechanisms in place.

## What Changed

- Register a Codex terminal-failure classifier with the existing ACP
engine hook.
- Recognize explicit usage exhaustion only in a typed `limit` failure.
- Reuse the Codex reset-time parser and existing provider-quota recovery
fields.
- Leave context, turn, rate, budget, storage-capacity, and unknown
failures on their existing paths.
- Test real ACP children, both dependency patches, privacy, and recovery
classification. Document the boundary.

## Verification

- 88 focused tests passed across Codex ACP, parsing, and server recovery
classification.
- An initial cross-adapter run passed 53 tests, including the existing
Claude quota suite.
- Removing only the classifier registration makes five new integration
tests fail; restoring it passes all 19 new tests.
- `pnpm -r typecheck` and `pnpm build` passed.
- Full Linux CI passed on `b10d60e002`, including all unit/integration
shards, browser shards, typecheck/build, runner checks, and
release/package gates.
- The duplicate local `pnpm test:run` reported three skill-cache
failures and two runner-suite failures. It is not claimed as a full
local pass.
- The three skill-cache failures reproduce in a clean worktree at the
unchanged base commit (3 failed, 107 passed, 3 skipped across the skills
and runner files). The cache rename reports `EACCES` on macOS.
- An isolated runner-suite run reports two embedded PostgreSQL startup
failures before its assertions (37 tests pass). No runner source is
changed; the Linux CI runner and server suites passed.
- Greptile reviewed the current head at 5/5 with no actionable findings
or unresolved threads.

## Risks

- Explicit usage-exhaustion failures now wait for quota recovery instead
of short generic retries.
- Unknown wording retains the existing behavior. The generic historical
terminal-limit message alone cannot establish quota exhaustion.
- Reset parsing keeps the existing Codex clock formats. A missing or
unsupported reset uses the existing quota backoff.
- No schema, credential, UI, dependency, or deployment changes. Provider
text stays in memory and is absent from results and logs.

## Model Used

OpenAI GPT-6 via Codex. Exact runtime model variant and context-window
size were not exposed. Used reasoning, repository inspection, editing,
and local tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:03:19 -07:00
DottaandPaperclip b0155a681a feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack conversations use the same tasks and agents as the Paperclip
board.
> - A board reply must reach that Slack conversation and let the agent
continue the work.
> - An assigned agent also needs its Slack tools during normal tasks and
scheduled routines.
> - Both paths must keep the linked user's authority, delivery rules,
and conversation history.
> - This pull request adds those paths and reduces setup friction for
Slack bots.

## Linked Issues or Issue Description

**Subsystem affected**

Server orchestration, Slack connector tools, shared contracts, and chat
setup UI.

**Problem or motivation**

Replies entered in Paperclip did not provide a complete round trip to
the linked Slack thread. Slack and board wakeups could select different
model sessions for the same task. Agents also lacked their assigned
Slack tools outside Slack-origin work, which prevented a routine from
sending its responsible user a briefing. Inviting a bot could leave the
new channel disabled.

**Proposed solution**

Mirror human board messages with author attribution and route the agent
result to the same thread. Use the same session key across both entry
points. Supply Slack tools to the connection's assigned agent in normal
tasks and routines, using the current responsible user's verified link.
Enable newly invited channels while preserving explicit disabled
choices. Add a browser-agent setup prompt to the Slack wizard.

**Alternatives considered**

A separate Slack scheduler or task dispatcher would duplicate existing
Paperclip workflows. Reusing the connection owner's identity would grant
the wrong authority. Replaying old channel history could start
unintended work. This change uses ordinary task wakeups, routine
dispatch, and Slack's original invitation mention event instead.

**Roadmap alignment**

Extends the shipped Scheduled Routines and governed Apps capabilities.
It does not add a separate task lifecycle. Related work: #13828 and
#13809. Related test stabilization: #13877. The existing plugin
Slack-control proposals are separate from this built-in connector
change.

## What Changed

- Queue human Paperclip messages for the original Slack thread with
display-name attribution and stable delivery identities. Require the
author’s current linked Slack identity and recheck access before
delivering messages or agent replies.
- Apply pause, dependency, cancellation, and closed-workspace guards
before explicit Board sends request work and again when the durable
outbox dispatches it.
- Route agent results back to Slack and preserve model-session
continuity, including replies that reopen completed tasks.
- Resolve assigned Slack connections for normal agent tasks and
routines. Recheck the responsible user's link, membership, and
permissions at execution.
- Add `slack_open_dm` for the responsible user's bot DM and request the
`im:write` scope.
- Enable newly discovered invited channels. Keep explicit OFF choices
and normal admission and deduplication rules.
- Add a copyable Slack setup prompt for a computer-use agent, with
Storybook coverage. Share the prompt-button component with GitHub.
- Update Slack tool documentation and runtime instructions.
- Stabilize the mobile project browser test by waiting for the final
canonical route before editing, preserving all persistence assertions.

## Verification

- Live staging: invited the bot after the first mention. The channel
became enabled and the bot answered that original mention.
- Live staging: a normal Paperclip reply appeared in Slack with author
attribution. The agent completed the calculation and replied once in the
original thread and in Paperclip.
- Live staging: a codeword entered in Slack was recalled from Paperclip.
A following Slack calculation used the result from the Paperclip turn.
Run metadata confirmed the same model session for both entry points.
- Live staging: a scheduled routine used `slack_open_dm` and
`slack_post_message` to deliver one DM. The existing app was reinstalled
with `im:write`. The test routine was paused after verification.
- Before the master merge: 397 focused feature tests passed. The
continuity fix passed all 76 issue comment/update route tests and six
focused route/integration cases. Typecheck, build, and token gates
passed.
- Review fixes: 47 focused integration cases passed, covering link
revocation/replacement, private membership removal, guarded outbox
dispatch, concurrent workers, lost scheduler responses, a real
one-connection pool, exact reply provenance, and attachment retries. All
99 issue-comment route tests and the Slack catalog browser test passed.
- Full local typecheck, production build, token gates, and
module-boundary checks passed. The full local test command passed 25,915
tests before a 15-second timeout in
`issue-thread-interaction-routes.test.ts`; that entire suite passed on
isolated rerun (81 tests). Remaining serialized coverage is provided by
the current-head CI shards.
- An unchanged Cursor adapter test hit its 10-second limit in CI; all
five tests in that file passed on a local rerun in 3.11 seconds, and the
failed CI shard passed on its single retry.
- The preview-server readiness test passed a local rerun (28 tests). The
mobile-project readiness fix passed three repetitions of both browser
tests (6/6).
- Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates
passed, including all eight browser shards, all server/chat suites,
typecheck, build, runner checks, and security checks. Greptile reviewed
this exact commit at 5/5; all review threads are resolved.

## Risks

- Human messages on a Slack-linked task now publish to its Slack thread.
The task banner states this behavior. Incoming Slack messages and
internal agent bookkeeping must not echo back.
- Normal tasks and routines can now use the assigned bot. Authority
remains bound to the current responsible user's link; it does not fall
back to the connection owner. Revocation, private-context limits, and
queued-write checks still apply.
- Existing Slack apps need `im:write` and a reinstall to open DMs. Other
existing capabilities remain available without that scope.
- New invited channels default to enabled. Explicit disabled choices
remain disabled. Channels created by bot tools still require a person to
enable responses.
- No database migration or new provider credentials are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI,
and browser tools. The runtime does not expose a more specific model
build identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:51:45 -05:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
DottaandPaperclip b41ccf097f fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agents request connections through cards in task threads.
> - MCP aggregators need provider-specific URLs and authentication.
> - The task dialog used the generic setup form and omitted these
fields.
> - This pull request uses the same provider setup controller in tasks
and Apps.
> - Users can configure a connection without leaving the task.

## Linked Issues or Issue Description

Related: #13755, #13855.

**What happened?**

An inline Executor request opened a very wide dialog with an empty
credential step. Connect failed because the MCP URL was missing. The
other aggregator cards also bypassed their provider setup.

**Expected behavior**

Each card shows its provider instructions, URL field, and authentication
options in a bounded dialog. Completing setup grants access only to the
requesting agent.

**Steps to reproduce**

1. Enable experimental MCP aggregators.
2. Have an agent request Zapier, Arcade, Composio, or Executor from a
task.
3. Open the card and continue past Access.

## What Changed

- Route page and task setup through the same provider controller.
- Bound the task dialog width and preserve the requesting agent's access
scope.
- Support existing accounts, saved drafts, URL/token setup, and
task-bound OAuth.
- Keep a sign-in link available when the browser cannot open a popup.
Verify completion through the existing durable callback path.
- Add inline Access, configuration, and narrow Storybooks for all four
providers.
- Document the shared setup requirement and correct Executor's URL
instructions.

## Verification

- Focused Vitest selection: 24 passed. Covers all four inline forms,
requester access, existing accounts, saved drafts, OAuth retry, callback
validation, popup cleanup, and generic reconnect endpoint preservation.
- UI typecheck, UI build, design-token gates, and Storybook build
passed.
- Live local browser: new Executor, Arcade, and Composio connections
completed provider consent from task cards. Each appeared Connected with
the requester selected.
- Real Test calls returned Executor output `4`, an Arcade public GitHub
star count, and Composio tool-discovery results. An ungranted agent was
denied access. Real Paperclip process-agent runs discovered each
provider catalog with only the requester’s connection installed.
- Deployed implementation commit `5d722e89c` to the isolated staging
tenant. The original failing Executor card now completes, discovers
seven actions, limits access to its requesting agent, and resumes that
agent. Its continuation completed real Executor calls and the provider
resume flow, then returned an upstream Airtable authorization link. A
real staging Test call returned `4` in 1.5 seconds. Later PR commits add
regression coverage and popup-unmount cleanup.
- Zapier fresh-token browser test remains pending a provider clipboard
handoff. Its URL and token flow passes focused tests.
- Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook
deployment and visual regression jobs skipped. Greptile 5/5, both review
threads resolved. Canceled runners and unrelated chat timeouts passed
the single retry on unchanged code.
- The full local test suite was not run, as requested by the maintainer.
CI runs the repository gates.

## Risks

- OAuth popup behavior differs by browser. The explicit sign-in link and
durable server completion checks provide recovery.
- Saved task drafts store only a connection ID in browser storage.
Credentials remain in the existing server vault.
- No database or server protocol changes.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code
execution, and browser tools. The exact context-window size is not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 09:19:18 -05:00
DottaandPaperclip 4721f55803 fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - MCP aggregator setup can return from provider consent to a saved
draft.
> - The branded setup did not explain failed or cancelled authorization.
> - It also offered identity changes that the server does not apply when
a saved connection resumes.
> - This pull request explains OAuth return outcomes and keeps the
displayed identity consistent with the saved policy.
> - Users can understand the outcome and retry the same connection.

## Linked Issues or Issue Description

Related: #13755 introduced the MCP aggregators. #13758 retired the
legacy Composio broker. #13584 proposes changes to the provider handoff
window; this fix retains the current handoff behavior.

**What happened?**

Cancelling Composio consent returned to the setup form with no
explanation. The same controller ignored failed OAuth callback outcomes.
Returning to Access on a saved draft also offered personal/shared
choices, although the server retains the saved identity. This could make
the OAuth request disagree with that identity.

**Expected behavior**

Explain cancellation or failure, preserve the draft, and offer Try
again. Display the retained credential identity and start OAuth with the
policy returned by the server.

**Steps to reproduce**

1. Enable MCP aggregators and start a Composio connection.
2. Continue to provider consent and cancel it.
3. Observe the return screen. Before this change, it showed the form
without cancellation feedback.
4. Go back to Access. Before this change, the form offered
personal/shared choices even though resume retains the original
identity.

**Paperclip version or commit**

Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`.
The original report of successful consent leaving setup unfinished did
not reproduce. This PR addresses the recovery defects observed during
that investigation.

**Deployment mode**

Isolated local development instance and authenticated staging
deployment, with real Composio consent and provider calls.

## What Changed

- Read the OAuth callback outcome in branded MCP setup and show
cancellation or failure feedback.
- Retry the same saved draft without displaying untrusted callback error
text.
- Keep saved personal/shared identity fixed in Access and select OAuth
identity from the returned credential policy.
- Add six focused regression cases and authorization-failure Storybook
states for Arcade, Composio, and Executor.
- Document return-screen recovery and retained identity in the connector
playbook.

## Verification

- `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth
return' --maxWorkers=1`: 6 passed, 143 skipped, including after
integration with current master.
- `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed.
- `pnpm check:token-gates`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm --filter @paperclipai/ui build-storybook`: passed for the
implementation commit.
- Browser recovery: cancelled real Composio consent, observed the new
feedback, returned to Access, retried the same personal draft, completed
consent, and ran a real tool call.
- Staging on implementation commit
`84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal
connections each completed on the first consent attempt and loaded 11
tools. Real discovery, execution, and schema calls succeeded. Both
connections stayed Connected after reload. A real agent used the shared
connection through the Paperclip gateway and returned the public
repository documentation hierarchy with one success and zero errors.
- All current-head CI gates passed on
`662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and
agent-chat browser shards each had an initial timeout; both passed on
one targeted rerun without code changes. Greptile reviewed this exact
head at 5/5, with no open review threads.
- No full local suite was run, as requested. CI provides the broader
checks. The PR adds a master merge and documentation after the
live-tested implementation commit.

## Risks

- This shared setup controller also serves Arcade and Executor. Their
callback rendering and retry behavior have focused test coverage; this
investigation used Composio for live provider testing.
- Saved identity remains fixed during resume. A different identity
requires a new connection, consistent with server behavior.
- No database, protocol, credential storage, or gateway policy changes.

## Model Used

OpenAI GPT-6 via Codex, with code editing, shell tools, and browser
testing. The exact runtime model ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 08:05:30 -05:00
DottaandPaperclip a10702a878 feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people start and continue agent tasks from other
services.
> - A Slack conversation needs access to its surrounding discussion and
Slack collaboration tools.
> - The agent must use the linked requester's access and keep private
material within its permitted audience.
> - This pull request adds Slack tools through the existing connector
contribution and approval framework.
> - People can ask an invited bot to read a discussion, create follow-up
tasks, and collaborate in Slack.

## Linked Issues or Issue Description

**Subsystem affected**

Chat connectors, connector runtime, tool gateway, and connection
Settings/Access.

**Problem or motivation**

Slack-origin tasks can receive messages but cannot inspect the rest of a
channel or act through the originating bot. People must paste context or
configure a separate integration.

**Proposed solution**

Supply typed Slack tools and a bundled skill only to the originating
task and assigned agent. Resolve the linked requester on the server.
Check bot and requester access before reads and writes. Use existing
durable actions and approvals. Retrieved messages remain source
material.

**Alternatives considered**

Slack's user-OAuth MCP server does not replace the customer-created chat
bot. An unrestricted Web API proxy would not provide suitable permission
or publication boundaries.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps work with a provider
contribution. It does not add a task dispatcher or a separate Slack task
lifecycle. Related: #11144 covers generic per-user MCP grant execution;
this change binds Slack bot operations to chat-origin tasks.

## What Changed

- Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and
shared native/HTTP execution.
- Bind tools to company, endpoint, task, run, assigned agent, and
admitted linked requester. Check membership and revocation on each call
and before queued writes.
- Add paginated reads, bounded history search, source links, messages,
file uploads, reactions, pins, bookmarks, topics, canvases, lists, and
approved channel operations.
- Restrict private-source publication, including automatic replies and
uploaded deliverables. Keep other people's bot DMs inaccessible.
- Reuse action receipts, idempotency, approvals, and reconciliation.
Suppress an identical explicit-send/final-reply duplicate. Return
governed results through their verified originating conversation.
- Add endpoint-bound personal search OAuth storage and lifecycle. Keep
native real-time search disabled until a runtime meets Slack's
transient-result requirements. Current runtimes use bounded history
search.
- Show capabilities, scope upgrades, and personal search authorization
in Settings/Access and Storybook. Document provider and runtime limits.

## Verification

- Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no
unresolved findings. One unchanged rapid-callback timing test passed on
a single CI retry.
- Approval presentation regressions cover board-comment precedence and
exact Slack publication; the expanded database assertion passed in CI.
The local PostgreSQL startup probe later became unavailable, so that
final assertion was verified in CI. Slack setup and failed-run retry
browser tests also passed locally.

- Full workspace typecheck and build passed. Server typecheck/build
passed again after the approval routing fix.
- Broad local suites passed in separate groups: server 12,958 tests, UI
6,555, shared 770, skills catalog 20, and other workspace packages
2,652. CLI and serialized server checks passed after environment/timeout
retries. These are composite results, not one uninterrupted green
full-suite invocation.
- PostgreSQL authority regression covers admitted identity,
cross-company/task/agent rejection, recovery, retained-session
revocation, OAuth refresh/disconnect races, approval execution, exact
publication lineage, retries, uncertain sends, and duplicate
suppression.
- Gateway/response regressions cover separate-origin approval batches
and durable continuation. Focused provider, access, search, native
runtime, route, and AgentMail regressions pass.
- Storybook capability, missing-scope, OAuth configuration,
authorization, and disconnect states were inspected in the browser.
- Live staging: read a channel decision and full thread, create exactly
two assigned backlog tasks, add a reaction, paginate discovery to
exhaustion, and return bounded search matches with source links and
coverage.
- Live staging: create/edit/read a canvas and list, inspect the canvas
in Slack, post/edit one message, and create a channel only after
approval. New channels remain disabled for responses.
- Live staging: read a response-disabled channel from the requester's
DM; writes to that channel were denied. The test setting was restored.
- Final live retest passed: explicit file upload and exact content
read-back; approved deletion of only the disposable bot message;
continuation confirmation returned to the original Slack thread without
repeating the action.
- Optional OAuth, private multi-user boundaries, native RTS, and CLI
provider execution are not fully live-qualified. The staging agent
initially supplied malformed tool arguments; valid arguments succeeded,
and the tool/skill descriptions now emphasize UUID write keys.

## Risks

- Existing Slack apps must add scopes and reinstall for new
capabilities. Provider plans and document permissions can still restrict
operations.
- Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET`
for governed tool actions. The staging instance was configured with
explicit operator approval; fleet provisioning is a separate gap.
- Native RTS is not exposed on current transcript-retaining runtimes.
Bounded history scans are deliberately reported as incomplete. Inline
file reads support text/canvas content up to 256 KiB; other types return
metadata.
- Private document edits fail closed when the full audience cannot be
verified. Uncertain effects other than posts/uploads require inspection
instead of blind retries.
- Shared approval-delivery code now separates outcomes by source run to
preserve origin boundaries. No database migration is required.
- A separate completion-validator gap remains when the agent cites a
prior run's registered artifact during finalization. It asked for
registration again even though Slack delivery was confirmed. This change
does not add a connector-specific task-completion policy.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
browser testing. The exact deployed model identifier and context-window
size were not exposed in the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:29:07 -05:00
DottaandPaperclip a959e47508 fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents controlled access to external services.
> - Google Chat uses OAuth profiles and reviewed MCP tools.
> - Those profiles request membership and read-state access that the
supported feature set does not need.
> - Removing read-state access also requires us to block unread search
filters, including on existing connections.
> - This pull request reduces both OAuth methods and enforces the
reduced search contract before dispatch.
> - Users retain conversation lookup, message history, ordinary search,
and approved message sending.

## Linked Issues or Issue Description

**What happened?**

Both Google Chat profiles request membership and read-state scopes. The
supported tool set does not include membership listing or read-state
updates. Message search still advertises an unread filter. Related work:
Refs #12619.

**Expected behavior**

Managed and customer-owned OAuth request only the scopes needed for
supported features. Unsupported unread filters fail clearly before any
provider call. Existing cached catalogs and broader grants must not
bypass that policy.

**Steps to reproduce**

Start Google Chat OAuth from either connection method and inspect the
requested scopes. Inspect the message-search tool schema, then submit a
search with `searchParameters.isUnread` set to true or false.

**Paperclip version or commit**

The scope change is based on master at `110d176fc`.

**Deployment mode**

Managed Cloud and self-hosted instances with Google Chat Apps enabled.

## What Changed

- Remove `chat.memberships.readonly` and `chat.users.readstate.readonly`
from shared profiles and all four Chat connection methods.
- Hide unsupported read-state fields and instructions in agent and board
Test tool schemas.
- Reject explicit unread filters, including false, null, snake-case
fields, and encoded filter objects, before provider dispatch.
- Recheck previously approved calls and support existing profile-bound
and URL-only Chat connections.
- Add scope, signed broker request, OAuth URL, allowlist, schema, and
dispatch regression tests.
- Document coordinated app/broker rollout, existing-grant reconnects,
and the remaining deployment checks.

## Verification

- All seven focused OAuth and Chat gateway test files pass: 498 tests on
the rebased branch.
- `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4.
- The complete sharded CI test matrix passes, including general server,
Chat, workspace, serialized server, Runner, and all eight browser e2e
shards. The duplicate unsharded local `pnpm test:run` was stopped after
CI passed; it did not complete locally.
- `git diff --check` passes.
- `node scripts/ingest-app-definitions.mjs` succeeds and leaves the
branch unchanged. Google Workspace JSON is the durable reviewed input
used by the generator.
- All current-head CI checks pass at
`7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build,
and canary dry run. Greptile is 5/5 with zero unresolved threads. The
generator concern was withdrawn after review of the source and
regeneration evidence.
- No production deployment or live Google consent test was performed.
After coordinated deployment, verify reduced consent scopes, normal
search/history, message sending, and rejection of unread filters.

## Risks

- Coordinate deployment with the companion Cloud broker scope change.
Mixed versions can reject exact-scope requests.
- Existing tokens are not narrowed or revoked. Grants with old scopes
need new consent. Do not revoke a shared Google client to migrate one
profile.
- Explicit unread filters now return an error instead of being sent to
Google. Ordinary search and the approved send tool remain available.
- No database, UI, lockfile, or workflow changes.

## Model Used

OpenAI Codex, a GPT-5-based coding agent, with tool use and code
execution. The exact runtime model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:12:30 -05:00
Devin FoleyandPaperclip a7d3b17a97 Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup
instructions from catalog and health routes. Bound response parsing and keep
unknown upstream errors reportable without exposing provider settings links.

Verified 361 focused tests, server typecheck, and authenticated discovery.
The three route regressions fail before this change and pass afterward.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:25:06 -07:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
DottaandPaperclip 57fd8b70d2 feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack connections let a team talk to those agents in Slack.
> - Agents now have a saved avatar, but Slack setup did not offer that
image.
> - A matching avatar helps a team recognize its agent.
> - This pull request adds an optional avatar step and a download in
connector Settings.
> - Users download a PNG and upload it directly in Slack with clear
instructions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Slack connector onboarding and its Settings page.

**Current behavior**
Setup does not offer the assigned agent's avatar or explain how to
upload it in Slack.

**Proposed behavior**
After Slack connection verification, users can download a 512 × 512 PNG
of their agent's saved avatar. They can upload it in Slack, confirm, or
skip. Settings keeps the download and upload instructions available
after onboarding.

**Reason and benefit**
The same avatar helps people recognize the agent across Paperclip and
Slack. Users who skip the optional step can return to it in Settings.

**Breaking changes**
None. No schema, authentication, Slack scope, or provider API change.
Completed connections keep their existing completion state. Searched
existing Slack avatar and Cliptoon PRs; no matching implementation was
found.

## What Changed

- Add an optional avatar step before personal Slack account linking.
Keep the numbered sidebar and shared footer.
- Resolve the selected agent's saved appearance for the preview and PNG
download.
- Add the same download and expandable upload instructions to connector
Settings.
- Remember uploaded or skipped per company and endpoint in browser
storage. Treat uploaded as user confirmation, not provider verification.
- Reject failed or non-PNG download responses and allow retry.
- Reuse the production avatar components in onboarding and Settings
stories.
- Test wizard progression, resume, Settings, download recovery, storage
isolation, and terminated assigned agents.
- Exercise real PNG downloads in the Slack browser flow and keep default
app names consistent with app creation.
- Fetch the assigned agent directly so its saved avatar remains
available after termination.

## Verification

- Focused chat suites: 48 passed; the two affected suites passed again
after the final naming fix (32 tests).
- Slack browser E2E passed through setup, avatar download, account
linking, and Settings download. PNG signature and 512 × 512 dimensions
verified.
- UI token gates passed.
- Browser: downloaded the real 512 × 512 PNG; checked confirmation,
return, mobile layout, and Settings instructions.
- Full workspace typecheck, application build, and production Storybook
build passed.
- All latest-head CI checks passed (54 passed, 2 skipped), including all
browser, chat, general, and serialized test groups. The unrelated Sentry
test failed once and passed on the single CI rerun; its suite also
passed locally.
- Local full-suite attempt encountered a rapid Slack callback ordering
failure under concurrent build load; that test passed in isolation, and
all three chat shards passed in CI. The remaining local run was not used
as the merge gate.
- Review the Connections / Slack / Add avatar and Avatar in Settings
stories.

## Risks

- Slack upload is manual. Confirmation does not claim to verify the
Slack icon.
- Optional step progress is browser-local. Clearing storage or changing
browsers can show it again. Setup still works when storage is
unavailable.
- The existing avatar API remains the image source. Download failures
show a retry message.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 09:38:48 -05:00
DottaandPaperclip 9f30eb10dd fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path

> - Paperclip manages AI agents and keeps their work attached to tasks.
> - Chat connectors carry user messages and agent replies between a
provider and those tasks.
> - Each extra startup and context reset delays a reply.
> - Managed account metadata was lost during adapter decoding, so
compatible follow-ups started fresh.
> - This branch fixes the reset and measures the remaining preparation,
execution, and delivery costs.
> - The changes must preserve account isolation, authorization, durable
output, and recovery ownership.

## Linked Issues or Issue Description

Refs #13699. The related service lifecycle work in #13410 and #13408 is
separate; this branch focuses on task-bound chat response latency.

**What happened?**

Managed AI follow-ups started new provider sessions even after their
configuration fingerprint stayed stable. The Codex codec removes unknown
fields. The resume check then read the removed credential identity and
treated it as a credential change.

**Expected behavior**

Compatible follow-ups resume the correct provider session. Changes to
credentials, responsible users, permissions, or task configuration
retain their reset behavior.

**Steps to reproduce**

1. Use a Slack connector with a managed AI connection.
2. Send a message, then send a same-thread follow-up.
3. Inspect the configuration reset reason and the provider session
identity.

**Paperclip version or commit**

Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`.

**Deployment mode**

Cloud staging with a native Codex runner.

## What Changed

- Read saved credential identity before adapter decoding discards it.
- Remove the internal credential identity from adapter-facing session
params.
- Test the real Codex codec and missing, changed, or unmanaged identity
cases.
- Preserve configured warm Codex runners and flush refreshed credentials
after every turn.
- Fence detached or closing session handles from successor credential
ownership.
- Stage current Codex launch credentials after restoring durable session
history, uploading launch assets only once.
- Reuse a runner binary already in the retained sandbox only when its
SHA-256 matches the controller-owned artifact; still verify required
capabilities before launch.
- Lock the task before the run when saving results, preventing deadlocks
with task updates.
- Scope reusable projectless sandboxes to the company, environment,
task, agent, and runtime configuration; verify Daytona sentinels for
that scope.
- Admit a new authorized chat message after a fully committed failed run
and verified process cleanup.
- Send compact deltas for verified plain-text Slack continuations. Match
the actual prior run and current comment identity/body; exclude edited
historical comments and prior agent output, preserve genuine brief edits
and the full bootstrap fallback.
- Keep attachments, omitted input, questions, approvals, recovery, and
other providers on their existing framing.
- Document managed session compatibility, credential lifecycle, and
compact continuation boundaries.

## Verification

- Workspace/session coverage: 156 tests passed.
- Native session and credential ownership coverage: 390 tests passed,
including exact artifact reuse, mismatches, failed probes, timeouts, and
explicit artifact overrides.
- Explicit continuation and durable chat authorization coverage: 172
tests passed.
- Session resume and launch preparation coverage: 416 tests passed.
- Result persistence coverage: 15 tests passed. The new concurrency test
reproduced a PostgreSQL deadlock before the lock-order fix.
- Environment lifecycle coverage: 92 tests passed, including projectless
reuse and task/agent isolation at both selection and atomic handoff.
- Daytona plugin coverage: 237 tests passed; 6 gated tests skipped.
Standalone plugin build passed.
- Compact Slack continuation and native resume coverage: 69 tests
passed, including full-bootstrap retention, matching message
authors/bodies, current-delivery selection, rejection of duplicate
identities and historical comments, brief edits, and attachment/recovery
fallbacks.
- Final frozen-head `pnpm test:run` on repository-supported Node 26: 668
suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped.
The sole failure was a local `socket hang up` in
`issue-recovery-actions.test.ts`, not an authorization assertion
mismatch. All 57 tests in that suite passed three fresh reruns, and the
suite passed latest-head CI. The full local invocation is therefore not
claimed green.
- An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup
failure; that 11-test catalog suite passes on Node 26 and in CI. No test
behavior or timeout was relaxed.
- Full local typecheck and build passed. Latest-head CI is green;
Greptile is 5/5 with no unresolved review threads.
- Two real Slack baseline replies took 25.1 and 24.6 seconds
(24.9-second mean). Three same-thread signed probes on this head took
23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample
and a modest wall-clock improvement, not a large or statistically
established speedup.
- In that same thread, uncached provider input fell from 8,514 tokens
before compact input to 694–765 tokens afterward. The current delivery
uses a 362-character delta; the full 19–21k-character bootstrap remains
available for failed resume. Verified runner artifact preparation fell
from about 1.2 seconds to 0.6 seconds.
- A fresh thread created a separate task, sandbox, and provider session
with full bootstrap (24.4 seconds). Its follow-up reused its own
sandbox/session and compact input (28.3 seconds, including 16 seconds of
model execution). Model variability and process startup remain
substantial.
- A signed duplicate webhook produced exactly one user comment, one
successful run, and one final Slack reply. Slack's API independently
confirmed the actual replies and a public task URL without an internal
or pool hostname.
- Earlier signed probes verified recovery after a failed run and reuse
across a server deployment. The final idle test observed Daytona report
the sandbox as stopped, then delivered a new reply in 19.9 seconds using
the same sandbox/provider-session identity and compact input. Slack’s
API confirmed that reply.
- Live probes use signed synthetic inbound webhooks and real outbound
Slack delivery, read back through Slack’s API. The final browser recheck
found the Mac locked and the Slack tab blocked by another extension, so
this is not claimed as full UI E2E proof.
- This is a review branch. Do not merge until the maintainer reviews it.

## Risks

- Incorrect session reuse could mix account or task context. Missing or
changed identities continue to reset, and existing authorization checks
remain in place.
- Warm mode remains opt-in. Remote warm mode requires a reusable sandbox
lease. Retained processes keep credentials until they close, so idle
expiry and ownership fences are required.
- A fresh user message may continue after a committed provider failure.
Approval, current authorization, process termination, and prior-result
checks remain required.
- Projectless sandbox reuse is task- and agent-scoped. Missing or
mismatched ownership cannot replace an existing lease; existing
workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox
per task/agent, so provider auto-stop and deletion policies still
determine idle compute and storage costs. Fleet defaults are unchanged.
- Compact prompts apply only after proven resume and a matching
prior-run delta. Missing or specialized context falls back to full
input; fresh sessions always receive the full bootstrap.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-20 13:12:44 -05:00