Commit Graph
611 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 6681c71b40 fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Browser tests check chat state across navigation, reload, and agent
runs.
> - Runtime tests check that a service is ready before Paperclip
publishes its address.
> - CI for #13891 failed in these paths, then passed attempt 3 with the
same code.
> - The browser traces stopped during development asset startup. A
separate sidebar assertion used an unstable focus path through the rich
editor.
> - This PR gives browser tests fresh built assets and a direct keyboard
path to the star button. It adds service-worker reload coverage under
CPU throttling.
> - A later browser failure exposed a real reload race: comments can
change between the first query and the first live subscription. The UI
now refreshes active queries when that subscription opens.
> - Runtime fixtures now have separate registry state and better failure
evidence. The readiness deadline remains unchanged.
> - A later CI run exposed a wall-clock backoff assertion and three
authorization cases sharing one test lifecycle. The PR anchors the
assertion to transport time and separates the cases.

## Linked Issues or Issue Description

Refs #13891. Related runtime ownership and cleanup work: #11791, #11389,
#11278.

Evidence: [original CI run, attempt
3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3).
Attempt 3 passed both affected shards without code changes. This PR is
separate from the wake-payload change, which has since merged.

| Failure | Diagnosis and classification |
| --- | --- |
| Sidebar blank page and retry text missing after reload in CI | Both
traces show a blank document before React startup. Vite connects, but
the failing page makes no application API requests. The service worker
forwards the unbundled development module graph. The retry response is
already stored and its process-adapter run succeeded before reload. This
places the failure in browser bootstrap, not wake payload or reply
persistence. The exact reason the development module graph stopped is
not established by the retained trace. The harness now serves built
assets, and the new test covers startup and controlled reload at 4x CPU
throttling. |
| Star opacity remains zero on macOS | Reproduced on unchanged master
`f55759942b`. The old test clicked the rich editor, focused the star,
then used Tab and Shift+Tab. An instrumented baseline run captured the
sequence: Tab moved from the star to the next sidebar link, then the
editor bundle called `focus()` on its contenteditable before Shift+Tab.
That key reached the composer, where it is a work-mode shortcut. This
confirms a pending editor selection update stole focus; it was not the
browser skipping sidebar buttons. The assertion did not prove the star
retained focus. The test now moves the pointer away, focuses the
preceding sidebar link, presses Tab once, and asserts both actual focus
and opacity. This is test synchronization and keyboard traversal, not a
demonstrated CSS defect. |
| Runtime readiness exceeds 10 seconds | The old error only says `fetch
failed`. There is no child startup output in that failure, so it cannot
distinguish slow process startup from a refused or stalled probe. It
passed unchanged locally and took 3.416 seconds in attempt 3. Resource
contention is plausible but unproved. This PR does not claim a proven
historical runtime root cause: it isolates fixture registry/log files,
checks ports before spawn and after stop, checks live backends at
publication, and preserves transport errors, probe count, elapsed time,
and fixture startup timestamps for the next occurrence. |
| Later CI: recovery reply disappears after reload | Product
synchronization bug, distinct from the blank-page bootstrap failure.
Reproduced locally with a trace: the comments request started at
`21:56:07.443`, the server saved the reply at `.520`, and the first live
subscription started at `.541`. The reload fetched a successful run and
its complete log, but missed the comment event. The provider refreshed
after reconnects only. It now refreshes active queries on the first
connection too. A deterministic regression test fails before the fix. No
browser assertion was changed. |
| Later CI: Slack backoff and authorization tests | The 30-second
backoff assertion required more than 25 seconds to remain when it read
the saved action. CI spent 10.311 seconds in the test, exceeding that
five-second allowance. A local six-second read delay reproduces the
failure; the new transport-anchored lower and upper bounds pass the same
fault injection. The neighbouring authorization test ran three
independent fixtures in one test and timed out at 15 seconds. Each case
now has its own fixture cleanup and the normal per-test deadline, so
earlier cases do not remain active during later global worker sweeps. No
specific production slow call was established. |
| Self-hosted runner loses communication or shuts down | The original
lost-communication failure has no assertion. On final-head [attempt
1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1),
Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet
instances. All received a runner shutdown signal at `22:04:55 UTC`,
within 35 milliseconds, then cancellation. Server shard 6/12 received
the same shutdown signal one minute later. All four were Spot
`m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build
error preceded those stops. This is infrastructure interruption; the
reason the fleet stopped the runners is not available in job logs. |

## What Changed

- Build the browser fixture UI into the static server's preferred
directory, `server/ui-dist`, and disable Vite middleware. Ignore these
generated assets. Start the source CLI directly from the repository
root, as required by the CLI invocation safety contract.
- Keep all existing chat assertions. Add first takeover and three
service-worker-controlled reloads under CPU throttling. Assert that the
page uses built module assets and that opening it creates no chat task.
- Use forward keyboard traversal from the Zeta link to its star. Check
focus before checking the reveal style.
- Give each runtime exposure test a temporary Paperclip home and restore
environment state after process cleanup.
- Check that reserved ports are free before spawn, serve the fixture
response before exposure, and become free after stop.
- Include the nested fetch error, probe count, and elapsed time in
readiness failures. Add a unit test for this diagnostic contract.
- Refresh active queries when the first live connection opens, covering
events missed during initial page loading. Keep reconnect toast
suppression unchanged.
- Measure Slack retry timing from the transport attempt and recovery
completion. This checks the full provider-requested delay without
spending a small wall-clock allowance on unrelated processing.
- Run each Slack authorization-revocation scenario as a separate test,
with cleanup between cases. All assertions remain.
- Document the browser fixture's build and serving mode. No configured
assertion deadline, readiness deadline, retry count, or skip was added.

## Verification

- Final-head [Linux CI, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2):
**green**. All 52 check runs passed; two conditional checks were
skipped. The legacy Snyk status also passed. No pending or failed checks
remain.
- `pnpm -r typecheck` passed.
- `pnpm build` passed again after the final UI fix.
- Focused runtime suites: 38 passed, 3 existing platform skips.
- CLI invocation safety suite: 39 passed.
- Slack timing negative control: inserting a six-second delay before
reading the saved action fails the old assertion. All four revised
authorization/backoff cases pass with that same delay. The diagnostic
delay is not committed.
- Live update suites: 92 passed, including the new regression. UI
typecheck and token gates passed.
- Recovery browser negative control: the unchanged tests reproduced the
missing reply (1 failed, 9 passed). After the first-connection fix, all
six recovery paths passed twice (12 passed). Browser assertions and
deadlines are unchanged.
- Server typecheck passed after the Slack test adjustment.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat
--shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong
to the remaining shards; collection verified exact coverage.
- Final direct-source CLI launch: all five chat session tests passed.
- `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed,
including nine controlled reloads at 4x CPU throttling.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group
general-server-without-chat --shard-index 10 --shard-count 12`: 55 files
passed, 867 tests passed, 21 existing skips. An earlier run could not
initialize PostgreSQL because this Mac exhausted its System V
shared-memory slots. After reclaiming the orphaned segment from this
task's stopped browser server, the full shard passed.
- On final commit `1f1fafc08d`, [browser
4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947)
passed all 16 tests; [browser
8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155)
passed all 23 tests; [server
11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328)
passed 887 tests with one existing skip. The readiness lifecycle case
took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by
runner shutdowns passed unchanged on their single rerun.
- Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments.
- Full local `pnpm test:run` was attempted and stopped after confirming
failures outside this patch: a sibling `skills` directory shadows
bundled Slack/AgentMail skills; macOS rejects rename of read-only
skill-cache directories (`EACCES`, also reproduced in an isolated
filesystem probe); and the host exhausts PostgreSQL System V
shared-memory slots. Focused reruns confirmed these limits. The complete
Linux CI run is the repository-wide verification; the full local run is
not green.
- Baseline evidence: the original sidebar test failed on unchanged
master; a development-mode run at 4x CPU throttling passed six selected
cases, so CPU pressure alone did not reproduce the CI bootstrap stall.

## Risks

- Default browser tests now exercise the shipped static UI. They no
longer implicitly cover Vite middleware or HMR; use the development
server for those checks.
- The new browser startup test uses Chromium CDP, matching the only
configured browser project.
- Opening a live subscription now causes one active-query refresh to
close the initial event gap. This adds startup API reads but no
recurring poll.
- Runtime behavior and deadlines are unchanged except for error details.
The historical readiness stall remains unconfirmed; a green rerun alone
cannot establish its cause.
- Local verification runs on macOS. The runtime lifecycle checks also
passed on Linux CI.

## Model Used

- OpenAI GPT-6 through Codex. The session identifies the model family as
GPT-6; an exact served model ID and context window size are not exposed.
Used reasoning, repository inspection, shell tools, code editing, and
test execution. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted suites passed;
full local host limits are listed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 15:22:21 -07:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
Devin FoleyandPaperclip 4b8ec588f3 Stop duplicating wake context in adapter environments (#13891)
## Thinking Path

> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.

## Linked Issues or Issue Description

Refs #13144, #13860, #13872, #13793.

Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.

Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.

## What Changed

- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.

## Verification

- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on c3b8191f97: 53 checks passed and two skipped. Attempt 3
passed the remaining shards without source changes. Earlier attempts hit
a lost runner, preview readiness, and chat browser failures. The preview
test passed unchanged locally. Local chat tests passed 3/4; the
remaining sidebar focus-style assertion also fails on unchanged master.
No tests were weakened or skipped to obtain the passing rerun.
- The staged diff passed `gitleaks stdin --redact` and `git diff
--check`.

## Risks

- Custom instructions or scripts that read `PAPERCLIP_WAKE_PAYLOAD_JSON`
must use the wake payload in the prompt instead. The variable is absent
even for small wakes.
- This removes one cause of `E2BIG`. Legacy CLI paths for Gemini, Grok,
Kimi, Pi, and Hermes still pass prompts as arguments and retain their
existing argument-size limits. Other large environment variables also
remain subject to OS limits.
- No new context truncation, API route, database migration,
authorization change, or production rollout is part of this PR. Existing
prompt windows and resume rendering remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with code execution and repository tools.
The exact model variant and context window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 13:50:10 -07:00
DottaandPaperclip f55759942b fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-23 11:27:43 -07:00
DottaandPaperclip b41ccf097f fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agents request connections through cards in task threads.
> - MCP aggregators need provider-specific URLs and authentication.
> - The task dialog used the generic setup form and omitted these
fields.
> - This pull request uses the same provider setup controller in tasks
and Apps.
> - Users can configure a connection without leaving the task.

## Linked Issues or Issue Description

Related: #13755, #13855.

**What happened?**

An inline Executor request opened a very wide dialog with an empty
credential step. Connect failed because the MCP URL was missing. The
other aggregator cards also bypassed their provider setup.

**Expected behavior**

Each card shows its provider instructions, URL field, and authentication
options in a bounded dialog. Completing setup grants access only to the
requesting agent.

**Steps to reproduce**

1. Enable experimental MCP aggregators.
2. Have an agent request Zapier, Arcade, Composio, or Executor from a
task.
3. Open the card and continue past Access.

## What Changed

- Route page and task setup through the same provider controller.
- Bound the task dialog width and preserve the requesting agent's access
scope.
- Support existing accounts, saved drafts, URL/token setup, and
task-bound OAuth.
- Keep a sign-in link available when the browser cannot open a popup.
Verify completion through the existing durable callback path.
- Add inline Access, configuration, and narrow Storybooks for all four
providers.
- Document the shared setup requirement and correct Executor's URL
instructions.

## Verification

- Focused Vitest selection: 24 passed. Covers all four inline forms,
requester access, existing accounts, saved drafts, OAuth retry, callback
validation, popup cleanup, and generic reconnect endpoint preservation.
- UI typecheck, UI build, design-token gates, and Storybook build
passed.
- Live local browser: new Executor, Arcade, and Composio connections
completed provider consent from task cards. Each appeared Connected with
the requester selected.
- Real Test calls returned Executor output `4`, an Arcade public GitHub
star count, and Composio tool-discovery results. An ungranted agent was
denied access. Real Paperclip process-agent runs discovered each
provider catalog with only the requester’s connection installed.
- Deployed implementation commit `5d722e89c` to the isolated staging
tenant. The original failing Executor card now completes, discovers
seven actions, limits access to its requesting agent, and resumes that
agent. Its continuation completed real Executor calls and the provider
resume flow, then returned an upstream Airtable authorization link. A
real staging Test call returned `4` in 1.5 seconds. Later PR commits add
regression coverage and popup-unmount cleanup.
- Zapier fresh-token browser test remains pending a provider clipboard
handoff. Its URL and token flow passes focused tests.
- Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook
deployment and visual regression jobs skipped. Greptile 5/5, both review
threads resolved. Canceled runners and unrelated chat timeouts passed
the single retry on unchanged code.
- The full local test suite was not run, as requested by the maintainer.
CI runs the repository gates.

## Risks

- OAuth popup behavior differs by browser. The explicit sign-in link and
durable server completion checks provide recovery.
- Saved task drafts store only a connection ID in browser storage.
Credentials remain in the existing server vault.
- No database or server protocol changes.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code
execution, and browser tools. The exact context-window size is not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 09:19:18 -05:00
DottaandPaperclip 4721f55803 fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - MCP aggregator setup can return from provider consent to a saved
draft.
> - The branded setup did not explain failed or cancelled authorization.
> - It also offered identity changes that the server does not apply when
a saved connection resumes.
> - This pull request explains OAuth return outcomes and keeps the
displayed identity consistent with the saved policy.
> - Users can understand the outcome and retry the same connection.

## Linked Issues or Issue Description

Related: #13755 introduced the MCP aggregators. #13758 retired the
legacy Composio broker. #13584 proposes changes to the provider handoff
window; this fix retains the current handoff behavior.

**What happened?**

Cancelling Composio consent returned to the setup form with no
explanation. The same controller ignored failed OAuth callback outcomes.
Returning to Access on a saved draft also offered personal/shared
choices, although the server retains the saved identity. This could make
the OAuth request disagree with that identity.

**Expected behavior**

Explain cancellation or failure, preserve the draft, and offer Try
again. Display the retained credential identity and start OAuth with the
policy returned by the server.

**Steps to reproduce**

1. Enable MCP aggregators and start a Composio connection.
2. Continue to provider consent and cancel it.
3. Observe the return screen. Before this change, it showed the form
without cancellation feedback.
4. Go back to Access. Before this change, the form offered
personal/shared choices even though resume retains the original
identity.

**Paperclip version or commit**

Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`.
The original report of successful consent leaving setup unfinished did
not reproduce. This PR addresses the recovery defects observed during
that investigation.

**Deployment mode**

Isolated local development instance and authenticated staging
deployment, with real Composio consent and provider calls.

## What Changed

- Read the OAuth callback outcome in branded MCP setup and show
cancellation or failure feedback.
- Retry the same saved draft without displaying untrusted callback error
text.
- Keep saved personal/shared identity fixed in Access and select OAuth
identity from the returned credential policy.
- Add six focused regression cases and authorization-failure Storybook
states for Arcade, Composio, and Executor.
- Document return-screen recovery and retained identity in the connector
playbook.

## Verification

- `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth
return' --maxWorkers=1`: 6 passed, 143 skipped, including after
integration with current master.
- `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed.
- `pnpm check:token-gates`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm --filter @paperclipai/ui build-storybook`: passed for the
implementation commit.
- Browser recovery: cancelled real Composio consent, observed the new
feedback, returned to Access, retried the same personal draft, completed
consent, and ran a real tool call.
- Staging on implementation commit
`84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal
connections each completed on the first consent attempt and loaded 11
tools. Real discovery, execution, and schema calls succeeded. Both
connections stayed Connected after reload. A real agent used the shared
connection through the Paperclip gateway and returned the public
repository documentation hierarchy with one success and zero errors.
- All current-head CI gates passed on
`662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and
agent-chat browser shards each had an initial timeout; both passed on
one targeted rerun without code changes. Greptile reviewed this exact
head at 5/5, with no open review threads.
- No full local suite was run, as requested. CI provides the broader
checks. The PR adds a master merge and documentation after the
live-tested implementation commit.

## Risks

- This shared setup controller also serves Arcade and Executor. Their
callback rendering and retry behavior have focused test coverage; this
investigation used Composio for live provider testing.
- Saved identity remains fixed during resume. A different identity
requires a new connection, consistent with server behavior.
- No database, protocol, credential storage, or gateway policy changes.

## Model Used

OpenAI GPT-6 via Codex, with code editing, shell tools, and browser
testing. The exact runtime model ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 08:05:30 -05:00
DottaandPaperclip 7944ed3d97 fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path

> - Paperclip is the open source control plane for companies of AI
agents.
> - Native runner agents need governed tools, durable runtime state, and
useful execution evidence.
> - A first activity trace waited 53.467 seconds even though tool
activity took 6.274 seconds; provider input arrived before the server
API call executed.
> - Native agents also need a safe way to hire teammates without asking
the model to rebuild runtime configuration.
> - This pull request separates the observed ACP input-stream window
from the actual server `tool.execute` span and adds a server-owned
native hire contract.
> - The benefit is clearer latency evidence and safer native teammates
with existing approval, auth, and company boundaries preserved.

## Linked Issues or Issue Description

Related Daytona provenance work is in
[#13814](https://github.com/paperclipai/paperclip/pull/13814). No
duplicate public PR was found for this combined timing and native-hire
change.

**What existing behavior does this improve?**

Native runner agents can use governed tools and request hires. The
server did not expose a safe native hire operation that reused the
caller's validated runtime settings. First-activity traces also mixed
provider input timing with server tool execution timing.

**Current behavior**

A native hire must construct a separate runner configuration. Full
configuration copying could expose paths, instructions, secrets, or
sessions. Timing evidence could make a provider or MCP identity join
appear proven when the trace did not contain that join.

**Proposed behavior**

The native `hire_agent` operation accepts identity and persona inputs.
The server sends `adapterType: "paperclip_runner"` with
`inheritRuntimeFrom: "caller"`, then copies only validated provider,
model, permission, lifecycle, and bounded execution settings. It
inherits and validates the default environment, derives the managed AI
binding through existing normalization, preserves approval and
permissions, and creates fresh child instructions. Caller secrets,
paths, prompts, and sessions are excluded.

Provider events now include the optional boolean `inputUpdated`, with
Rust forwarding support. Timing evidence separately records the ACP
input-stream window and the actual server `tool.execute` activity. It
does not claim a provider or MCP join without matching evidence.

**Reason and benefit**

Native agents can hire teammates that start with the caller's approved
execution policy. Operators retain company boundaries, auth rules,
approval gates, and requalification. Reviewers can distinguish provider
streaming time from server API execution time when diagnosing
first-activity delays.

**Breaking changes**

None for existing hires or tool calls. `inheritRuntimeFrom` is optional
and only applies to same-company native agent callers. Conflicting
explicit runtime settings are rejected. The provider event field is
optional for existing producers.

## What Changed

- Added the native `hire_agent` protocol action, catalog entry, API
contract, and runner authority checks.
- Added `inheritRuntimeFrom: "caller"` validation and a closed native
runtime inheritance allowlist.
- Preserved managed AI binding normalization, default-environment
validation, approval snapshots, permissions, requalification, and fresh
child instructions.
- Added provider `inputUpdated` schema support and Rust forwarding.
- Added first-activity and server tool timing evidence with conservative
identity-join handling.
- Added route, authority, provider-event, sidecar, API, catalog, and
Rust-focused tests.
- Kept private Honeycomb links, raw traces, and local result paths out
of this description.

## Verification

Focused checks passed:

- 458 timing/session checks.
- 61 native hire inheritance checks.
- 20 hire authority checks.
- 1,741 API checks.
- 106 catalog checks.
- 54 provider sidecar checks.
- 12 Rust provider checks.

Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2).
R1 stopped at missing Docker image setup. The final trace is available
at
https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces.
Latest-head CI passed all required build, typecheck, Rust, static,
Vitest, serialized-server, workspace, chat, and E2E jobs. The focused
local checks listed above passed; the broad local suite was not run
before the live evaluation, while CI provides the full repository
verification.

## Risks

- Timing fields describe separate observed windows. They do not prove a
provider or MCP owner without a valid trace join.
- The inheritance allowlist must stay synchronized with native runner
configuration fields.
- Approval snapshots include resolved safe inherited settings and should
be reviewed when native configuration fields change.
- The focused local suite is narrower than the full repository suite;
latest-head CI covers the broader repository checks.

> Roadmap review: `ROADMAP.md` places this work within Paperclip's
bring-your-own-agent direction. It extends existing native runner hiring
and observability behavior.

## Model Used

OpenAI GPT-6 (exact serving model ID is not exposed), with extended
reasoning and repository tool use; GPT-5.6 Luna assisted with focused
implementation and verification work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 23:16:11 -05:00
Devin FoleyandPaperclip e4237c45f3 fix(ui): prevent organization title flicker during plugin loading (#13854)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The sidebar identifies the current organization.
> - An optional plugin can replace this navigation surface.
> - The built-in title appears before discovery and module loading
finish.
> - This pull request reserves the trigger until its owner is known.
> - The organization name appears once, while failures retain built-in
navigation.

## Linked Issues or Issue Description

Refs #13832. Searched related pull requests and issues; no duplicate fix
found.

**What happened?**
The organization switcher renders a provisional built-in title before an
installed replacement loads. Unrelated plugin imports can also affect
its loading state.

**Expected behavior**
Reserve the trigger with a neutral placeholder, then show the resolved
navigation surface. Keep the built-in menu on failed or absent
contributions.

**Steps to reproduce**
Install an organization-switcher contribution. Delay session, company,
contribution, and module responses. Reload the page and watch the
trigger through each stage.

## What Changed

- Reserve the trigger through account, company selection, slot
discovery, and module loading.
- Distinguish failed session lookup from pending lookup so errors retain
usable navigation.
- Load and await only contributions matching the requested slots.
Observe completion of imports started by another consumer.
- Document loading behavior and add regression coverage for loading,
failures, unrelated modules, and identity transitions.

## Verification

- `pnpm -r typecheck` passed, including Rust checks.
- `pnpm build` passed.
- All 629 UI test files passed: 6,593 tests. The 42 focused
UI/API/plugin tests also passed.
- `pnpm check:token-gates` and `git diff --check` passed.
- `pnpm test:run` was also attempted. The broad local server run was
stopped after recording skill-cache/channel fixture failures outside
this diff (for example, runtime skill source status `missing` instead of
`available`). The original cause is not established. All latest-head
Linux CI gates pass; the complete UI suite and affected local checks
pass.
- Desktop (1440px) and mobile (390px) Chromium checks passed with real
host components, dynamic module loading, and the built Account bundle.
Delayed fixture responses produced exactly two title states: empty
placeholder, then the resolved name. A slow refresh preserved the title
and trigger dimensions; absent/failed plugin fallback and Escape
dismissal passed, with zero uncaught browser errors. This is browser
component integration, not a live signed-in tenant test.

## Risks

A cold load displays a neutral placeholder until discovery completes.
Absent, ambiguous, failed, and invalid contributions still use the
built-in menu. No migrations or authorization changes. Scoped module
loading changes when an unrelated contribution is imported; each surface
loads its own matching modules.

## Model Used

OpenAI Codex, GPT-6, with reasoning, code execution, and browser
verification. The exact deployment ID and context window are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 20:33:29 -07:00
Devin FoleyandPaperclip cdf04a33fa feat(adapters): refresh current coding models and reasoning controls (#13829)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply model catalogs and reasoning controls to agent
setup.
> - Several provider releases are missing from the fallback catalogs.
> - Some newer models also have effort levels that the UI does not
offer.
> - Operators need the exact supported IDs and controls when discovery
is unavailable.
> - This pull request updates the existing adapters from current
provider documentation.
> - Operators can select current coding models without entering custom
IDs.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Model selection and reasoning controls across the existing coding-agent
adapters.

**Subsystem affected**

Claude, Codex, Grok, Gemini, Cursor, Kimi, and OpenCode adapters; model
discovery tests; agent creation and editing.

**Current behavior**

The catalogs omit Opus 5.5, GPT-6 Sol/Luna, Grok 4.7/4.6/4.5, current
Gemini Flash models, and several Cursor/Kimi choices. Bedrock has
obsolete IDs. The UI omits supported effort levels and saves Grok effort
under a key the runtime does not read.

**Proposed behavior**

Offer verified current model IDs and model-specific efforts. Remove
retired Gemini 2.0 choices. Keep configured defaults and saved model
IDs. Keep runtime discovery for account-specific choices.

**Reason and benefit**

Catch up with provider releases through September 22, 2026. Correct the
picker and runtime controls together.

**Breaking changes**

No database or API change. Gemini 2.0 options leave the picker after
their June 1 shutdown. Existing saved IDs remain unchanged. Corrected
Bedrock catalog IDs do not rewrite saved configuration.

**Additional context**

Supersedes the separate GPT-6 Sol PR #13830. Fable 5.1 was already
merged in #12730, and GPT-6 Astra in #12851. The Grok 4.6/4.5 proposal
#11324 was closed and parked by its author. This change retains the
default-sentinel fix from #12062. Related discovery proposals #13127 and
#13565 do not supply these catalog and effort updates. Searches found no
open PR for the additional model IDs.

See [the dated
audit](https://github.com/paperclipai/paperclip/blob/feat/claude-opus-5-5/doc/adapter-model-audit-2026-09-22.md)
for exact scope, primary sources, runtime observations, and
account-specific limits. This updates existing adapters and does not
duplicate planned core work.

## What Changed

- Add Opus 5.5 for direct Claude and Bedrock, with a Claude Code 2.1.280
gate. Correct and extend Bedrock model IDs.
- Add GPT-6 Sol/Luna and Fast mode. Offer Ultra for Astra/Sol and
GPT-5.6 Sol/Terra, and Max for both Luna generations.
- Add Grok 4.7/4.6/4.5, expose supported Extra High effort, and save
Grok edits under `reasoningEffort`.
- Add Gemini Flash 3.8/3.7/3.6/3.5, Flash Lite 3.5/3.1, and 3 Flash
Preview. Remove retired 2.0 choices.
- Add the current documented Cursor fallback models, including Fable
5.1, Composer 2.5, and Muse Spark 1.3.
- Refresh OpenCode fallback IDs used in remote environments from its
installed provider registry.
- Add Kimi K3 256K. Update the existing coding alias to K2.8 Preview and
enable its CLI effort settings.
- Use model-specific Claude/Grok efforts in creation and editing. Clear
unsupported effort when switching models.
- Add catalog, CLI/ACP forwarding, compatibility, and UI persistence
coverage. Record the audit and sources.

## Verification

- Latest head `6e63c9ef53b54ba869cd4fb431a8570bebe289f4`: 53 CI checks
passed, 2 skipped. This includes full workspace typecheck, build, and
all test shards. Greptile is 5/5 with zero unresolved threads. GitHub
reports no merge conflicts.
- 340 focused tests passed across adapter metadata, CLI/ACP arguments,
Claude version checks, Kimi effort, Grok execution, server model
discovery, and UI effort selection/persistence.
- `pnpm --filter @paperclipai/adapter-claude-local --filter
@paperclipai/adapter-codex-local --filter
@paperclipai/adapter-grok-local --filter
@paperclipai/adapter-gemini-local --filter
@paperclipai/adapter-kimi-local --filter
@paperclipai/adapter-cursor-local --filter
@paperclipai/adapter-opencode-local typecheck` — passed. The same
filters with `build` passed.
- `pnpm check:token-gates` and `git diff --check` — passed.
- Full workspace and UI typechecks were attempted locally. They stop on
existing missing `three` dependencies in `packages/shared/src/cliplab`.
- Full `pnpm test:run` and `pnpm build` were not run locally. Worktree
creation exhausted disk space, so a clean dependency install is not
feasible on this host. Focused checks reuse existing dependencies. CI
supplies full workspace verification.
- No provider inference was run. Account-specific runtime model lists
were inspected where available.
- Manual check: select the new models in agent setup and editing.
Confirm Luna has Max but no Ultra, Grok 4.7 has Extra High, and Fable
5.1 has Extra High/Max. Save Grok effort and confirm
`adapterConfig.reasoningEffort` contains the selection.

## Risks

- Catalog presence does not grant account access. Older CLIs and
restricted accounts can reject a model. Opus 5.5 has an explicit upgrade
check.
- Higher effort can increase cost and latency. Existing agent defaults
are unchanged.
- Cursor fallback IDs come from public model documentation; the local
account exposed no live catalog. Runtime discovery still adds
account-specific variants.
- Kimi effort remains supported only on its explicit CLI engine. This
does not add effort support to its default ACP engine.
- Saved obsolete Bedrock or retired Gemini IDs are not migrated
automatically.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed to
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:56:00 -07:00
Devin FoleyandPaperclip 7badae6981 feat(plugins): add an optional organization switcher slot (#13832)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins can add UI surfaces to the board.
> - Organization navigation is still fixed in the host sidebar.
> - A distribution needs a supported way to supply its own organization
menu.
> - This pull request adds one optional React slot with host-owned
navigation controls.
> - The built-in menu stays available when the optional contribution
cannot render.

## Linked Issues or Issue Description

**Subsystem affected**

Plugin SDK, server capability validation, and sidebar UI.

**Problem or motivation**

An installed plugin cannot replace the organization switcher without
editing the host menu. Existing sidebar and overlay slots do not provide
this replacement surface.

**Proposed solution**

Add an `organizationSwitcher` slot that requires `ui.sidebar.register`.
Pass display state, an icon renderer, and navigation/logout callbacks.
Keep the built-in menu for absent, ambiguous, missing, failed, or
unsupported contributions.

**Roadmap alignment**

This keeps distribution UI in plugins and adds a small host contract. It
does not add an account system or change company authorization. The
maintainer requested this extension.

Related prior menu changes: #12788, #10917, and #10850. No duplicate
replacement-slot PR was found.

## What Changed

- Add the slot to shared validation, SDK types, and server capability
checks.
- Wrap both sidebar menu variants with the optional replacement.
- Resolve selection against the current account query before mounting,
and reset replacement state on account or company changes.
- Add host-specific component props and a fallback to `PluginSlotMount`.
- Document the React-only contract and its trust boundary.
- Report runtime-supervisor fixture startup details when CI readiness
fails.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- `pnpm check:token-gates` passed.
- `pnpm exec vitest run --project @paperclipai/ui`: 6,580 tests passed.
- Targeted manifest, replacement, and built-in menu tests: 26 passed,
including the incoming-account selection regression.
- Installed a local test contribution into an isolated server. CLI
inspection reported `ready`. Browser checks covered the loaded
production UI, keyboard dismissal, current-organization selection, and
an expired remote session. Remote account responses were fixtures.
- `pnpm test:run` was attempted. Its first server phase passed 8,323
tests but failed in 36 files due to embedded PostgreSQL startup and
filesystem permission errors on this Mac. Later phases did not run. The
current Linux CI run is green: 54 checks passed and two were skipped,
including build, typecheck, and browser gates. See
https://github.com/paperclipai/paperclip/actions/runs/35796722770.
- The earlier runtime-supervisor readiness failure did not reproduce
locally. The complete affected shard passed locally: 58 files and 843
tests. The six supervisor tests also passed on Node 24.21.0 with CI
flags. Added fixture startup diagnostics for the selected Node
executable and listener port. The affected Linux shard then passed all
843 tests. The original root cause remains unconfirmed; no production
runtime behavior or timeout was changed.

## Risks

- Plugin UI remains trusted same-origin code. Display props do not
authorize account requests.
- A replacement can change navigation behavior. The host retains the
built-in menu when discovery or rendering fails and keeps logout/session
cleanup host-owned.
- No database migration. Existing menus and portfolio behavior remain
available.
- Local full-suite verification is limited by the environment failures
listed above.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser testing. The runtime does not expose a more
specific model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:39:11 -07:00
DottaandPaperclip a10702a878 feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people start and continue agent tasks from other
services.
> - A Slack conversation needs access to its surrounding discussion and
Slack collaboration tools.
> - The agent must use the linked requester's access and keep private
material within its permitted audience.
> - This pull request adds Slack tools through the existing connector
contribution and approval framework.
> - People can ask an invited bot to read a discussion, create follow-up
tasks, and collaborate in Slack.

## Linked Issues or Issue Description

**Subsystem affected**

Chat connectors, connector runtime, tool gateway, and connection
Settings/Access.

**Problem or motivation**

Slack-origin tasks can receive messages but cannot inspect the rest of a
channel or act through the originating bot. People must paste context or
configure a separate integration.

**Proposed solution**

Supply typed Slack tools and a bundled skill only to the originating
task and assigned agent. Resolve the linked requester on the server.
Check bot and requester access before reads and writes. Use existing
durable actions and approvals. Retrieved messages remain source
material.

**Alternatives considered**

Slack's user-OAuth MCP server does not replace the customer-created chat
bot. An unrestricted Web API proxy would not provide suitable permission
or publication boundaries.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps work with a provider
contribution. It does not add a task dispatcher or a separate Slack task
lifecycle. Related: #11144 covers generic per-user MCP grant execution;
this change binds Slack bot operations to chat-origin tasks.

## What Changed

- Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and
shared native/HTTP execution.
- Bind tools to company, endpoint, task, run, assigned agent, and
admitted linked requester. Check membership and revocation on each call
and before queued writes.
- Add paginated reads, bounded history search, source links, messages,
file uploads, reactions, pins, bookmarks, topics, canvases, lists, and
approved channel operations.
- Restrict private-source publication, including automatic replies and
uploaded deliverables. Keep other people's bot DMs inaccessible.
- Reuse action receipts, idempotency, approvals, and reconciliation.
Suppress an identical explicit-send/final-reply duplicate. Return
governed results through their verified originating conversation.
- Add endpoint-bound personal search OAuth storage and lifecycle. Keep
native real-time search disabled until a runtime meets Slack's
transient-result requirements. Current runtimes use bounded history
search.
- Show capabilities, scope upgrades, and personal search authorization
in Settings/Access and Storybook. Document provider and runtime limits.

## Verification

- Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no
unresolved findings. One unchanged rapid-callback timing test passed on
a single CI retry.
- Approval presentation regressions cover board-comment precedence and
exact Slack publication; the expanded database assertion passed in CI.
The local PostgreSQL startup probe later became unavailable, so that
final assertion was verified in CI. Slack setup and failed-run retry
browser tests also passed locally.

- Full workspace typecheck and build passed. Server typecheck/build
passed again after the approval routing fix.
- Broad local suites passed in separate groups: server 12,958 tests, UI
6,555, shared 770, skills catalog 20, and other workspace packages
2,652. CLI and serialized server checks passed after environment/timeout
retries. These are composite results, not one uninterrupted green
full-suite invocation.
- PostgreSQL authority regression covers admitted identity,
cross-company/task/agent rejection, recovery, retained-session
revocation, OAuth refresh/disconnect races, approval execution, exact
publication lineage, retries, uncertain sends, and duplicate
suppression.
- Gateway/response regressions cover separate-origin approval batches
and durable continuation. Focused provider, access, search, native
runtime, route, and AgentMail regressions pass.
- Storybook capability, missing-scope, OAuth configuration,
authorization, and disconnect states were inspected in the browser.
- Live staging: read a channel decision and full thread, create exactly
two assigned backlog tasks, add a reaction, paginate discovery to
exhaustion, and return bounded search matches with source links and
coverage.
- Live staging: create/edit/read a canvas and list, inspect the canvas
in Slack, post/edit one message, and create a channel only after
approval. New channels remain disabled for responses.
- Live staging: read a response-disabled channel from the requester's
DM; writes to that channel were denied. The test setting was restored.
- Final live retest passed: explicit file upload and exact content
read-back; approved deletion of only the disposable bot message;
continuation confirmation returned to the original Slack thread without
repeating the action.
- Optional OAuth, private multi-user boundaries, native RTS, and CLI
provider execution are not fully live-qualified. The staging agent
initially supplied malformed tool arguments; valid arguments succeeded,
and the tool/skill descriptions now emphasize UUID write keys.

## Risks

- Existing Slack apps must add scopes and reinstall for new
capabilities. Provider plans and document permissions can still restrict
operations.
- Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET`
for governed tool actions. The staging instance was configured with
explicit operator approval; fleet provisioning is a separate gap.
- Native RTS is not exposed on current transcript-retaining runtimes.
Bounded history scans are deliberately reported as incomplete. Inline
file reads support text/canvas content up to 256 KiB; other types return
metadata.
- Private document edits fail closed when the full audience cannot be
verified. Uncertain effects other than posts/uploads require inspection
instead of blind retries.
- Shared approval-delivery code now separates outcomes by source run to
preserve origin boundaries. No database migration is required.
- A separate completion-validator gap remains when the agent cites a
prior run's registered artifact during finalization. It asked for
registration again even though Slack delivery was confirmed. This change
does not add a connector-specific task-completion policy.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
browser testing. The exact deployed model identifier and context-window
size were not exposed in the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:29:07 -05:00
DottaandPaperclip a959e47508 fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents controlled access to external services.
> - Google Chat uses OAuth profiles and reviewed MCP tools.
> - Those profiles request membership and read-state access that the
supported feature set does not need.
> - Removing read-state access also requires us to block unread search
filters, including on existing connections.
> - This pull request reduces both OAuth methods and enforces the
reduced search contract before dispatch.
> - Users retain conversation lookup, message history, ordinary search,
and approved message sending.

## Linked Issues or Issue Description

**What happened?**

Both Google Chat profiles request membership and read-state scopes. The
supported tool set does not include membership listing or read-state
updates. Message search still advertises an unread filter. Related work:
Refs #12619.

**Expected behavior**

Managed and customer-owned OAuth request only the scopes needed for
supported features. Unsupported unread filters fail clearly before any
provider call. Existing cached catalogs and broader grants must not
bypass that policy.

**Steps to reproduce**

Start Google Chat OAuth from either connection method and inspect the
requested scopes. Inspect the message-search tool schema, then submit a
search with `searchParameters.isUnread` set to true or false.

**Paperclip version or commit**

The scope change is based on master at `110d176fc`.

**Deployment mode**

Managed Cloud and self-hosted instances with Google Chat Apps enabled.

## What Changed

- Remove `chat.memberships.readonly` and `chat.users.readstate.readonly`
from shared profiles and all four Chat connection methods.
- Hide unsupported read-state fields and instructions in agent and board
Test tool schemas.
- Reject explicit unread filters, including false, null, snake-case
fields, and encoded filter objects, before provider dispatch.
- Recheck previously approved calls and support existing profile-bound
and URL-only Chat connections.
- Add scope, signed broker request, OAuth URL, allowlist, schema, and
dispatch regression tests.
- Document coordinated app/broker rollout, existing-grant reconnects,
and the remaining deployment checks.

## Verification

- All seven focused OAuth and Chat gateway test files pass: 498 tests on
the rebased branch.
- `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4.
- The complete sharded CI test matrix passes, including general server,
Chat, workspace, serialized server, Runner, and all eight browser e2e
shards. The duplicate unsharded local `pnpm test:run` was stopped after
CI passed; it did not complete locally.
- `git diff --check` passes.
- `node scripts/ingest-app-definitions.mjs` succeeds and leaves the
branch unchanged. Google Workspace JSON is the durable reviewed input
used by the generator.
- All current-head CI checks pass at
`7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build,
and canary dry run. Greptile is 5/5 with zero unresolved threads. The
generator concern was withdrawn after review of the source and
regeneration evidence.
- No production deployment or live Google consent test was performed.
After coordinated deployment, verify reduced consent scopes, normal
search/history, message sending, and rejection of unread filters.

## Risks

- Coordinate deployment with the companion Cloud broker scope change.
Mixed versions can reject exact-scope requests.
- Existing tokens are not narrowed or revoked. Grants with old scopes
need new consent. Do not revoke a shared Google client to migrate one
profile.
- Explicit unread filters now return an error instead of being sent to
Google. Ordinary search and the approved send tool remain available.
- No database, UI, lockfile, or workflow changes.

## Model Used

OpenAI Codex, a GPT-5-based coding agent, with tool use and code
execution. The exact runtime model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:12:30 -05:00
DottaandPaperclip 1ccae464c5 feat(ui): add task artifact media gallery and full-row links (#13825)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks collect the files and work products that agents create.
> - The Artifacts tab shows these outputs in rows, which makes videos
hard to compare.
> - Small text links also make artifact rows harder to open.
> - This pull request adds image and video tiles with previews and makes
the full artifact row clickable.
> - Users can compare outputs and open the existing media viewer with
one click.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The task Artifacts tab and media previews in task chat.

**Current behavior**

Video outputs appear as file rows or icons. Users must click a small
link to open a work product. A task with eight video outputs gives
little visual context.

**Proposed behavior**

Show images and videos in a responsive gallery. Show a paused video
frame as the thumbnail. Keep documents, links, and other files in rows
whose entire area opens the item.

**Reason and benefit**

Users can compare generated media without opening each item. Larger
click targets also make the sidebar easier to use.

**Breaking changes**

None. This uses the existing artifact URLs, media viewer, run grouping,
and attachment filters. No API or database changes.

Related work: #11226 added the task sidebar output surface, and #7361
added rich attachment previews. #3524 concerns a separate
reviewed-assets panel. This PR improves the existing task artifact
components. The duplicate search found no active PR for this change.
This is polish for the shipped Artifacts & Work Products roadmap item.

## What Changed

- Add a shared media tile for work products and agent attachments.
- Reuse video and image previews in task artifacts and chat. Seek up to
one second into videos and reset preview state when the source changes.
- Use the existing task gallery for playback and downloads. Preserve
grouping and attachment deduplication.
- Extend native links and buttons across work-product rows, including
keyboard focus indicators.
- Add eight offline Storybook examples for video outputs, mixed media,
clickable rows, narrow and wide panels, missing previews, empty state,
and light mode.
- Register the component in the design guide and document its use.

## Verification

- 108 focused component tests pass, including thumbnail seeking, source
changes, gallery activation, and attachment deduplication.
- `pnpm build`, `pnpm -r typecheck`, Storybook build, token gates, and
`git diff --check` pass locally.
- Reviewed the production components in the embedded browser. Checked
all eight video thumbnails, mixed media, narrow layout, light mode,
blank-area row clicks, keyboard gallery activation for generic-MIME
images, and playback from chat video thumbnails.
- Storybook: open **Tasks / Artifact Gallery** and select **Eight Video
Outputs**, **Mixed Media And Files**, or **Whole Row Clickable**. The
small local clips are synthetic fixtures.
- All build, typecheck, unit, runner, and end-to-end CI jobs pass for
`b8ace289b5e07df5b9f2c319b159f3923ce427b4`. Greptile gives 5/5 with both
review findings resolved. All 54 PR checks pass, including the external
security scan. The full test suite passed in CI. The duplicate serial
local test run was stopped after CI finished; the 108 focused tests,
full build, and recursive typecheck passed locally.

## Risks

- Video thumbnails require the browser to load metadata and a frame. A
slow server or unsupported codec can leave the fallback visible; opening
and downloading still use the existing viewer.
- Full-row click targets change pointer interaction with work-product
cards. Native link and button semantics remain in place.

## Model Used

OpenAI GPT-6 in Codex, with code execution and browser tools. The exact
deployment ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:03:26 -05:00
Devin FoleyandPaperclip 92d4868e79 fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.

## Linked Issues or Issue Description

Refs #13446 and #13719.

**What happened?**

After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.

**Expected behavior**

Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.

**Steps to reproduce**

1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.

## What Changed

- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.

## Verification

- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed c9db03bcab at 5/5. Its only thread is resolved. Full
PR CI passed on that same head:
https://github.com/paperclipai/paperclip/actions/runs/35774449020

## Risks

Small change to error attribution. Unrelated errors may now form their
correct Sentry groups instead of reopening a prior run group. No errors
are filtered or suppressed. No tracing is enabled and no new event
fields are added. No schema or runtime-execution changes. HTTP logs
retain their request and status diagnostics while masking the capability
value. The new SDK job has read-only permissions, no secrets, and an
in-memory Sentry transport.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 12:53:31 -07:00
Devin FoleyandPaperclip 5f1100e3b3 refactor(server): remove retired operator UI snippet injection (#13789)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can extend its interface through trusted plugin UI
contributions.
> - The server also accepts executable HTML through two legacy
environment settings.
> - This older path bypasses the plugin installation and lifecycle
model.
> - This pull request removes snippet injection from static and
development pages.
> - Operators must migrate existing integrations before upgrading.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Retire the operator HTML injection path from the server. Related public
changes: #13168, #13245 and #13496 introduced the legacy settings;
#13646 supplies the generic plugin host contract.

**Current behavior**

Managed instances append operator-supplied HTML or a decoded script body
to every page. The same integration can use supported trusted plugin UI
slots.

**Proposed behavior**

Ignore both retired settings. Serve normal branded static and
development HTML, and keep plugin contributions unchanged.

**Breaking changes**

Installations that rely on `PAPERCLIP_CLOUD_UI_SNIPPET` or
`PAPERCLIP_CLOUD_UI_SNIPPET_B64` must migrate before upgrading. The
maintainer-owned staging and production deployments have completed the
migration prerequisite. Other operators must migrate their integrations
before adopting this change.

## What Changed

- Remove the snippet injector and its static/dev rendering integration.
- Remove injector-specific tests and retain a regression that old
settings no longer change served HTML.
- Replace setup instructions with a retirement and plugin-migration
note.

## Verification

- Rebased onto current master; `pnpm exec vitest run
server/src/__tests__/static-index-html.test.ts
server/src/__tests__/vite-html-renderer.test.ts`: 5 tests passed.
- With repository-pinned Rust/Cargo installed, full `pnpm -r typecheck`
and `pnpm build` pass after rebase.
- Previous full local `pnpm test:run` encountered unrelated macOS
runtime-cache rename `EACCES` errors and a missing AgentMail skill path;
a focused reproduction confirmed 5 failures / 95 passes. That is not a
passing full-suite result. The unchanged focused suites, build and
typecheck were repeated after rebase; the full local suite was not
repeated. All 54 refreshed GitHub checks/contexts passed on
`bd62bff63f9d7980bfd10e54cb0102693d27ecfb`; fresh Greptile is 5/5 with
no unresolved threads.
- No UI component styling, database or API contract changed.
- Maintainer approved the remaining rollout and cleanup. Staging snippet
retirement and sleep/wake verification are complete. Production
migration and removal of the legacy settings are complete for serving
tenant instances. Remaining old warm inventory is excluded from new
signups until configuration reconciliation completes. The maintainer
authorized upgrading the remaining old deployments and clearing their
pins.

## Risks

- Removing the settings disables integrations that still depend on them;
operators outside the completed maintainer rollout must migrate before
upgrading. This is an intentional behavior change, documented at the
existing setup-doc path.
- Keep a previous image and its configuration for rollback. Existing
browser tabs need a refresh to unload already-injected code.
- Plugin UI remains trusted same-origin code. This does not add a
security sandbox or change ordinary branding.

## Model Used

- OpenAI GPT-6 (Codex; exact deployment variant and context-window size
are not exposed in this session). Reasoning, repository inspection,
local code execution and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused rendering tests;
unrelated local full-suite failures are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All refreshed checks pass on
`bd62bff63f9d7980bfd10e54cb0102693d27ecfb`
- [x] Fresh Greptile is 5/5 on the current head, with no unresolved
findings
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 11:18:46 -07:00
DottaandPaperclip 3d78e3a4ec fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The native Runner keeps a live provider process between task turns.
> - Managed GitHub access used a token tied to one run.
> - A new run forced Paperclip to replace that process to replace its
token.
> - This PR gives the session a stable credential transport and binds
each operation to the active run.
> - The agent can keep its process while Paperclip checks current
identity and grants.

## Linked Issues or Issue Description

Follow-up to #13738. Related credential-rotation work: #11770 and #8208
use process replacement for other adapter credentials; this change
applies to managed GitHub access in the native Runner.

**What happened?**

A configured GitHub connection forced a warm native provider process to
close at each new run. The saved conversation survived, but the live
process did not.

**Expected behavior**

Keep the warm provider process. Resolve GitHub access for the current
run when each command starts. Deny access while idle or after the run
ends.

**Steps to reproduce**

1. Configure managed GitHub access for a native Runner agent with a warm
session.
2. Complete a turn, then send another message to the same task.
3. Observe the provider process close with the reason `warm native
session configuration changed`.

**Paperclip version or commit**

Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before
submission.

**Deployment mode**

Local and remote native execution, including the sandbox callback
bridge.

## What Changed

- Move configured native GitHub transport and launcher ownership from
the run to the provider session.
- Bind the broker only after the executor acquires session ownership.
Clear that binding when the run exits.
- Keep the shared live-run, identity, grant, and trust-policy checks for
each credential request.
- Reject wrong scopes, idle requests, and credential responses that
arrive after their run binding changes.
- Retire transport and launcher files with the provider session. Keep
anonymous commands available if bridge startup fails.
- Add red/green executor tests, real subprocess and callback-bridge
tests, and database checks. Update the runtime documentation.

## Verification

- Before the fix, both new local and remote warm-session reuse tests
failed.
- After the fix, 435 targeted tests passed across the executor, broker,
launcher, token, and database suites.
- A real long-lived test process kept the same PID and original
environment across two runs, including through the production callback
bridge on local test processes.
- Server typecheck and TypeScript compilation passed.
- Full workspace typecheck and build passed. Server typecheck passed
again after the review fix.
- The fallback-logging regression failed before the fix; all 9 broker
tests pass afterward.
- The exact chat sidebar browser scenario passed locally. The initial CI
timeout showed failed Vite module downloads; all eight browser shards
pass on the latest commit.
- All 53 latest-head checks passed, including the full CI test matrix
and security checks (two unrelated conditional checks skipped).
- The duplicate full local test run was stopped after CI passed; it is
not claimed as a completed local pass. Targeted local tests, workspace
typecheck/build, and the browser scenario passed.
- Greptile reviewed the latest commit at 5/5 with no unresolved
findings.
- No fresh paid provider or Daytona campaign has run for this change.

## Risks

- The broker now lives as long as the provider session. Tests cover idle
denial, late cleanup, late responses, shutdown, and failed startup.
- Its in-memory authority does not survive a controller restart.
Existing checkpoint and process-recovery rules still apply.
- Raw GitHub credentials remain confined to individual command
processes. The session transport token cannot select a different task,
agent, company, or run.
- No database migration or public API change.

## Model Used

OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution.
The exact serving model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:07:13 -05:00
Devin FoleyandPaperclip 110d176fc9 fix: preserve chat message bindings in review recovery (#13818)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - External chat messages can start work on tasks that remain in review.
> - Paperclip queues one recovery run when a task loses its review path.
> - That recovery keeps the chat source but loses the admitted message IDs.
> - The authorization check then rejects the recovery before execution starts.
> - This change retains the message IDs so the existing check can verify current access.

## Linked Issues or Issue Description

Refs #13809.

**What happened?**

A successful external-chat run can leave an active task in review without a maintained review path. Its automatic recovery then fails with `reviewed_chat_execution_binding_not_authorized`. The recovery context retains `chat:slack` but drops `wakeCommentIds`.

**Expected behavior**

An eligible recovery should retain its admitted message references and pass a new authorization check. Missing or revoked access must still prevent execution.

**Steps to reproduce**

1. Finish a chat run whose task remains in review with no maintained review path.
2. Build the bounded review recovery from that run's context.
3. Dispatch the recovery through the reviewed-chat authorization check.

The new regression tests fail before this patch. Related PR #13809 handles answered conversations that become idle. This patch handles recovery when the conversation remains active.

## What Changed

- Retain the admitted message batch for external-chat review recovery. Derive the current comment reference from that batch.
- Share the existing supported-provider selector between recovery and run-bound chat authorization. AgentMail remains on its separate email inbox path.
- Keep the existing authorization check. Do not copy prior checkout, authorization, session, or prompt state.
- Test Slack and Discord recovery, missing batches, revoked access, and changed task or company bindings.
- Document the recovery authorization contract.

## Verification

- The new tests reproduced the missing-message failure before the fix.
- Six focused suites passed: 257 tests covering review recovery, reviewed-chat authorization, issue liveness, Slack lifecycle, comment-wake batching, and external-chat waits. PostgreSQL integration tests ran against a temporary local PostgreSQL database.
- Focused test files: `review-path-recovery.test.ts`, `heartbeat-reviewed-chat-binding.integration.test.ts`, `heartbeat-issue-liveness-escalation.test.ts`, `slack-conversation-lifecycle.test.ts`, `heartbeat-comment-wake-batching.test.ts`, and `external-chat-wait.integration.test.ts`.
- Server TypeScript check passed with scratch configuration that resolves this checkout's workspace packages. Existing dependency links point to another checkout; the default check reports stale shared-type errors.
- `node scripts/check-module-boundaries.mjs` and `git diff --check` passed.
- `git diff | gitleaks stdin --redact --no-banner` passed. The diff was also checked for private identifiers and user data.
- [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35760757895) passed on the latest commit, including build, workspace typecheck, general and serialized tests, Rust checks, and all eight browser shards. The unchanged local-service readiness test and agent-chat page-load assertion passed when their failed shards were retried. The local-service suite also passed locally (6 tests). The first, superseded run lost a chat runner; its stuck browser job was cancelled to unblock the current run.
- Full local workspace typecheck, tests, and build were not run. The machine has about 2 GiB free, and those commands include Rust builds. The full PR CI checks passed before merge.

## Risks

Small context-construction change. Dispatch still checks current execution ownership and access for every admitted message. The recovery remains bounded to one attempt per consumed path. No schema changes. Existing failed runs are not retried by this patch.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:46:13 -07:00
DottaandPaperclip 8725d6ce09 fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions.

Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:35:33 -05:00
Devin FoleyandPaperclip e3d8fb0876 feat: attest standard production images at full source commits (#13797)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators deploy its standard production container on several CPU
architectures.
> - Downstream image builders need to identify the exact source of their
base image.
> - A short commit tag does not provide signed source evidence.
> - This pull request adds a full commit tag and signed image digest for
canonical master pushes.
> - Consumers can verify the source and compose from the immutable
digest.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Publication of the standard multi-platform production image.

**Current behavior**

The Docker workflow publishes short commit tags and channel tags. It
does not provide a signed standard-image contract tied to the complete
master commit.

**Proposed behavior**

Canonical master pushes also publish `sha-<full-commit>` and attest the
exact index digest after platform validation and an immutable-image
orphan-reaping check. The signer certificate binds the source
repository, commit, workflow and ref. Existing tags and the separate
cloud producer remain available.

**Reason and benefit**

Downstream builders can prove the source of a standard base without
adding their dependencies or repository details to the public workflow.
No matching open issue or duplicate PR was found.

## What Changed

- Add the canonical full-SHA tag without changing existing tag mappings.
- Validate amd64 and arm64 descriptors and hash the exact registry
response bytes and require its digest header to match.
- Verify the immutable image and sign it with GitHub artifact
attestations.
- Run the contract tests in trusted PR verification and document the
consumer contract.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/standard-image-contract.test.mjs`: 37 passed.
- A read-only check against an existing published index returned its
exact expected digest.
- `actionlint -shellcheck='' .github/workflows/docker.yml
.github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports
only existing `ls` and word-splitting warnings.
- `pnpm build`: passed locally with Cargo available.
- `pnpm -r typecheck`: passed locally.
- Full local Vitest was attempted: 8,260 passed, 14 failed, with 34
failing suites. The failures were missing embedded-PostgreSQL library
aliases in this fresh install and existing macOS runtime-skill-cache
rename errors. Native aliases are now restored. Rerunning the 33
affected database suites produced 577 passes and two unrelated AgentMail
skill-root lookup failures (32 suites passed). The four directly failing
database tests also pass independently. This is not a claim that the
full local suite passed.
- All final-head CI checks pass. One unrelated routine-route mock
assertion passed on the single-shard retry; its 15 tests also pass
locally. Greptile is 5/5 on this exact head, with no unresolved threads.
- Actual signing requires a canonical master push. This draft PR does
not publish trusted provenance.

## Risks

The new attestation step requires OIDC and attestation write permissions
in the merge job. Signing failure leaves the image available but without
the new admission proof. Consumers must fail closed when proof is
missing. Existing release tags, the legacy producer, and image retention
remain unchanged. No database or application behavior changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell
execution and test tools. The session does not expose a more specific
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 07:33:30 -07:00
Devin FoleyandPaperclip 8326e33ada Fix oversized sandbox process launch payloads (#13793)
## Thinking Path

> - Paperclip manages agent work and preserves context across retries.
> - Sandbox ACP runs encode their command and environment into one
launch value.
> - A long continuation can make that encoded value exceed Linux's exec
limit.
> - The launch shell then exits before the agent can initialize.
> - This PR transfers large command envelopes through a private
temporary file.
> - The agent receives the complete environment and can start normally.

## Linked Issues or Issue Description

Refs #13777.

**Bug description**

Sandbox tasks with long retry context fail during ACP initialization
with exit 127. The streamed bridge also drops the shell error that
explains the failure.

**Steps to reproduce**

Launch the streamed sandbox process bridge on Linux with one valid
110,000-byte environment value. Its base64 command envelope exceeds the
limit on one exec argument or environment string. Daytona reports
`argument list too long: env`, then exit 127.

**Expected behavior**

The command envelope must not make a valid child environment too large
to launch. A shell startup failure must retain its diagnostic in the run
log.

## What Changed

- Keep envelopes up to 64 KiB on the existing launch path. Upload larger
envelopes in bounded chunks inside a mode-0700 session directory. Set
the final payload file to mode 0600.
- Read the file without following symlinks and delete it before spawning
the child. Remove incomplete uploads on failure. Both streamed and
polled bridges use the same envelope.
- Preserve stderr when the launch shell fails before the wrapper emits a
terminal event. Emit a fixed terminal error and shutdown acknowledgement
if a payload cannot be read or parsed, without exposing its contents.
- Cover large environments, file permissions, payload deletion,
interrupted uploads, missing or malformed payloads, and startup
diagnostics. Document the transfer and cleanup behavior.

## Verification

- Both large-envelope regressions fail before the fix and pass after it.
- Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass:
370 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- A gated live Daytona probe on the current sandbox image reproduced
exit 127 with the old bridge. The fixed bridge launched the same command
successfully. A second probe completed real Claude ACP initialization
with a 110,000-byte context value. It did not run an agent task. All
temporary sandboxes were deleted.
- `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner
step and stop because this machine has no `cargo` executable.
- Full CI passes on `ea816cf58d`: [run
35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754).
All 53 checks pass; two optional checks are skipped. The local
full-suite run was stopped after equivalent CI suites passed; it has no
final local result.
- Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads.
The branch is mergeable.

## Risks

Large envelopes require extra upload calls during startup. The temporary
data stays inside the private session directory and is removed before
child startup or during failure cleanup. Individual child environment
values still obey the operating system's native limits. No migration or
configuration change is required.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 20:54:07 -07:00
Devin FoleyandPaperclip c496acb570 Return client errors for known OAuth reconnect states (#13794)
Map missing OAuth refresh credentials and terminal reauthorization to HTTP 422. Preserve reconnect instructions and reporting of unexpected provider failures.

All 363 focused tests and server typecheck pass. Required CI and review checks passed; Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:50:01 -07:00
Devin FoleyandPaperclip a7d3b17a97 Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup
instructions from catalog and health routes. Bound response parsing and keep
unknown upstream errors reportable without exposing provider settings links.

Verified 361 focused tests, server typecheck, and authenticated discovery.
The three route regressions fail before this change and pass afterward.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:25:06 -07:00
Devin FoleyandPaperclip ded156a904 Keep interaction continuations scoped to the target task
Separate an interaction's producer from explicit resume history. Do not
import another task's comments or results. Filter newly captured foreign
origins and recover older inherited origins only when the saved producer
context and comment row prove their source. Keep missing context and
company boundaries fail-closed.

Verified 46 continuation tests, 201 related recovery tests, and server
typecheck. Added 20 database regressions for provenance and scope guards.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:06:05 -07:00
Devin FoleyandPaperclip 878734a061 fix(tools): treat OAuth sign-in challenges as client errors (#13786)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connected apps can require OAuth sign-in before they list their
tools.
> - Remote discovery recognizes this condition as `oauth_challenge`.
> - Discovery and catalog refresh currently return it as HTTP 502.
> - The server error handler reports that response as a crash.
> - This pull request returns HTTP 422 for the explicit sign-in
challenge.
> - The operator keeps the sign-in instructions, while unexpected
upstream failures remain reportable.

## Linked Issues or Issue Description

**What happened?**

Connecting a remote MCP app that answers with a recognized OAuth
challenge returns HTTP 502. Reading an empty catalog or explicitly
refreshing it does the same. The server error handler then sends the
expected sign-in condition to error monitoring.

**Expected behavior**

A known sign-in requirement returns HTTP 422 with the existing
`oauth_challenge` code, message, and setup/reconnect links. An
unexplained upstream HTTP 400 or an unavailable upstream service still
returns 502 and reaches error monitoring.

**Steps to reproduce**

1. Configure a remote MCP app that returns HTTP 401 with a Bearer
challenge.
2. Connect the app, read its empty catalog, or request a catalog
refresh.
3. Observe HTTP 502 and a server error report before this change.

**Paperclip version or commit**

Reproduced on `6de50ba594b15efaa3eae6ed869cd39b3a436456`.

**Deployment mode**

Server with remote MCP connections. The regression coverage uses local
PostgreSQL and mocked upstream HTTP responses.

Related: #9750 addresses MCP initialization and session recovery. It
does not change the classification of this recognized sign-in condition.
Targeted searches found no duplicate classification PR.

## What Changed

- Return 422 for `oauth_challenge` from discovery and from catalog
health-error normalization.
- Preserve the existing structured error and remediation links.
- Test automatic empty-catalog reads and explicit refreshes. Verify that
OAuth challenges produce no Sentry capture and that upstream 400/503
failures still do.
- Update the direct-connect and blocked-redirect expectations and
document the monitoring behavior.

## Verification

- Before the fix, three sign-in route regressions fail with 502 instead
of 422; all four upstream-error controls pass.
- After the fix, all 339 tool-access and error-handler tests pass,
including authorization and redirect protections.
- These suites ran against disposable Homebrew PostgreSQL 16.14 through
the existing test-constructor seam. The temporary setup and config
remain outside the repository. CI uses the ordinary embedded PostgreSQL
setup.
- Direct server `tsc --noEmit` passes.
- Full build and recursive typecheck were attempted; the Runner Rust
step cannot run because `cargo` is absent on this machine.
- Full `pnpm test:run`: 8,211 passed, 14 failed, 4,759 skipped. The 36
failed files match the existing embedded PostgreSQL startup/cleanup and
macOS runtime-cache `EACCES` limitations. The changed database-backed
service suite passed separately with local PostgreSQL.
- Greptile: 5/5 with no unresolved comments. Seven CI workers received a
simultaneous shutdown signal; the failed jobs are being retried through
the normal workflow. Other completed checks passed.

## Risks

Low risk. Clients now receive 422 instead of 502 for the explicit
`oauth_challenge` condition. The code, message, and remediation links
remain available. No permissions, credential handling, OAuth discovery
rules, retry policy, or schema change. Other upstream failures retain
their existing behavior.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (339 service and
error-handler tests; full workspace limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 00:34:45 +00:00
Devin FoleyandPaperclip 6de50ba594 fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can enable Sentry for the server and the signed-in
browser.
> - The server SDK reads `SENTRY_ENVIRONMENT` from the process
environment.
> - The browser receives its DSN through the session response, but
receives no environment.
> - A browser in staging therefore reports errors under the SDK's
production default.
> - This pull request passes the configured environment through the
existing session and monitoring gate.
> - Browser errors then identify the deployment environment while
preserving the existing privacy settings.

## Linked Issues or Issue Description

**What happened?**

With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged
`production`. This can send errors to the wrong environment's alerts and
makes deployment follow-up unreliable.

**Expected behavior**

The browser uses the server's configured Sentry environment. A reused
image works in either staging or production. A signed-out browser still
sends no events.

**Steps to reproduce**

1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`.
2. Sign in and capture a browser exception.
3. Inspect the event environment. Before this change, it is
`production`.

**Paperclip version or commit**

Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real
browser SDK and a local test transport.

**Deployment mode**

Authenticated server and browser with optional Sentry monitoring
enabled.

No duplicate environment-attribution issue or pull request was found in
the targeted GitHub search.

## What Changed

- Add `sentryEnvironment` to the authenticated session response and
shared schema. The optional field supports a newer browser reading an
older server response.
- Pass the environment to the browser SDK. An environment change
restarts the client through its existing serialized lifecycle.
- Cover environment attribution with a real SDK event, session
authorization, unchanged-session refetches, environment changes, and
legacy responses.
- Document configuration and compatibility. Keep the loaded bundle's
release identity and existing privacy filters.

## Verification

- The regression test emits `production` for a requested staging
environment before the fix.
- Focused route, schema, browser lifecycle and real-SDK tests: 69 pass.
- UI and shared-package typechecks, direct server `tsc --noEmit`, and
token gates pass.
- Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at
the Runner Rust step because `cargo` is absent on this machine.
- Complete UI suite: 6,540 tests pass in 626 files.
- Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files
fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache
`EACCES`. These match the existing local baseline; none touch the
changed behavior.
- Greptile: 5/5, no unresolved review threads. Linux CI has passed
Build, Typecheck + Release Registry, and the completed test jobs so far.
Remaining jobs are running or queued: the AWS runner provisioner is
retrying EC2 CreateFleet `InternalError` responses. Full results will be
recorded before merge.

## Risks

Low risk. This adds one optional session field and changes Sentry
attribution only. No migration or new monitoring opt-in is introduced.
Missing settings keep the browser SDK default. Agent and unauthenticated
requests still receive 401 without monitoring settings. Existing loaded
browser bundles keep their old behavior until refreshed.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused and full UI suites
pass; full-root environment failures documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 23:40:46 +00:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 57fd8b70d2 feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack connections let a team talk to those agents in Slack.
> - Agents now have a saved avatar, but Slack setup did not offer that
image.
> - A matching avatar helps a team recognize its agent.
> - This pull request adds an optional avatar step and a download in
connector Settings.
> - Users download a PNG and upload it directly in Slack with clear
instructions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Slack connector onboarding and its Settings page.

**Current behavior**
Setup does not offer the assigned agent's avatar or explain how to
upload it in Slack.

**Proposed behavior**
After Slack connection verification, users can download a 512 × 512 PNG
of their agent's saved avatar. They can upload it in Slack, confirm, or
skip. Settings keeps the download and upload instructions available
after onboarding.

**Reason and benefit**
The same avatar helps people recognize the agent across Paperclip and
Slack. Users who skip the optional step can return to it in Settings.

**Breaking changes**
None. No schema, authentication, Slack scope, or provider API change.
Completed connections keep their existing completion state. Searched
existing Slack avatar and Cliptoon PRs; no matching implementation was
found.

## What Changed

- Add an optional avatar step before personal Slack account linking.
Keep the numbered sidebar and shared footer.
- Resolve the selected agent's saved appearance for the preview and PNG
download.
- Add the same download and expandable upload instructions to connector
Settings.
- Remember uploaded or skipped per company and endpoint in browser
storage. Treat uploaded as user confirmation, not provider verification.
- Reject failed or non-PNG download responses and allow retry.
- Reuse the production avatar components in onboarding and Settings
stories.
- Test wizard progression, resume, Settings, download recovery, storage
isolation, and terminated assigned agents.
- Exercise real PNG downloads in the Slack browser flow and keep default
app names consistent with app creation.
- Fetch the assigned agent directly so its saved avatar remains
available after termination.

## Verification

- Focused chat suites: 48 passed; the two affected suites passed again
after the final naming fix (32 tests).
- Slack browser E2E passed through setup, avatar download, account
linking, and Settings download. PNG signature and 512 × 512 dimensions
verified.
- UI token gates passed.
- Browser: downloaded the real 512 × 512 PNG; checked confirmation,
return, mobile layout, and Settings instructions.
- Full workspace typecheck, application build, and production Storybook
build passed.
- All latest-head CI checks passed (54 passed, 2 skipped), including all
browser, chat, general, and serialized test groups. The unrelated Sentry
test failed once and passed on the single CI rerun; its suite also
passed locally.
- Local full-suite attempt encountered a rapid Slack callback ordering
failure under concurrent build load; that test passed in isolation, and
all three chat shards passed in CI. The remaining local run was not used
as the merge gate.
- Review the Connections / Slack / Add avatar and Avatar in Settings
stories.

## Risks

- Slack upload is manual. Confirmation does not claim to verify the
Slack icon.
- Optional step progress is browser-local. Clearing storage or changing
browsers can show it again. Setup still works when storage is
unavailable.
- The existing avatar API remains the image source. Download failures
show a retry message.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 09:38:48 -05:00
Devin FoleyandPaperclip c65fc9e3c8 fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures

Include connect timeouts in the bounded retry policy for idempotent actor
synchronization. Handle WebSocket constructor failures through existing
reconnect paths and preserve HTTP polling while realtime is unavailable.
Refresh visible company queries until the socket recovers and clear all
fallback timers on hiding or unmount.

Verify 172 focused tests, server/UI typechecks, UI build, and design token
gates. Full workspace build/typecheck require the unavailable Rust toolchain;
the full test run is tracked separately.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 12:43:25 -07:00
DottaandPaperclip 9f30eb10dd fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path

> - Paperclip manages AI agents and keeps their work attached to tasks.
> - Chat connectors carry user messages and agent replies between a
provider and those tasks.
> - Each extra startup and context reset delays a reply.
> - Managed account metadata was lost during adapter decoding, so
compatible follow-ups started fresh.
> - This branch fixes the reset and measures the remaining preparation,
execution, and delivery costs.
> - The changes must preserve account isolation, authorization, durable
output, and recovery ownership.

## Linked Issues or Issue Description

Refs #13699. The related service lifecycle work in #13410 and #13408 is
separate; this branch focuses on task-bound chat response latency.

**What happened?**

Managed AI follow-ups started new provider sessions even after their
configuration fingerprint stayed stable. The Codex codec removes unknown
fields. The resume check then read the removed credential identity and
treated it as a credential change.

**Expected behavior**

Compatible follow-ups resume the correct provider session. Changes to
credentials, responsible users, permissions, or task configuration
retain their reset behavior.

**Steps to reproduce**

1. Use a Slack connector with a managed AI connection.
2. Send a message, then send a same-thread follow-up.
3. Inspect the configuration reset reason and the provider session
identity.

**Paperclip version or commit**

Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`.

**Deployment mode**

Cloud staging with a native Codex runner.

## What Changed

- Read saved credential identity before adapter decoding discards it.
- Remove the internal credential identity from adapter-facing session
params.
- Test the real Codex codec and missing, changed, or unmanaged identity
cases.
- Preserve configured warm Codex runners and flush refreshed credentials
after every turn.
- Fence detached or closing session handles from successor credential
ownership.
- Stage current Codex launch credentials after restoring durable session
history, uploading launch assets only once.
- Reuse a runner binary already in the retained sandbox only when its
SHA-256 matches the controller-owned artifact; still verify required
capabilities before launch.
- Lock the task before the run when saving results, preventing deadlocks
with task updates.
- Scope reusable projectless sandboxes to the company, environment,
task, agent, and runtime configuration; verify Daytona sentinels for
that scope.
- Admit a new authorized chat message after a fully committed failed run
and verified process cleanup.
- Send compact deltas for verified plain-text Slack continuations. Match
the actual prior run and current comment identity/body; exclude edited
historical comments and prior agent output, preserve genuine brief edits
and the full bootstrap fallback.
- Keep attachments, omitted input, questions, approvals, recovery, and
other providers on their existing framing.
- Document managed session compatibility, credential lifecycle, and
compact continuation boundaries.

## Verification

- Workspace/session coverage: 156 tests passed.
- Native session and credential ownership coverage: 390 tests passed,
including exact artifact reuse, mismatches, failed probes, timeouts, and
explicit artifact overrides.
- Explicit continuation and durable chat authorization coverage: 172
tests passed.
- Session resume and launch preparation coverage: 416 tests passed.
- Result persistence coverage: 15 tests passed. The new concurrency test
reproduced a PostgreSQL deadlock before the lock-order fix.
- Environment lifecycle coverage: 92 tests passed, including projectless
reuse and task/agent isolation at both selection and atomic handoff.
- Daytona plugin coverage: 237 tests passed; 6 gated tests skipped.
Standalone plugin build passed.
- Compact Slack continuation and native resume coverage: 69 tests
passed, including full-bootstrap retention, matching message
authors/bodies, current-delivery selection, rejection of duplicate
identities and historical comments, brief edits, and attachment/recovery
fallbacks.
- Final frozen-head `pnpm test:run` on repository-supported Node 26: 668
suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped.
The sole failure was a local `socket hang up` in
`issue-recovery-actions.test.ts`, not an authorization assertion
mismatch. All 57 tests in that suite passed three fresh reruns, and the
suite passed latest-head CI. The full local invocation is therefore not
claimed green.
- An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup
failure; that 11-test catalog suite passes on Node 26 and in CI. No test
behavior or timeout was relaxed.
- Full local typecheck and build passed. Latest-head CI is green;
Greptile is 5/5 with no unresolved review threads.
- Two real Slack baseline replies took 25.1 and 24.6 seconds
(24.9-second mean). Three same-thread signed probes on this head took
23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample
and a modest wall-clock improvement, not a large or statistically
established speedup.
- In that same thread, uncached provider input fell from 8,514 tokens
before compact input to 694–765 tokens afterward. The current delivery
uses a 362-character delta; the full 19–21k-character bootstrap remains
available for failed resume. Verified runner artifact preparation fell
from about 1.2 seconds to 0.6 seconds.
- A fresh thread created a separate task, sandbox, and provider session
with full bootstrap (24.4 seconds). Its follow-up reused its own
sandbox/session and compact input (28.3 seconds, including 16 seconds of
model execution). Model variability and process startup remain
substantial.
- A signed duplicate webhook produced exactly one user comment, one
successful run, and one final Slack reply. Slack's API independently
confirmed the actual replies and a public task URL without an internal
or pool hostname.
- Earlier signed probes verified recovery after a failed run and reuse
across a server deployment. The final idle test observed Daytona report
the sandbox as stopped, then delivered a new reply in 19.9 seconds using
the same sandbox/provider-session identity and compact input. Slack’s
API confirmed that reply.
- Live probes use signed synthetic inbound webhooks and real outbound
Slack delivery, read back through Slack’s API. The final browser recheck
found the Mac locked and the Slack tab blocked by another extension, so
this is not claimed as full UI E2E proof.
- This is a review branch. Do not merge until the maintainer reviews it.

## Risks

- Incorrect session reuse could mix account or task context. Missing or
changed identities continue to reset, and existing authorization checks
remain in place.
- Warm mode remains opt-in. Remote warm mode requires a reusable sandbox
lease. Retained processes keep credentials until they close, so idle
expiry and ownership fences are required.
- A fresh user message may continue after a committed provider failure.
Approval, current authorization, process termination, and prior-result
checks remain required.
- Projectless sandbox reuse is task- and agent-scoped. Missing or
mismatched ownership cannot replace an existing lease; existing
workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox
per task/agent, so provider auto-stop and deletion policies still
determine idle compute and storage costs. Fleet defaults are unchanged.
- Compact prompts apply only after proven resume and a matching
prior-run delta. Missing or specialized context falls back to full
input; fresh sessions always receive the full bootstrap.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-20 13:12:44 -05:00
Devin FoleyandPaperclip 600e552d7b fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build.
Use validated build commits for Docker and source/npm artifacts, preserve
explicit server release overrides, and keep cached browser bundles tied
to the commit they loaded.

Verify 127 focused tests, server/UI typechecks, Docker and source build
stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 08:08:33 -07:00
DottaandPaperclip 2a99de80ec fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.

## Linked Issues or Issue Description

Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.

**What happened?**

An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.

**Expected behavior**

Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.

**Steps to reproduce**

1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.

**Paperclip version or commit**

Reproduced on the source at b193077582.

**Deployment mode**

Cloud staging with a native runner.

## What Changed

- Supply an exact public task URL in fresh and resumed external-chat
context.
- Return a nullable `url` on task records in `get_task_context` and
`search_tasks`, including child tasks.
- Resolve the current claimed Cloud origin at use time. Preserve the
existing safe-URL filter and omit unsafe links.
- Normalize only the exact credential-home paths created by the managed
AI runtime when comparing session configuration. Do not change the
actual execution environment.
- Add regression tests and update the connector playbook.

## Verification

- 255 focused tests passed for public URLs, prompt context, AI
connections, and session configuration.
- 19 native task-tool tests passed, including public URL lookup, origin
changes, and unsafe URL rejection.
- Full workspace typecheck and build passed locally. The full local test
run is still in progress; all CI test lanes passed at the final head.
- All 54 CI checks passed, with two optional jobs skipped. Greptile
scored 5/5 with no review comments.
- After merge, deploy the exact merged commit to the authorized staging
stack. Check a fresh Slack mention, a same-thread reply, the returned
task link, and the measured session reuse and response time.

## Risks

- Existing saved fingerprints can reset once after this change. Future
runs should reuse a compatible session.
- The fingerprint change applies to managed AI connections. Tests
preserve resets for real account, credential, model, permission, and
environment changes.
- A public task URL still requires normal Paperclip access. This change
grants no access and keeps internal URL filtering.
- This change does not remove all runner startup or checkpoint latency.
Staging measurements are still required.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size were not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 20:40:05 -05:00
DottaandPaperclip b193077582 fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack connections turn messages into governed agent runs and return
replies to Slack.
> - Remote runs must replace an incompatible sandbox runner with the
controller's packaged binary.
> - The fallback used a package-relative path that does not match the
vendored server layout.
> - Slack callback health also compared the internal proxy address with
the public callback address.
> - Master now contains the runner fallback fix; this pull request fixes
callback health and extends missing-runner regression coverage.

## Linked Issues or Issue Description

Refs #13677, #13680, #13691, and #13686.

**What happened?**

A packaged controller stopped a Slack-triggered remote run with
`runner_remote_artifact_unavailable` when the sandbox runner needed
replacement. Working Slack callbacks also showed a stale URL warning
behind the Cloud gateway.

**Expected behavior**

The controller stages its packaged runner when needed. Callback health
uses the observed public address and still detects real address changes.

**Steps to reproduce**

1. Run a packaged server with a sandbox that has an older runner or no
runner.
2. Send a Slack mention to an agent that uses that sandbox.
3. Route signed Slack callbacks through a claimed Cloud gateway that
rewrites the upstream host.
4. Check the run and the Slack callback health panel.

**Paperclip version or commit**

Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The
fix branch also includes #13691.

**Deployment mode**

Packaged server with a Cloud gateway and a Daytona sandbox.

## What Changed

- Extend the controller-owned runner fallback tests with a missing
sandbox binary case; retain the fix now merged in #13686.
- Prefer dedicated gateway diagnostic headers that survive provider
rewrites of standard forwarded headers. Use validated host hints only
for callback-health evidence on claimed Cloud instances after provider
acceptance. Preserve request bodies, routing, authentication, and
configured callback URLs.
- Cover current, stale, and missing sandbox runners, all Slack callback
surfaces, rejected callbacks, malformed proxy hints, real host and port
changes, and the existing self-hosted behavior.
- Document the artifact lookup and callback-health boundaries.

## Verification

- Native session executor and binary resolver suites: 379 tests passed
after merging current master (`45c99a0d0`).
- Targeted callback integration suite with disposable PostgreSQL: 4
tests passed before rebase.
- Full `pnpm -r typecheck` and `pnpm build` passed after merging current
master. The full local suite passed 12,668 tests; one suite failed to
start its disposable PostgreSQL. Rerunning that suite alone passed all
31 tests.
- The built server resolver selected the executable under
`server/dist/vendor/paperclip-runner/bin/`.
- Final callback regression: all 4 targeted integration tests pass,
covering provider header rewrites, default ports, uppercase/trailing-dot
hosts, and ignored self-hosted hints.
- Live staging proof of the runner fix: a previously failed Slack thread
recovered, a new mention received its requested response, and the
account-connect command succeeded. Unsigned callbacks returned 401.
- Live browser and Slack acceptance passed: generated callback URLs,
account linking, a new mention, an interactive question and answer, and
all three callback-health indicators. The connector was activated
through the onboarding UI.
- All 54 latest-head checks pass, two optional checks are skipped, and
Greptile is 5/5 with no unresolved threads.

## Risks

- Runner fallback must select a binary for the remote platform. This
preserves explicit remote artifact overrides and the existing capability
checks.
- Proxy headers are not identity proof. They are used only for
diagnostics on claimed Cloud instances after the provider accepts the
request. Self-hosted instances ignore them. Wrong public hosts and ports
still warn.
- No database migration or new public API contract.

## Model Used

OpenAI GPT-6 via Codex. Used reasoning, repository tools, code
execution, and browser/native-app testing. Exact model variant and
context window were not exposed by the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 17:32:51 -05:00
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00
DottaandPaperclip 04546c82d5 fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner can execute a task inside a Daytona sandbox.
> - The sandbox can keep running when the Paperclip controller restarts.
> - Recovery treated sandbox process IDs as local process IDs and
selected the wrong recovery path.
> - Live verification also found races between startup, shutdown, and
queued task cleanup.
> - This pull request verifies the existing remote owner and orders
those transitions.
> - Users can continue the same task and provider session after a
controller restart.

## Linked Issues or Issue Description

**What happened?**

The Daytona `recover-controller` cases failed with
`runner_state_identity_mismatch`. Remote process IDs can be absent on
the controller or collide with unrelated local processes. Recovery then
looked for remote state in the local runner directory. Later turns could
also start before the previous executor released its sandbox resources.

**Expected behavior**

Reconnect to the original sandbox and authenticated runner. Preserve the
task, provider session, and queued comments. Reject a replacement
sandbox or mismatched identity. Do not start another provider during
reattachment.

**Steps to reproduce**

Run the `everyday-workflows` `recover-controller` case for
`runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a
Python tool, requests a revision, restarts the controller during
execution, and queues another revision. It then downloads and tests the
final ZIP.

Related: #13682 is the preceding operational fix. #13291 addresses
legacy sandbox conversation recovery, a different execution path. #13666
includes broader run-capacity work; this change guards cleanup of an
existing native task executor.

## What Changed

- Add remote runner recovery without interpreting sandbox PIDs on the
controller.
- Verify the original provider lease, remote workspace, durable state,
process marker, and authenticated PRP authority before adoption.
- Compare the process marker with live Linux boot identity and start
ticks to reject PID reuse. Read virtual proc files through the
guaranteed Node runtime; unavailable proof blocks adoption without
blocking a fresh launch.
- Make the E2E supervisor own the actual server process so forced
restart cannot leave a late database closer behind.
- Scope the chat delivery lease test to its own fixture instead of
draining other tests’ pending deliveries.
- Preserve provider-attempt counts and recorded evidence during
reattachment.
- Serialize an idle-session checkpoint with admission of the next native
turn.
- Wait for an in-progress startup to acknowledge restart detachment.
Fail after a bounded deadline if it cannot.
- Keep a queued comment waiting until the previous native task executor
releases its resources. Allow unrelated tasks to continue.
- Update the Daytona image's resolved lock digest to match current
dependency manifests.
- Add classifier, ownership, process, startup, checkpoint, and
queued-admission regression tests. Document recovery behavior.

## Verification

- 415 focused tests passed across native execution, restart recovery,
workspace synchronization, queued admission, and real-process restart
tests. The final Node-based fingerprint change passed all 375
native-session tests.
- Runner harness unit tests: 394 passed. Chat integration shard 2: 335
passed after fixture isolation.
- The exact fingerprint command succeeded twice in a disposable Daytona
sandbox and returned the same identity; the sandbox was deleted.
- 11 real-process restart integration tests passed, including absent and
colliding remote PIDs.
- Repository typecheck and final build passed. Broad local checks found
machine-dependent database startup and timing failures; focused retries
passed. The final-revision PR pipeline is green. One unrelated browser
shard hit a five-second blank-page timeout on the first run and passed
its targeted retry.
- Final-revision local headed browser E2E:
`everyday-workflows.runner-acpx-claude.daytona.recover-controller`
passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser
inspection confirmed Done, all three ZIPs, and delivery of the queued
follow-up. All three runs succeeded using the same provider session. The
harness downloaded and independently tested the final artifact.
- Final-revision Daytona campaign:
https://github.com/paperclipai/paperclip/actions/runs/35463999611 —
**Codex passed first attempt (4.8 minutes); ACPX Claude passed first
attempt (6.1 minutes)**. Campaign aggregation/publication is finishing;
both test jobs succeeded.
- Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**,
no open findings.
- Staging browser verification is pending selection of a disposable
staging instance and removal of a Chrome extension UI block.

## Risks

- Recovery now depends on the original sandbox remaining available. A
replacement or mismatched identity still blocks adoption.
- Shutdown waits up to 30 seconds for a native startup to reach a safe
detach point. An unfinished startup returns a clear failure instead of a
false detach receipt.
- Queued native work on the same task waits for cleanup. Unrelated tasks
remain eligible.
- The image digest update rebuilds the Daytona runtime image. No
database migration or public API change is included.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code execution, and
browser tools. The runtime does not expose the exact deployed model ID
or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:13:34 -05:00
Nicky LeachandClaude Opus 5 aeef493f4a chore(db): keep only the newest 5 drizzle snapshots and stop shipping them (#13687)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - `@paperclipai/db` owns the Drizzle schema and the migration history
> - Drizzle writes a full copy of the schema as a snapshot for each
generated migration. Each snapshot is now about 1.3 MB.
> - The `meta/` folder is about 103 MB. That is about half of each
checkout and each worktree. The build also copies it into `dist`, so the
published `@paperclipai/db` package is 112.8 MB unpacked.
> - `drizzle-kit generate` reads only the newest snapshot. The runtime
migrator reads only the `.sql` files and `_journal.json`.
> - This pull request keeps the newest 5 snapshots and removes snapshots
from `dist`.
> - The benefit is a checkout that is about 100 MB smaller, and a
published package that is about 1.5 MB instead of 113 MB.

## Linked Issues or Issue Description

Refs #11240, #11254, #12333 (earlier snapshot work: diff collapse,
binary diffs, drift repair)

**What existing behavior does this improve?**
The size of the Drizzle migration snapshots in the repository and in the
published `@paperclipai/db` package.

**Subsystem affected**
`packages/db`: migrations and the build.

**Current behavior**
`packages/db/src/migrations/meta/` holds 141 snapshots (102.8 MB). The
size grows faster than the number of migrations, because each snapshot
is a full copy of the schema. `build` runs `cp -r src/migrations
dist/migrations`. `@paperclipai/db@2026.916.0` contains 147 snapshot
files. It is 112.8 MB unpacked and 5.3 MB as a tarball.

**Proposed behavior**
Keep the newest 5 snapshots. `generate` deletes older snapshots after it
runs. `dist` gets only `*.sql` and `meta/_journal.json`.

**Reason and benefit**
- In drizzle-kit 0.31.10, `generate` sorts `meta/*` and diffs against
the last snapshot only (`bin.cjs`, `preparePrevSnapshot`). The [generate
docs](https://orm.drizzle.team/docs/drizzle-kit-generate) also say it
compares against "the most recent" snapshot.
- Gaps in the snapshot history already work. 138 of the 280 migrations
never had a snapshot, because they were written by hand before
`doc/DATABASE.md` required `generate`.
- We keep 5 snapshots instead of 1. This lets a developer undo the
latest generated migration, and it keeps the `prevId` chain for recent
branches.
- Git keeps the history cheaply. The 631 snapshot versions use only 2.2
MB of the pack, because git stores each version as a delta of the
previous one. The cost is in the checked-out files, not the clone
download. Therefore this change does not use git-lfs and does not
rewrite history. Old snapshots stay available with `git show
<rev>:<path>`.

## What Changed

- `packages/db/package.json`: a new `prune:snapshots` script keeps the
newest 5 `*_snapshot.json` files. It is `ls | sort -r | tail -n +6 |
xargs rm -f`, which works with the BSD tools on macOS and the GNU tools
on Linux. `generate` runs this script after `drizzle-kit generate`.
- `packages/db/package.json`: `build` copies only `src/migrations/*.sql`
and `meta/_journal.json` into `dist/migrations`.
- Deleted 137 older snapshots. `0277`–`0281` remain.
- `chat-identity-migration-reconciliation.test.ts`: removed the walk
over snapshots `0254`–`0268`. Those files do not change after merge, and
the walk would fail after pruning. The journal-order assertions in the
same test remain. `migration-snapshot-drift.test.ts` still makes sure
that the newest snapshot matches the schema.
- `doc/DATABASE.md`: documented the retention rule.
- The snapshots were already marked `linguist-generated=true -diff
-merge` by `packages/db/.gitattributes` (#11240). No change there. `git
check-attr` confirms it.

## Verification

- `pnpm --filter @paperclipai/db exec vitest run`: 43 files, 160 tests
pass.
- `pnpm --filter @paperclipai/db build`: `dist/migrations` contains 280
`.sql` files and `meta/_journal.json`. It is 1.5 MB, compared with about
105 MB before.
- `prune:snapshots` was run on macOS (BSD) and in `debian:stable-slim`
(GNU findutils 4.10). With 7 fixture snapshots, it keeps the newest 5.
When 5 or fewer are present, it deletes nothing and exits 0 on both.

## Risks

- Low risk. Runtime migration does not read snapshots. The only commands
that read older snapshots are `drizzle-kit check` and `drizzle-kit
drop`. No script or CI job calls them, and they still have the newest 5
snapshots.
- A branch that is open now can still add its own snapshot. If the
branch conflicts, the rule is the same as today: renumber the migration
and run `generate` again.
- When we upgrade to drizzle-kit v1 (folder per migration), check
whether its new cross-branch "commutativity" checks need a longer
snapshot history.

## Model Used

- Claude Opus 5 (`claude-opus-5`) in Claude Code, with tool use (shell,
file edits, web fetch).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 10:40:38 -07:00
DottaandPaperclip da257c3069 Warn when routine webhook URLs may not be publicly reachable (#13684)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Routines can start that work when another app sends a webhook.
> - Local and private URLs often cannot receive events from public
services.
> - HTTPS alone does not make a Tailscale address public.
> - This pull request explains these limits during setup and editing.
> - Users can still finish setup for senders on their own network.

## Linked Issues or Issue Description

Refs #13637. Webhook setup needs a clear warning when the generated URL
appears local, private, or unencrypted. The warning must explain how to
make the endpoint reachable without blocking private-network use.

## What Changed

- Add a shared warning banner to the Connect, Check connection, and Edit
webhook views.
- Distinguish localhost, private network addresses and domains, HTTP,
and Tailscale hostnames.
- Explain the difference between Tailscale Serve and Funnel. Link to the
Paperclip HTTPS guide.
- Add five full-page Storybook examples, design guide examples, and
documentation.
- Add URL classification tests and a regression test that finishes setup
despite the warning.

## Verification

- Passed 41 focused URL and trigger-flow tests.
- Passed workspace typecheck, workspace build, token gates, and
Storybook build.
- Browser-tested the Tailscale story through Check connection, Finish
setup, and Edit webhook. The warning stays visible and does not block
setup.
- Open Product / Routines / Webhooks stories 11–15 to review the warning
states.
- All 54 PR checks passed, including the full test matrix and eight
browser shards; two optional Storybook jobs were skipped by workflow
policy.
- The server supervisor readiness test timed out once in CI, then passed
on rerun and locally (6 tests).
- The duplicate local full-suite run was stopped after the complete CI
matrix passed.
- Greptile: 5/5 on the current commit, with no unresolved review
comments.

## Risks

- URL checks are hints. They do not test DNS, firewall rules, or actual
reachability.
- A Tailscale hostname can serve either private Serve traffic or public
Funnel traffic. The warning explains this uncertainty and permits both.
- No API, schema, authentication, or webhook delivery behavior changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 12:24:06 -05:00
Devin FoleyandPaperclip b70641f23f feat(plugins): support image catalogs and persistent application overlays (#13646)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins extend the application without adding each integration to
Core.
> - A downstream image needs a way to supply prebuilt plugins.
> - Some plugin UI must stay mounted as users move between pages.
> - This change adds an image catalog and a persistent application slot.
> - Operators can upgrade or remove these plugins through their image
and configuration.

## Linked Issues or Issue Description

**Subsystem affected**

Plugin packaging, activation and application UI.

**Problem or motivation**

The built-in plugin catalog is fixed in Core source. Downstream images
cannot add entries through an explicit catalog. Existing page slots also
cannot preserve a small application overlay across route changes.

**Proposed solution**

Read a bounded catalog of prebuilt plugins from the image. Verify its
files before importing manifests. Use the existing managed selection and
plugin lifecycle. Add an `appShellOverlay` slot with account and company
cleanup.

**Alternatives considered**

A downstream fork adds merge work. Script injection provides no plugin
lifecycle. A separate runtime download system adds a second distribution
channel.

**Roadmap alignment**

This extends the existing plugin system. Related PR #9006 covers runtime
install replication; this change covers immutable image contents. PR
#12555 covers CLI scaffolding. Neither provides this catalog or
application slot. The maintainer requested this work directly.

## What Changed

- Validate catalog identities, confined paths, package versions and
bundle hashes before importing code.
- Apply image selection to persisted plugin installs, including removal
and rollback. Adopt the verified image path from existing npm/local
installs and bind runtime worker/UI entrypoints to verified package
declarations.
- Mount application overlays in both UI shells. Preserve route state and
clear it on account, company and onboarding changes.
- Restrict service-worker offline storage/fallback to hashed public
assets in a separate cache namespace; exclude application HTML and
extension/API data, including after worker restart.
- Document the packaging contract, trust model and rollback
requirements.

## Verification

- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Affected server/UI typechecks and builds, plus token
gates, passed again after rebasing onto current master; the 124 focused
tests also passed after rebase.
- Latest focused verification: 124 tests in nine files passed for
catalog/reconciliation/loader, overlay lifecycle, Layout and
service-worker policy. The broader UI/shared/SDK run passed 7,204 tests
in 690 files with canonical `TMPDIR`.
- Real disposable Core/PostgreSQL: catalog install, selection removal,
0.1.0→0.1.1→0.1.0, same-version npm/legacy-path adoption, and
preservation of disabled status passed. Added permissions entered
`upgrade_pending`, withheld UI across restart, and activated only after
explicit operator enable.
- Real Chromium: desktop/mobile layout, route draft retention and
Escape/focus passed with mocked extension responses. A persistent
browser restart retained public hashed-asset offline fallback while
refusing seeded legacy/current private entries and legacy HTML.
- Full `pnpm test:run`: 12,539 passed; 17 failed across six existing
files, stopping later phases. macOS read-only directory renames fail in
runtime-skill-cache and company-skills-service; email tests require an
absent local AgentMail fixture. Native runner/comment-redaction passed
in isolation after temporary Rust setup; agent-conversations also passed
in isolation. No unrelated source was changed to hide failures.
- After rebase, two unchanged chat timing tests failed in CI and passed
locally in isolation. Their CI shard passed on its single retry. All
other current-head CI jobs passed on the initial run; review is 5/5 with
no unresolved threads.
- No live deployment or external plugin service was used.

## Risks

- Plugins are trusted code. The catalog detects packaging errors; it
does not authenticate an untrusted image builder.
- Invalid catalogs fail startup. Images must contain the catalog and
bundles together, with stable directories.
- A host older than this contract lacks the activation guard. Disable
added plugins and remove their configuration keys before reverting to
it.
- Offline navigation now returns 503 instead of replaying cached
application HTML. Only public build assets have offline fallback.
- Rolling back an unapproved permission change retains the approval
gate; review the current manifest and explicitly enable it. A reduced
permission set cannot establish prior approval or prior enabled status.
- Plugin data migrations need their own rollback policy. This change
retains installed records and does not reverse migrations.

## Model Used

- OpenAI GPT-6 (Codex), model ID `gpt-6`, with repository inspection,
code execution and browser verification. The runtime does not expose an
exact context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites; broad
macOS server-run exceptions are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (fresh run on 488b3754ae; chat
shard passed its single retry)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(fresh review on 488b3754ae)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:18:10 -07:00
DottaandPaperclip f589660ec0 feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Routines turn scheduled work and external events into tasks for an
assigned agent.
> - Webhook setup was disabled, and actor authentication rejected valid
webhook bearer keys.
> - Operators need to connect and test a sending app before events can
start work.
> - This pull request adds a guided setup with durable connection tests
that cannot dispatch a task.
> - It keeps trigger management, execution tasks, and activity within
the routine.
> - The benefit is a webhook that can be configured, verified, and
operated from one place.

## Linked Issues or Issue Description

Fixes #11937.

Related: #13216 adds provider-specific Sentry support. This PR addresses
general routine setup and ingress. #6841 addresses legacy secret
bindings; this PR retains the existing secret service.

**Current behavior**

Webhook creation is disabled. Bearer deliveries can fail in agent
authentication before the routine checks its key. Setup has no safe
connection test. Runs and Activity send the operator away from the
routine.

**Proposed behavior**

Choose a schedule or a webhook. Follow the setup steps, copy credentials
or complete agent instructions, and test delivery without creating work.
Finish setup to allow future events to start tasks. Edit or remove
compact trigger cards, undo removal, and inspect tasks and activity
inside the routine.

**Reason and benefit**

An operator can verify credentials and delivery before enabling
automatic work. Durable setup state survives refreshes and restarts.
Retry receipts prevent an old test event from starting work after
activation.

## What Changed

- Add a production trigger wizard using reusable Slack setup navigation
and footer components.
- Add schedule and webhook choices, one-time credentials, agent
instructions, and live connection feedback.
- Persist pending setup, test delivery receipts, connection status, and
reversible trigger removal.
- Keep setup checks free of routine runs, tasks, and agent wakeups.
Preserve delivery idempotency after activation.
- Add compact trigger cards, inline editing, key rotation, pause
controls, removal, and Undo.
- Keep Runs and Activity in the routine. Use the shared task list and
compact activity rows.
- Permit only exact public delivery POSTs through actor authentication.
Retain webhook authentication, JSON-object validation, and log
redaction.
- Add production-backed Storybook states and focused server, database,
and UI coverage.
- Document signing modes, setup checks, retries, rotation, HTTPS
ingress, and navigation.

## Verification

- Full workspace typecheck, build, and token gates passed on the rebased
branch. Storybook also builds.
- Focused routine, middleware, logging, shared wizard, and UI coverage
passes on the rebased branch: 195 tests across 14 files. The migration
passed on a fresh PostgreSQL database and on two repeated applications.
- Browser testing used the real app, database, and a deterministic
process worker through Tailscale HTTPS and the current Cloud proxy code.
- Verified rejected keys, safe setup deliveries, persisted state after
restart, activation, retry deduplication, key rotation, schedule
editing, removal, and Undo.
- Fresh bearer and GitHub-signed deliveries created tasks that the
worker checked out and completed. Runs and Activity stayed within the
routine.
- Current Cloud ingress tests passed. Public delivery POSTs passed
through without a browser session; management routes remained gated.
- All 54 current-head PR checks pass, including general and serialized
tests, all eight browser E2E shards, typecheck, build, runner checks,
security checks, and the canary dry run. Two optional Storybook jobs are
skipped by workflow conditions.
- Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review
threads. The stale connection-status finding is fixed and covered by a
regression test.
- No production deployment was performed.

## Risks

- Migration 0281 adds three trigger columns and a test-receipt table. It
is additive and safe to reapply. Apply it before running the new server.
Existing triggers remain live by default.
- Requests without delivery IDs are new events after activation. Senders
must reuse an event's delivery ID for retries.
- Completed webhooks keep normal dispatch behavior. Their management
connection check can start work; the UI states this.
- Removing a trigger archives it. Undo restores the URL and credentials.
Permanent deletion remains available through the existing API.
- Public ingress must remain restricted to the delivery POST route. The
tenant verifies credentials. Cloud sleeping-stack behavior is unchanged.
- Shared setup components also serve Slack. Existing setup contracts and
navigation tests cover that integration.
- Senders must use application/json with an object. Other media types
receive 415.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 10:49:05 -05:00
DottaandPaperclip 36dbb7ed1c fix: harden agent chat runner tools and recovery (#13678)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat turns discussion into plans, tasks, reviews, and hires.
> - These workflows need reliable tool results and task context on the
native runner.
> - Live Claude and Codex tests exposed lost retry requests, invalid
project inputs, and a child startup crash.
> - Recovery also exposed a misleading retry action and missing child
task context.
> - This pull request fixes those paths and adds regression coverage.
> - Agents can continue the original request and operators can inspect a
stopped run.

## Linked Issues or Issue Description

**What happened?**

A failed Agent Chat retry could lose the user's question. Project
creation accepted unsupported icons in its tool schema. Codex could stop
when a helper's MCP startup event arrived before its thread lineage. A
stopped task offered Retry even when the server required execution
reconciliation. Resumed agents could miss existing delegated tasks.
Hiring and review instructions did not describe the native runner's
available tools and source requirements.

**Expected behavior**

Retries retain the selected request. Tool schemas match the API. Child
startup information does not gain authority over the parent or stop it.
Recovery actions match the server's requirements. Task context exposes
existing child work. Handoffs contain the material the assignee needs.

**Steps to reproduce**

1. Enable experimental Agent Chat in an isolated development instance.
2. Configure native Codex and ACPX Claude agents on Paperclip Runner.
3. Ask for a plan, revise it, approve task creation, and request a hire
and status report.
4. Retry a failed chat turn and check that it answers the original
request.
5. Start a Codex helper before its thread lineage arrives.
6. Resume a delegated task and inspect its existing children and saved
output.

**Paperclip version or commit**

The live failures were found at `f2c5e54dc`. This branch is rebased onto
`86b7ee992`.

**Deployment mode**

Isolated local development instance with native Codex and ACPX Claude.
No database migration or default permission change.

Related work: Refs #13284 for Agent Chat. Refs #13438 for the
server-side API receipt fix, which this branch preserves. The transport
also accepts the earlier HTTP receipt format. Refs #13655 for the
current Codex continuation and helper lineage handling, which this
branch also preserves.

## What Changed

- Preserve failed Agent Chat wake-comment IDs and session generation
from the authorized source run. Reject pre-reset retries.
- Wrap API receipts with the correct semantic call identity. Test
current and earlier receipt formats through real HTTP and runnerd.
- Classify early child MCP startup notifications as information. Keep
foreign completion and result events rejected.
- Constrain project icons on both tool surfaces and regenerate the
protocol contracts.
- Include bounded, company-scoped visible direct child tasks in task
context. Filter hidden tasks before applying the limit.
- Replace the rejected Retry action with Inspect run for native
continuation reconciliation.
- Update hiring, review handoff, status reporting, and development
guidance.

## Verification

- Live tests covered Claude and Codex questions, plan revisions,
approval, task creation, hiring, status, chat reset, failures, and
recovery.
- The recovered task produced its saved checklist and example. A later
follow-up read the existing child tasks and document without creating
more work.
- Full build, repository type checks, token gates, 142 focused tests,
188 runner TypeScript tests, and the Rust notification/descendant
regressions passed after rebase. The separate local full-suite run was
stopped after the complete CI suite passed.
- Review fixes passed the updated route, tool-authority, and icon
regression tests plus server type checking.
- Required commands: `pnpm build`, `pnpm -r typecheck`,
`PAPERCLIP_IN_WORKTREE=false pnpm test:run`, and `pnpm
check:token-gates`.
- At `4ce8047b0`, all 55 applicable GitHub checks pass (two Storybook
checks are intentionally skipped), including the complete
general/serialized test matrix, runner tests, browser tests, build, type
checks, Docker checks, and canary dry run.
- Fresh Greptile review is 5/5 on `4ce8047b0`; all three findings were
fixed with regressions and there are no unresolved review threads.
- Two initial CI service-startup timeouts passed unchanged in local
reproductions and in the latest CI run.

## Risks

- The new event classification is limited to MCP startup information. It
does not authorize foreign task completion, results, or tool requests.
- Task context returns at most 100 direct child tasks and reports
truncation. It excludes hidden tasks and other companies. This improves
delegation context but does not enforce semantic duplicate detection.
- Native reconciliation still requires an operator to inspect and record
prior outcomes. The new link does not replace the recovery API.
- API tools remain opt-in. Claude permission choices remain explicit. No
default permission, schema, or workflow changes.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository editing, code
execution, API tools, and browser testing. The exact deployment
identifier and context-window size are not exposed in this session. Live
acceptance agents used OpenAI `gpt-5.6-sol` and Anthropic
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:11:04 -05:00
DottaandPaperclip 9335b7db10 fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance.

Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale.

Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-19 08:51:18 -05:00
86b7ee992c feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents have a persistent visual identity (#13171): a ClipLab
character in one of 17 palettes, rendered as cached PNGs in lists and as
a live character in larger placements.
> - The onboarding wizard is where a person meets that identity first,
and it showed a stock ClipLab expression on the previous engine while
the rest of the app would show a different character on a newer one.
> - The wizard's steps also cut from one screen to the next, so the arc
read as separate pages rather than one walk.
> - This pull request puts one character on one engine everywhere, gives
the wizard's hero the studio's sleepy → wink → idle sequence on Review,
and hands the steps over inside one presence.
> - The benefit is that what wakes on Review is exactly what the agent
looks like on the dashboard afterwards, and the walk to it reads as one
screen changing.

## Linked Issues or Issue Description

Refs #13171, now merged into master. This PR contains the onboarding and
ClipLab update on top of that foundation. Original feature work by
@tonio-alucema; merge preparation preserves the original commits.

**Problem or motivation**

The onboarding hero and the app's avatars were two different characters
on two different ClipLab engines. Steps 1 → 4 of the wizard cut between
screens, and the wizard mounted cold when a cloud-managed workspace
arrived from Cloud's naming screen.

**Proposed solution**

Vendor ClipLab v0.2.0 as the shared engine and render one
studio-exported character from it in every palette, for every pose and
size. Play the export's one-shot wake on Review with the palette fading
in over the gray dormant loop. Hand steps over inside one presence so
the footer slides instead of jumping, and play the arrival half of that
hand-off when the wizard opens directly on the agent step.

**Alternatives considered**

Exporting mp4/webm loops per size: no cursor following, no clean alpha,
and the palette "colour in" is a runtime blend. Minting a `cap-v2`
character version: nothing had shipped `cap-v1`, so the artwork is
regenerated in place instead of migrated. Keeping the separately
vendored runtime bundle for the hero: two engines and two characters in
one app.

## What Changed

- `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0
(`987b6db0`) with the Paperclip adaptations replayed (optional graphics
backend for the Node SVG snapshot path, supersampled live textures,
character framing, deterministic SVG id prefixes); new upstream
`particles.ts`.
- `packages/shared/src/cliplab/character.ts`: the studio export,
mirrored from `ui/src/assets/cliplab/onboarding.character.json` by
`scripts/sync-cliplab-character.mjs` (drift caught by
`check:token-gates`). `characterDefinition` builds every palette from
it; the resting portrait is its idle beat.
- `OnboardingCharacter`: gray `sleepy` loop through the agent and
connect steps; on Review the one-shot sleepy → wink → idle plays on two
lock-step canvases while the palette fades in, then the `idle` loop.
Body-follows the pointer, page-scoped. 160px in the wizard.
- `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one
`AnimatePresence` (departing content fades and gives its room back;
arriving content opens its room then fills); the hero has a room that
opens on the walk into the agent step; opening directly on the agent
step plays the arrival half; the self-hosted naming step uses the arc's
label and field.
- Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`,
`ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`).
- Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent
arc` walkable from the naming step plus `Arrive From Cloud`; the
companies fixture answers the wizard's create call with a company.
- Uses the shared runtime for onboarding; `doc/agent-personas.md`
documents the shared character.
- Releases both onboarding canvases after partial startup or transition
failure. Registers each canvas before seeking so synchronous render
errors can release it. Six component tests cover these failures and
palette changes before or during wake.
- Refreshes both sleeping canvases when the palette changes, including a
palette change in the same render as wake.
- Moves choreography values into the CSS token layer and preserves the
shared motion catalog drift check across the imported stylesheet.
- Repairs the static Storybook avatar route and uses accessible heading
names/current button labels in the wizard play functions.
- Closes the lazy avatar worker pool during application shutdown.

## Verification

- Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm
build-storybook`, and `pnpm check:token-gates` pass. The final UI
typecheck and 123 focused onboarding, lifecycle, and token catalog tests
pass. All 55 checks on final head
`b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full
sharded test suite, runner verification, and all eight browser shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)).
The duplicate monolithic local `pnpm test:run` was stopped after CI
completed; it is not claimed as a separate completed local run.
- Chromium walkthrough: palette change, wake, return to sleep, WebGL
failure fallback, Review step hand-offs and cloud arrival pass with
normal and reduced motion; no browser errors. The signoff happy-path
browser test also passes against a disposable instance.
- The final CI run confirms the catalog fix and a passing signoff
browser shard. The earlier signoff failure was a heartbeat-run
availability timeout; the focused local reproduction and final CI passed
without signoff code changes.
- Original author verification:
- `pnpm check:token-gates` (includes the new character sync check);
shared, server avatar/persona (17) and UI onboarding/persona (137)
suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs
through the worker pipeline.
- Storybook: `Agents / Personas` Sizes, Expressions and Palettes render
the studio character at every size and pose; `Onboarding / Character →
Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc`
walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the
arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by
~560ms, footer travel continuous).
- The original author walked the agent → connect → review flow and wake
after a real sign-in on staging.
- Not done here: the Linux Storybook visual baselines
(`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for
the new engine, hero size and naming-step changes.

## Risks

- Every avatar's pixels change (new engine, new character) under the
unchanged `cap-v1` name. Stacks that rendered avatars on the previous
engine keep those PNGs in their cache
(`generated-agent-avatars/cap-v1/...`, served immutable) until cleared;
only the two pinned staging stacks ever did.
- The one-shot handoff to the idle loop is timed from the sequence's
authored duration (the engine reports completion by continuing into idle
itself); presentation only, nothing in the wizard's state waits on it.
- Reduced motion skips the wake and the hand-offs; jsdom is treated the
same way, so the wizard tests see the next step's content immediately.
- The committed export differs from the studio by one animation (Loop
off, leading idle step removed); a re-export without that fix would play
a 5.6s idle before the wake.

## Model Used

Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in
Claude Code, with shell, browser, and file tools. The original context
window was not recorded.

Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex,
with reasoning, shell execution, file editing, GitHub CLI, and automated
tests. The session does not expose an exact runtime model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 08:30:14 -05:00
DottaandPaperclip c9e8677979 docs: document chat connector UX and make the runbook self-contained (#13675)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Connections let agents work with external services.
> - Contributors use the connection runbook to add and review providers.
> - The runbook depended on private issue references and did not capture
the chat setup UX conventions.
> - This pull request adds a chat connector UX guide and puts the
missing requirements in the runbook.
> - Contributors can apply the guidance without access to the internal
issue tracker or a personal skill installation.

## Linked Issues or Issue Description

**Issue type**

Missing documentation and unclear contributor instructions.

**Where is the issue?**

`doc/connections/CONNECTOR-PLAYBOOK.md` and the chat connector setup
guidance.

**What's wrong?**

The runbook sent contributors to private issues for validation, OAuth
ownership rules, and catalog review. The Slack setup work also
established useful UX rules that other providers should share.

**Suggested fix**

Add a companion UX document. Link it from the runbook. Include
validation and review requirements directly in the public documentation.
Use portable company and instance examples.

Related implementation:
https://github.com/paperclipai/paperclip/pull/13638. A search of related
PRs found no duplicate documentation change.

## What Changed

- Add `CHAT-CONNECTOR-UX.md` with setup, credential, identity, footer,
test, and management conventions.
- Include adaptation examples for Discord, Telegram, and email
providers.
- Link the companion from the runbook introduction, contents, and UX
section.
- Replace private issue references with inline architecture boundaries,
risk classification, validation evidence, and per-tool review
requirements.
- Replace personal deployment examples with sample company and instance
addresses.
- Mark the recorded Notion provider observations as a dated snapshot.

## Verification

- Passed `git diff origin/master --check`.
- Passed local validation of all 35 relative links and heading anchors
in the two documents.
- Passed code-fence and private-reference scans.
- Reviewed the guide against the Slack setup decisions and the existing
connection docs.
- Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. They
could not complete because this fresh worktree has no installed
dependencies (`@types/node` and the CLI `tsx` entry point are missing).
No application code changed.
- GitHub CI passed on commit `070fefc8b2513cb9469bae76e66940c33c1465ea`,
including typecheck, build, general and serialized tests, runner
verification, browser tests, and canary dry run.
- Security checks passed. Greptile scored the current commit 5/5 with no
findings or unresolved review threads.

## Risks

Low risk. This changes two documentation files only. Provider APIs and
screens can change. The guide requires authors to verify provider
capabilities, and the Notion example identifies its observation date.
The documentation does not assert new runtime support.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and context
window are not exposed in this session. Used reasoning, repository
inspection, shell execution, and documentation editing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal issue
id or instance-derived details
- [x] I have run the applicable documentation checks locally and they
pass; application checks were attempted and the environment limitation
is recorded above
- [x] I have added or updated tests where applicable (documentation
validation; no runtime changes)
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 08:16:56 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00