Commit Graph
1935 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 83abfa46f6 feat(ui): enable Grok in Cloud agent setup (#13791)
## Thinking Path

> - Paperclip manages AI agents and their execution settings.
> - The new-agent flow selects an adapter before configuring credentials
and a model.
> - Cloud uses one adapter policy for the picker and direct setup links.
> - That policy excludes Grok despite its existing adapter and xAI
connection support.
> - This change adds Grok to the Cloud policy.
> - Cloud users can configure Grok with the existing subscription or
API-key flow.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent creation on Cloud.

**Current behavior**

The Cloud picker offers Claude, Codex, and OpenCode. Direct Grok setup
links also fail the shared adapter check.

**Proposed behavior**

Offer Grok alongside the existing choices. Use the existing xAI
connection and managed sandbox setup.

**Reason and benefit**

Users can select an already-supported adapter through the Cloud creation
flow.

**Breaking changes**

None. Loaded and enabled checks still apply. Other excluded adapters
remain excluded.

## What Changed

- Add `grok_local` to the shared Cloud creation policy.
- Display four adapter choices in a 2×2 grid on desktop and mobile, and
reuse the theme-aware provider mark on the connection step so Grok is
visible in dark mode.
- Verify picker navigation and both Grok authentication methods through
sandbox setup, model testing, and agent creation.
- Verify that probe and hire payloads carry the xAI connection binding
without the entered API key.
- Update the agent configuration specification.

## Verification

- `cd ui && pnpm exec vitest run src/components/NewAgentDialog.test.tsx
src/pages/NewAgent.test.tsx
src/components/new-agent/AgentProviderConnection.test.tsx` — 69 tests
passed.
- Chromium checks against the production components and built stylesheet
— four cards occupy two rows and two columns at 1280px and 390px; the
connection step loads and displays the white Grok logo in dark mode and
the black logo in light mode.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/ui build` — passed.
- `pnpm check:token-gates` — passed.
- `git diff origin/master...HEAD | gitleaks stdin --redact --no-banner`
— no leaks found; manual diff review found no private identifiers or
user data.
- `pnpm -r typecheck` and `pnpm build` — blocked at the existing runner
package because `cargo` is not installed locally.
- `pnpm test:run` — started locally, then stopped after the equivalent
CI suites passed.
- Initial [PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35681351752)
passed, including full typecheck/build, general and serialized tests,
Rust checks, and all eight browser-test shards. One unchanged Cursor
test timed out on the first attempt; its five-test file passed locally
and the failed CI shard passed on retry.
- The [latest CI
run](https://github.com/paperclipai/paperclip/actions/runs/35683642812)
passed build, typecheck, all general and serialized tests, Rust checks,
and all eight browser-test shards. One unchanged
local-service-supervisor test failed its HTTP readiness check on the
first attempt; its six-test file passed locally, and the failed server
shard passed on retry.
- Live xAI login and model execution were not run; the setup tests mock
provider calls.

## Risks

Small UI policy change. The existing Grok adapter, authentication, and
secret storage paths remain in use. No schema or control-plane change is
required. Cloud must deploy a tenant-app release containing this change.
Revert the policy entry to hide Grok from new-agent setup again.

## Model Used

- OpenAI GPT-6 (Codex), with repository inspection, code editing, and
shell-based verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 09:17:59 -07:00
DottaandPaperclip 8725d6ce09 fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions.

Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:35:33 -05:00
Devin FoleyandPaperclip 6de50ba594 fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can enable Sentry for the server and the signed-in
browser.
> - The server SDK reads `SENTRY_ENVIRONMENT` from the process
environment.
> - The browser receives its DSN through the session response, but
receives no environment.
> - A browser in staging therefore reports errors under the SDK's
production default.
> - This pull request passes the configured environment through the
existing session and monitoring gate.
> - Browser errors then identify the deployment environment while
preserving the existing privacy settings.

## Linked Issues or Issue Description

**What happened?**

With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged
`production`. This can send errors to the wrong environment's alerts and
makes deployment follow-up unreliable.

**Expected behavior**

The browser uses the server's configured Sentry environment. A reused
image works in either staging or production. A signed-out browser still
sends no events.

**Steps to reproduce**

1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`.
2. Sign in and capture a browser exception.
3. Inspect the event environment. Before this change, it is
`production`.

**Paperclip version or commit**

Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real
browser SDK and a local test transport.

**Deployment mode**

Authenticated server and browser with optional Sentry monitoring
enabled.

No duplicate environment-attribution issue or pull request was found in
the targeted GitHub search.

## What Changed

- Add `sentryEnvironment` to the authenticated session response and
shared schema. The optional field supports a newer browser reading an
older server response.
- Pass the environment to the browser SDK. An environment change
restarts the client through its existing serialized lifecycle.
- Cover environment attribution with a real SDK event, session
authorization, unchanged-session refetches, environment changes, and
legacy responses.
- Document configuration and compatibility. Keep the loaded bundle's
release identity and existing privacy filters.

## Verification

- The regression test emits `production` for a requested staging
environment before the fix.
- Focused route, schema, browser lifecycle and real-SDK tests: 69 pass.
- UI and shared-package typechecks, direct server `tsc --noEmit`, and
token gates pass.
- Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at
the Runner Rust step because `cargo` is absent on this machine.
- Complete UI suite: 6,540 tests pass in 626 files.
- Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files
fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache
`EACCES`. These match the existing local baseline; none touch the
changed behavior.
- Greptile: 5/5, no unresolved review threads. Linux CI has passed
Build, Typecheck + Release Registry, and the completed test jobs so far.
Remaining jobs are running or queued: the AWS runner provisioner is
retrying EC2 CreateFleet `InternalError` responses. Full results will be
recorded before merge.

## Risks

Low risk. This adds one optional session field and changes Sentry
attribution only. No migration or new monitoring opt-in is introduced.
Missing settings keep the browser SDK default. Agent and unauthenticated
requests still receive 401 without monitoring settings. Existing loaded
browser bundles keep their old behavior until refreshed.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused and full UI suites
pass; full-root environment failures documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 23:40:46 +00:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00
Devin FoleyandPaperclip a3749aac46 fix(deps): share Lezer node properties across editor languages
fix(deps): share Lezer node properties across editor languages

The installed graph gave syntax highlighters @lezer/common 1.5.1 and
language parsers 1.5.2. Their independent NodeProp counters collided,
so highlighting ordinary code read unrelated metadata as tags and
crashed with tags-is-not-iterable.

Override @lezer/common to one compatible version in both manifests.
Extend the installed-graph check and exercise Python, JavaScript, HTML
and SQL highlighting through the editor dependencies. All four examples
failed before the override and pass with it.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 15:04:50 -07:00
Devin FoleyandPaperclip df66219780 fix(ui): retry the latest failed task attempt (#13765)
## Thinking Path

> - Paperclip manages work done by AI agents.
> - The task thread lets an operator retry a failed run.
> - Legacy runs with transcript output did not get a failure marker.
> - The thread could therefore offer Try again for an older failure.
> - The server correctly reused that failure's existing retry, even when
it had already failed.
> - This change keeps the latest failure actionable and reports stopped
retry responses to the operator.

## Linked Issues or Issue Description

**What happened?**

Try again could return success without starting work. An initial setup
failure had an empty transcript. Its later retry produced output and
failed. Only the initial failure had a retry marker, so the button kept
requesting the initial failure's already-failed successor.

**Expected behavior**

Try again targets the latest failed attempt. An already-stopped retry
response shows an error and refreshes the task's run state.

**Steps to reproduce**

1. Start a legacy adapter task that fails before producing transcript
output.
2. Retry it. Let this attempt produce output and a final failure comment
before it fails.
3. Click Try again in the task thread.
4. Before this fix, the click targets the original failure and replays
the stopped successor.

**Paperclip version or commit**

Reproduced against `1483bb8bcf`; the regression is also present on the
branch base `8813a50105`.

**Deployment mode**

Authenticated server with a legacy adapter. The bug is in the shared
task UI and retry API client.

Related: #11650 adds a different recovery-notice action. This change
fixes failed-run markers and retry response handling. Searches found no
duplicate of this failure case.

## What Changed

- Render legacy failure markers even when the run has a transcript or
final comment.
- Keep later cancelled automatic retries from replacing the failed run's
retry action.
- Reject already-stopped retry responses in the API client so existing
error feedback appears.
- Refresh task run queries after both successful and failed retry
requests.
- Add regression coverage for failed and timed-out attempts, execution
gates with output, and retry response states.

## Verification

- Before the fix, the new regression tests failed: two selected the
original failure, and four accepted a stopped successor as success.
- Targeted task-thread, retry API, marker, and issue-page tests: 279
passed.
- `pnpm --filter @paperclipai/ui typecheck`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm check:token-gates`: passed.
- Full UI suite: 6,529 tests passed across 626 files.
- [CI run
35650385023](https://github.com/paperclipai/paperclip/actions/runs/35650385023):
all 53 checks passed, including full workspace build, typecheck,
unit/integration suites, and browser tests. The redundant local
full-workspace test run was stopped after CI passed; it is not counted
as a completed local pass.
- Greptile: 5/5 on `608ee58c99`, with no review threads or unresolved
comments. The branch is mergeable.
- `pnpm -r typecheck` and `pnpm build` were attempted. Both stop at the
Runner's Rust checks because this host has no `cargo`. The full
workspace checks passed in CI.

## Risks

Low risk. This changes UI presentation and response handling only. The
server's exact-retry idempotency, authorization, execution ownership,
and recovery gates remain in place. No schema changes or live task
mutations. Existing documentation describes this retry action; the fix
restores that behavior.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(existing behavior restored; no documentation change needed)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:31:43 -07:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
DottaandPaperclip 57fd8b70d2 feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack connections let a team talk to those agents in Slack.
> - Agents now have a saved avatar, but Slack setup did not offer that
image.
> - A matching avatar helps a team recognize its agent.
> - This pull request adds an optional avatar step and a download in
connector Settings.
> - Users download a PNG and upload it directly in Slack with clear
instructions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Slack connector onboarding and its Settings page.

**Current behavior**
Setup does not offer the assigned agent's avatar or explain how to
upload it in Slack.

**Proposed behavior**
After Slack connection verification, users can download a 512 × 512 PNG
of their agent's saved avatar. They can upload it in Slack, confirm, or
skip. Settings keeps the download and upload instructions available
after onboarding.

**Reason and benefit**
The same avatar helps people recognize the agent across Paperclip and
Slack. Users who skip the optional step can return to it in Settings.

**Breaking changes**
None. No schema, authentication, Slack scope, or provider API change.
Completed connections keep their existing completion state. Searched
existing Slack avatar and Cliptoon PRs; no matching implementation was
found.

## What Changed

- Add an optional avatar step before personal Slack account linking.
Keep the numbered sidebar and shared footer.
- Resolve the selected agent's saved appearance for the preview and PNG
download.
- Add the same download and expandable upload instructions to connector
Settings.
- Remember uploaded or skipped per company and endpoint in browser
storage. Treat uploaded as user confirmation, not provider verification.
- Reject failed or non-PNG download responses and allow retry.
- Reuse the production avatar components in onboarding and Settings
stories.
- Test wizard progression, resume, Settings, download recovery, storage
isolation, and terminated assigned agents.
- Exercise real PNG downloads in the Slack browser flow and keep default
app names consistent with app creation.
- Fetch the assigned agent directly so its saved avatar remains
available after termination.

## Verification

- Focused chat suites: 48 passed; the two affected suites passed again
after the final naming fix (32 tests).
- Slack browser E2E passed through setup, avatar download, account
linking, and Settings download. PNG signature and 512 × 512 dimensions
verified.
- UI token gates passed.
- Browser: downloaded the real 512 × 512 PNG; checked confirmation,
return, mobile layout, and Settings instructions.
- Full workspace typecheck, application build, and production Storybook
build passed.
- All latest-head CI checks passed (54 passed, 2 skipped), including all
browser, chat, general, and serialized test groups. The unrelated Sentry
test failed once and passed on the single CI rerun; its suite also
passed locally.
- Local full-suite attempt encountered a rapid Slack callback ordering
failure under concurrent build load; that test passed in isolation, and
all three chat shards passed in CI. The remaining local run was not used
as the merge gate.
- Review the Connections / Slack / Add avatar and Avatar in Settings
stories.

## Risks

- Slack upload is manual. Confirmation does not claim to verify the
Slack icon.
- Optional step progress is browser-local. Clearing storage or changing
browsers can show it again. Setup still works when storage is
unavailable.
- The existing avatar API remains the image source. Download failures
show a retry message.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 09:38:48 -05:00
DottaandPaperclip fd071748ee fix: stop repeated notifications for finalized run failures (#13739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native run reconciliation repairs saved execution outcomes after
interruptions.
> - The sweep also visits runs whose final results are already
committed.
> - An unchanged failed run still received a new status delivery ID on
every sweep.
> - The browser treated each delivery as a new failure after its short
duplicate window expired.
> - This pull request makes unchanged projections a no-op and suppresses
repeated or historical run toasts.
> - Operators receive fresh failure alerts without repeated alerts for
old work.

## Linked Issues or Issue Description

**What happened?**

An old failed run repeatedly produced failure toasts while the browser
remained open. The task could already be cancelled. Reconciliation
rewrote the same failed outcome and queued another status broadcast.

**Expected behavior**

An unchanged committed run must not queue a new status notification.
Repeated deliveries must still refresh cached state without another
toast.

**Steps to reproduce**

1. Finalize a native run with a failed result and successful workspace
finalization.
2. Deliver its pending execution status and cancel its task.
3. Replay finalization and status delivery on each periodic sweep.
4. Observe another failure broadcast for every sweep before this fix.

**Paperclip version or commit**

Reproduced against source commit c65fc9e3c. The new regression produced
three extra broadcasts in three sweeps before the fix.

**Deployment mode**

Authenticated hosted deployment. The database regression also reproduces
with isolated embedded PostgreSQL.

Related context: #13104 established contextual run notifications. This
change fixes duplicate delivery after finalization. No matching open fix
was found.

## What Changed

- Add a database predicate that skips unchanged committed run
projections. Preserve real status repairs, pending delivery IDs, and
ownership checks.
- Track observed company/run/status outcomes in the live update
provider. Keep this bounded history across socket reconnects.
- Suppress terminal toasts delivered more than five minutes after
completion, using server timestamps.
- Add regressions for cancelled tasks, concurrent replay, incomplete
projection repair, repeated broadcasts, historical failures, and
reopening a window.
- Document the notification behavior in DESIGN.md.

## Verification

- 103 focused server and UI tests passed before rebase. The new server
regression failed on the original implementation and passed with the
fix.
- Server and UI TypeScript checks passed before rebase.
- UI production build and `pnpm check:token-gates` passed before rebase.
- Full workspace `pnpm -r typecheck` passed locally after rebase. Full
tests, typecheck, and build passed in CI on the exact PR head; the
redundant local full-suite run was stopped after CI completed.
- All current-head CI checks passed, including all eight browser shards.
Browser shard 4 passed on one unchanged-code rerun after an initial
chat-page navigation timeout (blank screenshot, slow loads, no
JavaScript crash).
- Greptile reviewed the current head at 5/5 with no actionable findings
or open review threads.

## Risks

- The UI intentionally omits transient toasts for outcomes delivered
more than five minutes after completion. Run history and cache updates
remain available.
- The observed-outcome cache holds at most 2,000 identities. Historical
timestamps also protect fresh page loads.
- No schema changes. Recovery still repairs missing completion metadata
and stale successful-run errors. Existing ownership gates remain
enforced.

## Model Used

OpenAI GPT-6 through Codex, with code editing, shell tools, browser
inspection, and test execution. The exact deployment identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 09:31:50 -05:00
Devin FoleyandPaperclip c65fc9e3c8 fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures

Include connect timeouts in the bounded retry policy for idempotent actor
synchronization. Handle WebSocket constructor failures through existing
reconnect paths and preserve HTTP polling while realtime is unavailable.
Refresh visible company queries until the socket recovers and clear all
fallback timers on hiding or unmount.

Verify 172 focused tests, server/UI typechecks, UI build, and design token
gates. Full workspace build/typecheck require the unavailable Rust toolchain;
the full test run is tracked separately.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 12:43:25 -07:00
Devin FoleyandPaperclip 600e552d7b fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build.
Use validated build commits for Docker and source/npm artifacts, preserve
explicit server release overrides, and keep cached browser bundles tied
to the commit they loaded.

Verify 127 focused tests, server/UI typechecks, Docker and source build
stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 08:08:33 -07:00
DottaandPaperclip da257c3069 Warn when routine webhook URLs may not be publicly reachable (#13684)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Routines can start that work when another app sends a webhook.
> - Local and private URLs often cannot receive events from public
services.
> - HTTPS alone does not make a Tailscale address public.
> - This pull request explains these limits during setup and editing.
> - Users can still finish setup for senders on their own network.

## Linked Issues or Issue Description

Refs #13637. Webhook setup needs a clear warning when the generated URL
appears local, private, or unencrypted. The warning must explain how to
make the endpoint reachable without blocking private-network use.

## What Changed

- Add a shared warning banner to the Connect, Check connection, and Edit
webhook views.
- Distinguish localhost, private network addresses and domains, HTTP,
and Tailscale hostnames.
- Explain the difference between Tailscale Serve and Funnel. Link to the
Paperclip HTTPS guide.
- Add five full-page Storybook examples, design guide examples, and
documentation.
- Add URL classification tests and a regression test that finishes setup
despite the warning.

## Verification

- Passed 41 focused URL and trigger-flow tests.
- Passed workspace typecheck, workspace build, token gates, and
Storybook build.
- Browser-tested the Tailscale story through Check connection, Finish
setup, and Edit webhook. The warning stays visible and does not block
setup.
- Open Product / Routines / Webhooks stories 11–15 to review the warning
states.
- All 54 PR checks passed, including the full test matrix and eight
browser shards; two optional Storybook jobs were skipped by workflow
policy.
- The server supervisor readiness test timed out once in CI, then passed
on rerun and locally (6 tests).
- The duplicate local full-suite run was stopped after the complete CI
matrix passed.
- Greptile: 5/5 on the current commit, with no unresolved review
comments.

## Risks

- URL checks are hints. They do not test DNS, firewall rules, or actual
reachability.
- A Tailscale hostname can serve either private Serve traffic or public
Funnel traffic. The warning explains this uncertainty and permits both.
- No API, schema, authentication, or webhook delivery behavior changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 12:24:06 -05:00
Devin FoleyandPaperclip b70641f23f feat(plugins): support image catalogs and persistent application overlays (#13646)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins extend the application without adding each integration to
Core.
> - A downstream image needs a way to supply prebuilt plugins.
> - Some plugin UI must stay mounted as users move between pages.
> - This change adds an image catalog and a persistent application slot.
> - Operators can upgrade or remove these plugins through their image
and configuration.

## Linked Issues or Issue Description

**Subsystem affected**

Plugin packaging, activation and application UI.

**Problem or motivation**

The built-in plugin catalog is fixed in Core source. Downstream images
cannot add entries through an explicit catalog. Existing page slots also
cannot preserve a small application overlay across route changes.

**Proposed solution**

Read a bounded catalog of prebuilt plugins from the image. Verify its
files before importing manifests. Use the existing managed selection and
plugin lifecycle. Add an `appShellOverlay` slot with account and company
cleanup.

**Alternatives considered**

A downstream fork adds merge work. Script injection provides no plugin
lifecycle. A separate runtime download system adds a second distribution
channel.

**Roadmap alignment**

This extends the existing plugin system. Related PR #9006 covers runtime
install replication; this change covers immutable image contents. PR
#12555 covers CLI scaffolding. Neither provides this catalog or
application slot. The maintainer requested this work directly.

## What Changed

- Validate catalog identities, confined paths, package versions and
bundle hashes before importing code.
- Apply image selection to persisted plugin installs, including removal
and rollback. Adopt the verified image path from existing npm/local
installs and bind runtime worker/UI entrypoints to verified package
declarations.
- Mount application overlays in both UI shells. Preserve route state and
clear it on account, company and onboarding changes.
- Restrict service-worker offline storage/fallback to hashed public
assets in a separate cache namespace; exclude application HTML and
extension/API data, including after worker restart.
- Document the packaging contract, trust model and rollback
requirements.

## Verification

- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Affected server/UI typechecks and builds, plus token
gates, passed again after rebasing onto current master; the 124 focused
tests also passed after rebase.
- Latest focused verification: 124 tests in nine files passed for
catalog/reconciliation/loader, overlay lifecycle, Layout and
service-worker policy. The broader UI/shared/SDK run passed 7,204 tests
in 690 files with canonical `TMPDIR`.
- Real disposable Core/PostgreSQL: catalog install, selection removal,
0.1.0→0.1.1→0.1.0, same-version npm/legacy-path adoption, and
preservation of disabled status passed. Added permissions entered
`upgrade_pending`, withheld UI across restart, and activated only after
explicit operator enable.
- Real Chromium: desktop/mobile layout, route draft retention and
Escape/focus passed with mocked extension responses. A persistent
browser restart retained public hashed-asset offline fallback while
refusing seeded legacy/current private entries and legacy HTML.
- Full `pnpm test:run`: 12,539 passed; 17 failed across six existing
files, stopping later phases. macOS read-only directory renames fail in
runtime-skill-cache and company-skills-service; email tests require an
absent local AgentMail fixture. Native runner/comment-redaction passed
in isolation after temporary Rust setup; agent-conversations also passed
in isolation. No unrelated source was changed to hide failures.
- After rebase, two unchanged chat timing tests failed in CI and passed
locally in isolation. Their CI shard passed on its single retry. All
other current-head CI jobs passed on the initial run; review is 5/5 with
no unresolved threads.
- No live deployment or external plugin service was used.

## Risks

- Plugins are trusted code. The catalog detects packaging errors; it
does not authenticate an untrusted image builder.
- Invalid catalogs fail startup. Images must contain the catalog and
bundles together, with stable directories.
- A host older than this contract lacks the activation guard. Disable
added plugins and remove their configuration keys before reverting to
it.
- Offline navigation now returns 503 instead of replaying cached
application HTML. Only public build assets have offline fallback.
- Rolling back an unapproved permission change retains the approval
gate; review the current manifest and explicitly enable it. A reduced
permission set cannot establish prior approval or prior enabled status.
- Plugin data migrations need their own rollback policy. This change
retains installed records and does not reverse migrations.

## Model Used

- OpenAI GPT-6 (Codex), model ID `gpt-6`, with repository inspection,
code execution and browser verification. The runtime does not expose an
exact context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites; broad
macOS server-run exceptions are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (fresh run on 488b3754ae; chat
shard passed its single retry)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(fresh review on 488b3754ae)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:18:10 -07:00
DottaandPaperclip f589660ec0 feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Routines turn scheduled work and external events into tasks for an
assigned agent.
> - Webhook setup was disabled, and actor authentication rejected valid
webhook bearer keys.
> - Operators need to connect and test a sending app before events can
start work.
> - This pull request adds a guided setup with durable connection tests
that cannot dispatch a task.
> - It keeps trigger management, execution tasks, and activity within
the routine.
> - The benefit is a webhook that can be configured, verified, and
operated from one place.

## Linked Issues or Issue Description

Fixes #11937.

Related: #13216 adds provider-specific Sentry support. This PR addresses
general routine setup and ingress. #6841 addresses legacy secret
bindings; this PR retains the existing secret service.

**Current behavior**

Webhook creation is disabled. Bearer deliveries can fail in agent
authentication before the routine checks its key. Setup has no safe
connection test. Runs and Activity send the operator away from the
routine.

**Proposed behavior**

Choose a schedule or a webhook. Follow the setup steps, copy credentials
or complete agent instructions, and test delivery without creating work.
Finish setup to allow future events to start tasks. Edit or remove
compact trigger cards, undo removal, and inspect tasks and activity
inside the routine.

**Reason and benefit**

An operator can verify credentials and delivery before enabling
automatic work. Durable setup state survives refreshes and restarts.
Retry receipts prevent an old test event from starting work after
activation.

## What Changed

- Add a production trigger wizard using reusable Slack setup navigation
and footer components.
- Add schedule and webhook choices, one-time credentials, agent
instructions, and live connection feedback.
- Persist pending setup, test delivery receipts, connection status, and
reversible trigger removal.
- Keep setup checks free of routine runs, tasks, and agent wakeups.
Preserve delivery idempotency after activation.
- Add compact trigger cards, inline editing, key rotation, pause
controls, removal, and Undo.
- Keep Runs and Activity in the routine. Use the shared task list and
compact activity rows.
- Permit only exact public delivery POSTs through actor authentication.
Retain webhook authentication, JSON-object validation, and log
redaction.
- Add production-backed Storybook states and focused server, database,
and UI coverage.
- Document signing modes, setup checks, retries, rotation, HTTPS
ingress, and navigation.

## Verification

- Full workspace typecheck, build, and token gates passed on the rebased
branch. Storybook also builds.
- Focused routine, middleware, logging, shared wizard, and UI coverage
passes on the rebased branch: 195 tests across 14 files. The migration
passed on a fresh PostgreSQL database and on two repeated applications.
- Browser testing used the real app, database, and a deterministic
process worker through Tailscale HTTPS and the current Cloud proxy code.
- Verified rejected keys, safe setup deliveries, persisted state after
restart, activation, retry deduplication, key rotation, schedule
editing, removal, and Undo.
- Fresh bearer and GitHub-signed deliveries created tasks that the
worker checked out and completed. Runs and Activity stayed within the
routine.
- Current Cloud ingress tests passed. Public delivery POSTs passed
through without a browser session; management routes remained gated.
- All 54 current-head PR checks pass, including general and serialized
tests, all eight browser E2E shards, typecheck, build, runner checks,
security checks, and the canary dry run. Two optional Storybook jobs are
skipped by workflow conditions.
- Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review
threads. The stale connection-status finding is fixed and covered by a
regression test.
- No production deployment was performed.

## Risks

- Migration 0281 adds three trigger columns and a test-receipt table. It
is additive and safe to reapply. Apply it before running the new server.
Existing triggers remain live by default.
- Requests without delivery IDs are new events after activation. Senders
must reuse an event's delivery ID for retries.
- Completed webhooks keep normal dispatch behavior. Their management
connection check can start work; the UI states this.
- Removing a trigger archives it. Undo restores the URL and credentials.
Permanent deletion remains available through the existing API.
- Public ingress must remain restricted to the delivery POST route. The
tenant verifies credentials. Cloud sleeping-stack behavior is unchanged.
- Shared setup components also serve Slack. Existing setup contracts and
navigation tests cover that integration.
- Senders must use application/json with an object. Other media types
receive 415.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 10:49:05 -05:00
DottaandPaperclip 36dbb7ed1c fix: harden agent chat runner tools and recovery (#13678)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat turns discussion into plans, tasks, reviews, and hires.
> - These workflows need reliable tool results and task context on the
native runner.
> - Live Claude and Codex tests exposed lost retry requests, invalid
project inputs, and a child startup crash.
> - Recovery also exposed a misleading retry action and missing child
task context.
> - This pull request fixes those paths and adds regression coverage.
> - Agents can continue the original request and operators can inspect a
stopped run.

## Linked Issues or Issue Description

**What happened?**

A failed Agent Chat retry could lose the user's question. Project
creation accepted unsupported icons in its tool schema. Codex could stop
when a helper's MCP startup event arrived before its thread lineage. A
stopped task offered Retry even when the server required execution
reconciliation. Resumed agents could miss existing delegated tasks.
Hiring and review instructions did not describe the native runner's
available tools and source requirements.

**Expected behavior**

Retries retain the selected request. Tool schemas match the API. Child
startup information does not gain authority over the parent or stop it.
Recovery actions match the server's requirements. Task context exposes
existing child work. Handoffs contain the material the assignee needs.

**Steps to reproduce**

1. Enable experimental Agent Chat in an isolated development instance.
2. Configure native Codex and ACPX Claude agents on Paperclip Runner.
3. Ask for a plan, revise it, approve task creation, and request a hire
and status report.
4. Retry a failed chat turn and check that it answers the original
request.
5. Start a Codex helper before its thread lineage arrives.
6. Resume a delegated task and inspect its existing children and saved
output.

**Paperclip version or commit**

The live failures were found at `f2c5e54dc`. This branch is rebased onto
`86b7ee992`.

**Deployment mode**

Isolated local development instance with native Codex and ACPX Claude.
No database migration or default permission change.

Related work: Refs #13284 for Agent Chat. Refs #13438 for the
server-side API receipt fix, which this branch preserves. The transport
also accepts the earlier HTTP receipt format. Refs #13655 for the
current Codex continuation and helper lineage handling, which this
branch also preserves.

## What Changed

- Preserve failed Agent Chat wake-comment IDs and session generation
from the authorized source run. Reject pre-reset retries.
- Wrap API receipts with the correct semantic call identity. Test
current and earlier receipt formats through real HTTP and runnerd.
- Classify early child MCP startup notifications as information. Keep
foreign completion and result events rejected.
- Constrain project icons on both tool surfaces and regenerate the
protocol contracts.
- Include bounded, company-scoped visible direct child tasks in task
context. Filter hidden tasks before applying the limit.
- Replace the rejected Retry action with Inspect run for native
continuation reconciliation.
- Update hiring, review handoff, status reporting, and development
guidance.

## Verification

- Live tests covered Claude and Codex questions, plan revisions,
approval, task creation, hiring, status, chat reset, failures, and
recovery.
- The recovered task produced its saved checklist and example. A later
follow-up read the existing child tasks and document without creating
more work.
- Full build, repository type checks, token gates, 142 focused tests,
188 runner TypeScript tests, and the Rust notification/descendant
regressions passed after rebase. The separate local full-suite run was
stopped after the complete CI suite passed.
- Review fixes passed the updated route, tool-authority, and icon
regression tests plus server type checking.
- Required commands: `pnpm build`, `pnpm -r typecheck`,
`PAPERCLIP_IN_WORKTREE=false pnpm test:run`, and `pnpm
check:token-gates`.
- At `4ce8047b0`, all 55 applicable GitHub checks pass (two Storybook
checks are intentionally skipped), including the complete
general/serialized test matrix, runner tests, browser tests, build, type
checks, Docker checks, and canary dry run.
- Fresh Greptile review is 5/5 on `4ce8047b0`; all three findings were
fixed with regressions and there are no unresolved review threads.
- Two initial CI service-startup timeouts passed unchanged in local
reproductions and in the latest CI run.

## Risks

- The new event classification is limited to MCP startup information. It
does not authorize foreign task completion, results, or tool requests.
- Task context returns at most 100 direct child tasks and reports
truncation. It excludes hidden tasks and other companies. This improves
delegation context but does not enforce semantic duplicate detection.
- Native reconciliation still requires an operator to inspect and record
prior outcomes. The new link does not replace the recovery API.
- API tools remain opt-in. Claude permission choices remain explicit. No
default permission, schema, or workflow changes.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository editing, code
execution, API tools, and browser testing. The exact deployment
identifier and context-window size are not exposed in this session. Live
acceptance agents used OpenAI `gpt-5.6-sol` and Anthropic
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:11:04 -05:00
86b7ee992c feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents have a persistent visual identity (#13171): a ClipLab
character in one of 17 palettes, rendered as cached PNGs in lists and as
a live character in larger placements.
> - The onboarding wizard is where a person meets that identity first,
and it showed a stock ClipLab expression on the previous engine while
the rest of the app would show a different character on a newer one.
> - The wizard's steps also cut from one screen to the next, so the arc
read as separate pages rather than one walk.
> - This pull request puts one character on one engine everywhere, gives
the wizard's hero the studio's sleepy → wink → idle sequence on Review,
and hands the steps over inside one presence.
> - The benefit is that what wakes on Review is exactly what the agent
looks like on the dashboard afterwards, and the walk to it reads as one
screen changing.

## Linked Issues or Issue Description

Refs #13171, now merged into master. This PR contains the onboarding and
ClipLab update on top of that foundation. Original feature work by
@tonio-alucema; merge preparation preserves the original commits.

**Problem or motivation**

The onboarding hero and the app's avatars were two different characters
on two different ClipLab engines. Steps 1 → 4 of the wizard cut between
screens, and the wizard mounted cold when a cloud-managed workspace
arrived from Cloud's naming screen.

**Proposed solution**

Vendor ClipLab v0.2.0 as the shared engine and render one
studio-exported character from it in every palette, for every pose and
size. Play the export's one-shot wake on Review with the palette fading
in over the gray dormant loop. Hand steps over inside one presence so
the footer slides instead of jumping, and play the arrival half of that
hand-off when the wizard opens directly on the agent step.

**Alternatives considered**

Exporting mp4/webm loops per size: no cursor following, no clean alpha,
and the palette "colour in" is a runtime blend. Minting a `cap-v2`
character version: nothing had shipped `cap-v1`, so the artwork is
regenerated in place instead of migrated. Keeping the separately
vendored runtime bundle for the hero: two engines and two characters in
one app.

## What Changed

- `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0
(`987b6db0`) with the Paperclip adaptations replayed (optional graphics
backend for the Node SVG snapshot path, supersampled live textures,
character framing, deterministic SVG id prefixes); new upstream
`particles.ts`.
- `packages/shared/src/cliplab/character.ts`: the studio export,
mirrored from `ui/src/assets/cliplab/onboarding.character.json` by
`scripts/sync-cliplab-character.mjs` (drift caught by
`check:token-gates`). `characterDefinition` builds every palette from
it; the resting portrait is its idle beat.
- `OnboardingCharacter`: gray `sleepy` loop through the agent and
connect steps; on Review the one-shot sleepy → wink → idle plays on two
lock-step canvases while the palette fades in, then the `idle` loop.
Body-follows the pointer, page-scoped. 160px in the wizard.
- `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one
`AnimatePresence` (departing content fades and gives its room back;
arriving content opens its room then fills); the hero has a room that
opens on the walk into the agent step; opening directly on the agent
step plays the arrival half; the self-hosted naming step uses the arc's
label and field.
- Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`,
`ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`).
- Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent
arc` walkable from the naming step plus `Arrive From Cloud`; the
companies fixture answers the wizard's create call with a company.
- Uses the shared runtime for onboarding; `doc/agent-personas.md`
documents the shared character.
- Releases both onboarding canvases after partial startup or transition
failure. Registers each canvas before seeking so synchronous render
errors can release it. Six component tests cover these failures and
palette changes before or during wake.
- Refreshes both sleeping canvases when the palette changes, including a
palette change in the same render as wake.
- Moves choreography values into the CSS token layer and preserves the
shared motion catalog drift check across the imported stylesheet.
- Repairs the static Storybook avatar route and uses accessible heading
names/current button labels in the wizard play functions.
- Closes the lazy avatar worker pool during application shutdown.

## Verification

- Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm
build-storybook`, and `pnpm check:token-gates` pass. The final UI
typecheck and 123 focused onboarding, lifecycle, and token catalog tests
pass. All 55 checks on final head
`b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full
sharded test suite, runner verification, and all eight browser shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)).
The duplicate monolithic local `pnpm test:run` was stopped after CI
completed; it is not claimed as a separate completed local run.
- Chromium walkthrough: palette change, wake, return to sleep, WebGL
failure fallback, Review step hand-offs and cloud arrival pass with
normal and reduced motion; no browser errors. The signoff happy-path
browser test also passes against a disposable instance.
- The final CI run confirms the catalog fix and a passing signoff
browser shard. The earlier signoff failure was a heartbeat-run
availability timeout; the focused local reproduction and final CI passed
without signoff code changes.
- Original author verification:
- `pnpm check:token-gates` (includes the new character sync check);
shared, server avatar/persona (17) and UI onboarding/persona (137)
suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs
through the worker pipeline.
- Storybook: `Agents / Personas` Sizes, Expressions and Palettes render
the studio character at every size and pose; `Onboarding / Character →
Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc`
walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the
arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by
~560ms, footer travel continuous).
- The original author walked the agent → connect → review flow and wake
after a real sign-in on staging.
- Not done here: the Linux Storybook visual baselines
(`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for
the new engine, hero size and naming-step changes.

## Risks

- Every avatar's pixels change (new engine, new character) under the
unchanged `cap-v1` name. Stacks that rendered avatars on the previous
engine keep those PNGs in their cache
(`generated-agent-avatars/cap-v1/...`, served immutable) until cleared;
only the two pinned staging stacks ever did.
- The one-shot handoff to the idle loop is timed from the sequence's
authored duration (the engine reports completion by continuing into idle
itself); presentation only, nothing in the wizard's state waits on it.
- Reduced motion skips the wake and the hand-offs; jsdom is treated the
same way, so the wizard tests see the next step's content immediately.
- The committed export differs from the studio by one animation (Loop
off, leading idle step removed); a re-export without that fix would play
a 5.6s idle before the wake.

## Model Used

Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in
Claude Code, with shell, browser, and file tools. The original context
window was not recorded.

Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex,
with reasoning, shell execution, file editing, GitHub CLI, and automated
tests. The session does not expose an exact runtime model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 08:30:14 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
Devin FoleyandPaperclip 6d03428682 Skip task-only connector reads for agent chat views (#13654)
Agent chat views reuse the task surface with synthetic chat-prefixed IDs.
Skip their task-only email and external chat-binding queries, and reject
invalid UUIDs after existing authentication checks on both read routes.
Preserve normal task reads and company isolation.

Verified failing regressions before the fix, all 6,377 UI tests, route and
OpenAPI regressions, server/UI TypeScript checks, all Linux PR CI gates,
and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 20:14:03 -07:00
DottaandPaperclip 924f07be8c feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connections let people start and continue that work from Slack.
> - Setup mixed app creation, credentials, URL verification, account
linking, and testing on the same screens.
> - People also needed a safe way to link their own Slack identity after
the first operator finished setup.
> - This pull request gives each step a clear place and keeps membership
approval separate from identity linking.
> - It also makes connection details easier to use and fixes misleading
callback health behind HTTPS proxies.

## Linked Issues or Issue Description

**Subsystem affected**
Cross-cutting: chat routes and services, shared contracts, and the Apps
board UI.

**Problem or motivation**
Slack onboarding made users find settings without enough guidance. A
second user needed operator help to link their account. Activity stopped
at 100 records, and TLS termination could mark working callbacks as
stale.

**Proposed solution**
Use six setup steps with editable app names, a generated manifest,
credential guidance, URL verification, account linking, and an optional
message test. Send each Slack user a private, expiring confirmation
link. Require company membership or an approved access request before
linking. Add cursor pagination and tolerate the internal HTTP hop in
callback diagnostics.

**Roadmap alignment**
This improves the existing connected-app surface and supports CEO Chat
without changing the task-and-comments model. The maintainer requested
and reviewed the flow during a live Slack test drive.

**Additional context**
Related work: #7, #3349, #13000, and #13620. Those cover broader chat
capabilities, older webhook paths, or plugins. This PR improves the
existing native connector's setup and account-linking flow. HTTPS
documentation was published separately in
paperclipai/paperclip-docs#128.

## What Changed

- Split Slack onboarding into six clickable sidebar steps. Keep
secondary and primary actions on one row.
- Generate the Slack creation link and read-only manifest from editable
app, bot, and command names. Add credential prefix validation and direct
instructions.
- Add live account-link status and an optional mention-based message
test.
- Add private, single-use Slack account invitations and membership
access requests. Retain cloud authentication/bootstrap checks and
enforce the chat rollout flag in all identity APIs. Default new Slack
connections to linked users only.
- Put Settings, Access, Conversations, and Activity in the sidebar.
Simplify conversation rows and remove active header badges.
- Add 25-item activity pages, stable timestamp/ID cursors, and replay
safety across pages. Preserve the legacy array API for clients without
pagination parameters.
- Fix false callback warnings when HTTPS terminates at a proxy. Keep
host, port, and path drift detection.
- Document the setup flow, pagination, callback diagnostics, and shared
wizard footer rule.

## Verification

- Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed: focused Slack callback and pagination integration tests; UI
clipboard, wizard, pagination, and activity tests; OpenAPI route tests.
The final access-gate fix also passes 27 focused tests covering cloud
authentication/bootstrap, nonmember invitations, token validity, and the
server-enforced rollout flag.
- Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11
provider browser scenarios (including mobile light/dark navigation).
After rebase, the identity route, sidebar, and 25 clipboard tests pass.
- The full local `pnpm test:run` was attempted. The first run found 14
Slack fixtures that needed explicit guest access; those are fixed and
the complete chat suite passes. Unrelated embedded PostgreSQL
startup/resource failures and timeouts prevented a clean full local run.
All CI checks pass on `2d858b036`, including the full chat, server,
workspace, build, typecheck, and browser suites.
- Live test drive: Slack app creation, credential setup, URL
verification, private account confirmation, mention messages, and thread
replies. Verified the callback warning clears for the existing proxied
connection.
- Review: create a Slack connection, follow the six steps, link a second
user's account, and browse older activity with Next and Previous.

## Risks

- Identity invitations carry a temporary capability. Tokens are hashed,
expire after 15 minutes, work once, and require explicit confirmation by
a company member. Access requests do not grant membership.
- New Slack connections reject unlinked people by default. Existing
connection settings remain intact.
- Activity is a live ledger. Updated action rows can move forward in
time. Older pages do not poll.
- Proxy tolerance affects health display only. Slack signature checks
and proxy authentication settings remain unchanged.
- No database migration or package-lock changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose an
exact model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; full
local-run limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-18 17:23:53 -05:00
scotttongandClaude Opus 5 8f69e0af9d fix(ui): confirm every copy action (#13603)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Much of the work means moving an opaque value somewhere else: an
agent id, a callback URL, a webhook secret, an error payload for a bug
report
> - Copy buttons all over the app do that, and most already say whether
it worked — an icon that flips to a check, a label that reads "Copied"
> - Eleven did not. They passed the promise to `.catch(() => {})` and
showed nothing at all
> - A copy that says nothing looks exactly like a copy the browser
refused, and the only way to tell is to paste somewhere and look
> - The clipboard really does refuse: over plain HTTP on a non-secure
host the async Clipboard API is unavailable, and the fallback can still
fail
> - This pull request gives every copy action an affordance, through two
shared hooks
> - The benefit is that a person can tell a copy from a failure without
leaving the page

## Linked Issues or Issue Description

**What happened?**

Eleven copy buttons gave no feedback. They called `copyTextToClipboard`
and discarded the result:

```
onClick={() => void copyTextToClipboard(robotEmail).catch(() => {})}
```

Nothing changes on screen, whether the write succeeded or failed. This
is the same surface where most other copy buttons do show a check or a
"Copied" label, so the silent ones read as broken.

**Expected behavior**

Every copy action says whether the clipboard took the value. A rejected
write says so rather than staying silent or claiming success.

**Steps to reproduce**

1. Open the connector setup flow for an app that shows a sharing email
or a callback URL.
2. Press **Copy** next to that value.
3. Before this change nothing on screen changes. Compare with the copy
button on a login panel, which flips to a check.
4. To see the failure case, open Paperclip over plain HTTP on a
non-localhost host, where the async Clipboard API is not available.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

Commit `1fa36be35` moved every copy through `lib/clipboard.ts`, so the
call sites are findable. Of 62 call sites, 51 already showed feedback
and 11 did not.

## What Changed

- New `ui/src/lib/use-copy-action.ts` with two hooks:
- `useCopyAction` for a control that stays on screen. It returns
`copied` and `failed` for the inline swap the rest of the app already
uses, and resets itself.
- `useCopyToast` for a menu item, which unmounts with its menu before an
inline state could be read, so its confirmation goes to the toast
viewport.
- Both wait for the write to resolve before reporting success, and both
report a rejection as a failure. `AdapterLoginChrome.test.tsx` already
held that line for one button; the helper now makes it the default.
- The eleven silent sites now use one of the two: the connector setup
flow's sharing-email and callback-URL buttons, the runner inspector's
copy-value and copy-path buttons, the agent bubble and skill studio menu
items, annotation "Copy link" on both its hosts, both secret-error
detail buttons, the chat webhook secret, and the motion tweak panel's
export box.
- The chat webhook secret follows its own file's convention: a label
that reads "Webhook secret copied", matching the manifest buttons beside
it.
- The 51 sites that already had feedback are untouched. Converting them
would be a large diff with no visible change.
- `clipboard-usage.test.ts` gains a static check, sibling to the one
that already keeps copies on the shared helper. It fails if a new copy
site ships with no affordance.

## Verification

- `npx vitest run ui/src/lib/use-copy-action.test.tsx
ui/src/lib/clipboard-usage.test.ts ui/src/lib/clipboard.test.ts
ui/src/components/AdapterLoginChrome.test.tsx` — 18 tests pass.
- The new hook tests cover the three cases that matter: no confirmation
while the write is still in flight, a failure state on a rejected write,
and a return to rest after the reset delay. The toast hook is covered
for both tones.
- `npx vitest run ui/src/pages/apps/AppsConnect.test.tsx
ui/src/components/task-chat ui/src/components/RunnerInspector.test.tsx
ui/src/components/DocumentAnnotation` — the suites over the touched
components pass.
- `npx tsc -b ui` — clean.

## Risks

Low. Each change is additive at its own call site and the copy path
itself is unchanged. The static check is the only part that touches
future contributors: it is one assertion, and its pattern list is easy
to extend or drop.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 14:32:19 -07:00
scotttongandClaude Opus 5 34355e5109 fix(task-chat): selecting an option no longer jumps to the next question (#13602)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent that needs a decision asks in the task composer, which
renders a question set one question per page
> - Single-choice questions use radio options, and the composer also
ships a "Next" button and pagination arrows
> - Selecting a radio option also set a pending advance, waited for a
short confirm animation, and turned the page on its own
> - That takes the page away from the reader while they are still
reading it, and a misclick costs them the question
> - Nothing on the option says that clicking it will navigate
> - This pull request makes selection answer the question and nothing
else
> - The benefit is that moving between questions is always something the
reader chose to do

## Linked Issues or Issue Description

**What happened?**

In the task composer, selecting an option with a radio button jumps to
the next question. `QuestionForm.toggleOption()` set `pendingAdvance`
for any single-select question that was not on the last page, and an
effect then read `--motion-question-confirm`, waited that long, and
called `setPage(page + 1)`. With reduced motion it advanced immediately,
with no pause at all.

The result is that a click meant to answer a question also navigates
away from it. There is no way to read the rest of the page after
choosing, and no way to undo the jump other than pressing the back
arrow.

**Expected behavior**

Selecting an option records the answer and stays on the question. Moving
to the next question stays an explicit act.

**Steps to reproduce**

1. Get an agent to ask a question set with two or more single-choice
questions.
2. Open the question in the task composer.
3. Click one of the radio options.
4. Before this change the composer moves to the next question on its
own.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

Every route forward already shipped beside the options, so removing the
shortcut traps no one:

- a footer button that reads "Next" on any page but the last, and the
submit label on the last;
- previous and next pagination arrows;
- "Skip" on questions that are not required.

The same question sets render outside the composer in
`IssueThreadInteractionCard`, which never had the auto-advance. This
change brings the two surfaces to the same behavior.

## What Changed

- `QuestionForm` no longer sets a pending advance when a single-choice
option is selected. The effect that turned the page is removed with it.
- Selecting an option still clears a submit error about a missing
answer, which is the case that error is about.
- The confirm-before-advance animation existed only to soften the jump.
Its `--motion-question-confirm` token, its keyframes, its class, and the
now-unused `confirming` prop on the option button are removed.
- Tests that asserted the jump now assert the opposite: selection leaves
the page number unchanged, "Next" advances, and number-key selection
stays on the same question.

## Verification

- `npx vitest run ui/src/components/task-chat` — the composer,
interaction card, protocol card, and motion token suites pass.
- `npx vitest run ui/src/components/task-chat/TaskChatComposer.test.tsx`
— 87 tests pass, including four new tests for the stationary behavior.
- `npx tsc -b ui` — clean.
- Keyboard: options keep `role="radio"` inside a `radiogroup`, and a
test asserts that number-key selection selects without navigating.

## Risks

Low. One deliberate behavior is removed. It costs a reader one extra
click per question on multi-question sets, and the explicit control for
that click already existed and is already tested.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 14:18:39 -07:00
scotttongandClaude Opus 5 378f11994b fix(connections): stop the sign-in screen from adding a phantom step (#13601)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents reach other products through connections, which people create
in the connector setup wizard
> - That wizard shows a stepper: a row of dots, a label under each, and
"Step 1 of 2" in the header
> - When a curated OAuth app hands off to the provider, a waiting screen
replaces the wizard while Paperclip prepares sign-in
> - That waiting screen drew its own stepper with three hardcoded
labels, so the stepper grew a dot at the exact moment the person pressed
Connect
> - A stepper that grows mid-flow tells the person they missed a screen
> - This pull request makes the waiting screen show the step model of
the flow that opened it
> - The benefit is that the stepper keeps the same shape from the first
screen to the handoff

## Linked Issues or Issue Description

**What happened?**

The connector setup wizard shows a two-step stepper for a curated OAuth
app: two dots, labelled "Access" and "Sign in", with "Step 1 of 2" and
then "Step 2 of 2" in the header. When you press Connect, a third dot
appears, labelled "Ready".

The extra dot comes from the OAuth waiting screen.
`OAuthConnectStateScreen` rendered its own header with
`labels={["Access", "Sign in", "Ready"]}`, a constant that did not
depend on the wizard that opened it.

"Ready" is also not a step the wizard can reach. The success screen
hides the step header, so no one ever sees a third step become active.

**Expected behavior**

The stepper keeps the same number of steps from the first screen through
the provider handoff. Waiting for browser sign-in is part of the last
step, not a step of its own.

**Steps to reproduce**

1. Open **Apps → Connect** and pick a curated app that signs in through
OAuth, for example Notion.
2. Look at the stepper. It shows two dots and reads "Step 1 of 2".
3. Complete the Access step and press the button that starts sign-in.
4. Look at the stepper on the "Preparing secure sign-in" screen. Before
this change it shows three dots and adds a "Ready" label.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

The reporter suggested removing the stepper, on the condition that no
connector flow has three or more steps. One does.
`ConnectionSetupFlow.tsx` keeps `STEP_LABELS = ["Pick app", "Access",
"Add your key"]` for the generic path, which a person walks when they
paste an MCP endpoint instead of choosing a curated app. The two-step
counts apply only after an app is selected. The condition fails, so this
pull request keeps the stepper and repairs the phantom step alone.

## What Changed

- `OAuthConnectStateScreen` takes a `steps` prop: the labels and active
index of the flow that opened it.
- The curated OAuth path passes its own two-step model, so the count
does not change at the handoff.
- The generic pasted-endpoint path passes the three-step model it was
already walking, so its count does not change either.
- The default for hosts without a wizard of their own, such as the
paste-a-config tab, is `["Access", "Sign in"]`. The invented "Ready"
step is gone.
- `AppsConnect.test.tsx` gains a test that reads the stepper before and
after the handoff and fails if the two differ.

## Verification

- `npx vitest run ui/src/pages/apps/AppsConnect.test.tsx
ui/src/features/connections ui/src/pages/tools/PasteConfigTab.test.tsx`
— 185 tests pass.
- `npx vitest run ui/src/pages/apps/generic-mcp-connect.test.ts` — 27
tests pass.
- `npx tsc -b ui` — clean.
- The new test fails on the previous code. With the old three-label
constant restored it reports `["Access", "Sign in", "Ready"]` where it
expects two labels.

## Risks

Low. The change is limited to which labels the waiting screen draws. No
connection logic, no network call, and no navigation changes.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 14:11:35 -07:00
Devin FoleyandPaperclip 45586170e1 fix(ui): reveal task history while retries wait to start (#13597)
## Thinking Path

> - Paperclip lets people manage AI agents and inspect their work.
> - Task conversations combine comments with run transcripts.
> - The first reveal waits for the relevant history to load.
> - Scheduled retries have not started, so the log reader does not
hydrate them.
> - Waiting for those retries keeps the whole conversation hidden after
its header loads.
> - This change reveals available history while a retry waits to start.

## Linked Issues or Issue Description

**What happened?**

A task header loads, but the conversation stays behind the loading
overlay while a linked run has `scheduled_retry` status. The readiness
check expects that run in `hydratedRunIds`, although the log reader does
not read its log. A retry delay can therefore become a task-loading
delay.

**Expected behavior**

Show existing comments and transcripts once their data is ready. A retry
that has not started must not block the conversation.

**Steps to reproduce**

1. Open a task with existing comments and a linked run waiting in
`scheduled_retry`.
2. Let the comments and other run transcripts finish loading.
3. Observe that the header loads but the conversation stays hidden until
the retry changes status.

**Paperclip version or commit**

Reproduced on master at `3f1d897a7` with regression tests.

**Deployment mode**

Web UI, built from source. This is a client readiness bug.

Searched public issues and PRs for scheduled retries, conversation
loading, and history loading. No duplicate found. Related: #10255
changes the server response for active runs with no log; it does not
address this scheduled-retry readiness gate. This focused bug fix does
not add a roadmap feature.

## What Changed

- Exclude `scheduled_retry` runs from the initial task-history readiness
gate.
- Extend transcript mocks to expose per-run hydration state.
- Add six regression cases across legacy and native runs. Existing
history loads during retry waits, while unhydrated running and completed
runs still block the first reveal.
- Document the exception in a code comment. No user command or
configuration changes need documentation.

## Verification

- All six new cases fail before the production fix.
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/TaskChatThread.test.tsx
src/components/transcript/useLiveRunTranscripts.test.tsx
src/components/transcript/useNativeRunTranscripts.test.tsx`: 164 passed.
- `pnpm --filter @paperclipai/ui typecheck`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm check:token-gates` and `git diff --check`: passed.
- `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the
native runner package because the local machine has no Rust `cargo`
executable. UI checks pass separately.
- `pnpm test:run`: started locally, then stopped before completion after
full CI passed on the same commit. The local full-suite result is
incomplete.
- CI on `46daa8b06`: all 53 checks passed; two optional Storybook jobs
were skipped. This includes the full test matrix, build, typecheck,
native runner checks, browser tests, and canary dry run.
- Greptile: 5/5 on `46daa8b06`, with no review threads or actionable
findings. The branch has no merge conflicts with master.

## Risks

Low risk. The exception applies only to runs waiting in
`scheduled_retry`. Running and completed runs retain their existing
readiness checks. Retry scheduling, transcript fetching, and server
behavior stay the same. No migration is needed.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The exact backend revision and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 22:07:27 -07:00
scotttong 3f1d897a7c copy(connectors): say "organization" in the connector setup flow (#13589) 2026-09-17 19:37:14 -07:00
DottaandPaperclip 84fe89906d fix: complete native agent review handoffs (#13581)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Native execution uses durable runs, issue locks, wake requests, and
typed tool authority
> - A child can finish with a native agent review request while its
original assignee stays responsible for the work
> - The reviewer then needs a bounded execution path that can inspect
the child, record one decision, and finish safely
> - Before this change, assignee-only gates rejected the reviewer or
left the parent waiting after the child review ended
> - This pull request adds typed reviewer admission, scoped reviewer
tools, durable wake and recovery handling, and parent continuation
evidence
> - The benefit is that native review handoffs complete without changing
child ownership or granting broad mutation access

## Linked Issues or Issue Description

Refs: #13314
Refs: #13574

**What happened?**

A native child run could report `needs_review` for an agent reviewer.
The reviewer wake then failed assignee and execution-lock checks. The
child remained in review and the parent remained waiting.

**Expected behavior**

The named reviewer should receive one durable wake. The reviewer should
inspect the child and resolve the exact review card. The child assignee
should stay unchanged. The parent should receive the recorded review
outcome after the child reaches its terminal state.

**Steps to reproduce**

1. Run a native task with a different named agent reviewer.
2. Keep the child assigned to its original worker.
3. Let the worker finish with a native completion review request.
4. Start the durable reviewer wake.
5. Resolve the review and finish the reviewer run.
6. Observe the child and parent state.

**Paperclip version or commit**

Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification
source: `eea171aae`.

**Deployment mode**

Built from source.

**Installation method**

Built from source (pnpm build).

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Access context**

Both.

**Database mode**

Embedded PostgreSQL in the isolated live test fixtures.

## What Changed

- Add server-validated native review assignment facts.
- Admit only the exact company, issue, source run, decision, revision,
addressee, and resolver policy.
- Give reviewer runs a narrow set of Paperclip read and resolve tools.
File and shell access follow the configured agent and environment
policy, so reviewers can run tests.
- Separate server-owned reviewer instructions from untrusted persisted
review data. Escape the data boundary; retain server-enforced
authorization.
- Keep the child assignee unchanged. Atomically claim the reviewer run,
wake request, and issue execution lock. A competing lock prevents
provider startup.
- Require the exact running reviewer session and current issue lock to
resolve its assigned card. Reject missing, unrelated, or terminal
reviewer runs.
- Add durable reviewer wake, lock, stale-card, and abandoned-run
recovery handling.
- Prevent duplicate native wake dispatches during deferred admission and
recovery.
- Carry accepted or rejected child review outcomes into parent task
context and continuation evidence.
- Add focused server, runner, and native protocol coverage.
- Preserve upstream continuation rules. Add child review decisions as
separate evidence, while keeping real human answers in their own field.
- Return actionable completion validation feedback to both providers.
Permit a corrected completion after rejection. Keep strict terminal
acknowledgment validation.
- Apply exclusive shared-workspace locks to sandbox environments. Local
and SSH folders can run concurrently, including when old settings
request serialization.
- Repair test timing, native event parsing, and the review artifact
assertion. Allow a valid reject, correct, and accept review sequence.
Check the accepted card against its reviewer run and decision. Keep
polling within the existing deadline when review acceptance precedes the
parent wake projection; report a specific missing-continuation error at
timeout.
- Apply the ACPX pending-call limit to reserved finish/block calls, with
capacity-release and cancellation tests.

## Verification

- `pnpm build`: passed on `eea171aae`.
- `pnpm -r typecheck`: passed on `eea171aae`.
- `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on
`b31ad9ab8`; runner E2E typecheck also passed.
- `pnpm check:token-gates`: passed.
- Focused DB review, reviewer authority, and prompt-boundary checks: 31
tests passed. They cover invalid reviewer runs, competing locks, atomic
admission, duplicate claims, and valid resolution.
- Heartbeat, workspace, and recovery checks: 30 tests passed.
- ACPX sidecar suite: 27 tests passed. Moving the capacity guard back
below reserved handling makes both new regression cases fail.
- Four focused live continuation checks passed on their first attempt at
`f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and
question tool guidance (12/12 each). These cases do not use the reviewer
prompt path changed afterward.
- Fresh Codex and Claude review-handoff checks passed all 29 native
checks each on their first attempt at `eea171aae`. Both runs received
the expected fixed prompt and completed cleanup. Only the six selected
live flows were tested; no full paid provider catalog run.
- The final commit only extracts the existing test-harness timeout
diagnostic into a shared helper and adds positive and negative coverage.
Removing the accepted-review guard makes two regression assertions fail;
restoring it passes all six timeout tests. Production runtime code,
prompts, deadlines, and grading criteria are unchanged by this final
commit.
- Deadline regressions: a valid continuation delayed 20 seconds succeeds
within its 30-second unit-test deadline; an absent wake returns a
specific candidate-failure diagnostic at that same deadline. Both
assertions failed before the fix. Production E2E deadlines remain
unchanged.
- Historical native failures remain recorded: Docker availability
failures; a valid reject/correct/accept sequence that the first-card
grader misread; and a test that rejected the gap between accepted child
review and parent wake projection. No failed result was regraded. The
latest tests use a protected reference to the pinned Docker image and
the unchanged artifact oracle and time limits.
- Full repository verification runs in GitHub CI. Local verification
uses the focused suites above, full build, and full typecheck. An
unchanged Codex shutdown timing test failed once in CI, passed in
isolation, and its full shard passed on the final commit without changes
to that test or its causal code path. The original failure is retained
in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no
outstanding actionable findings. All review threads are resolved. All
current-head CI gates passed, including the isolated native runner
Docker build (55 successful checks; two skipped by the workflow).

## Risks

- Reviewer admission depends on exact persisted decision and interaction
bindings. A stale or changed card is rejected.
- Paperclip control-plane tools are limited to inspection and review
resolution. This is not a filesystem permission boundary; provider file
and shell access retain the configured policy.
- Deferred wake recovery changes dispatch receipt coalescing. A
scheduler regression could delay a continuation if the receipt state is
wrong.
- Parent review outcomes are evidence for the model. They do not grant
tool authority or change issue ownership.
- This change does not address legacy lease-hold handoff behavior.

> Roadmap review: native execution, review gates, and durable recovery
are existing roadmap capabilities. This PR completes a narrow
reliability path for those capabilities.

## Model Used

OpenAI `gpt-6-astra` with reasoning, tool use, and code execution.
OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and
journal work. Context window size is not exposed by this session. Live
test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the
PR authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 15:52:19 -05:00
mouse-value-add fcdb3f2499 feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents that do research work need live information from the web
> - Paperclip reaches external systems through governed, catalog-based
MCP connections
> - The Apps catalog is data-driven: a researched provider with a hosted
remote MCP server becomes a connectable app with no runtime code change
> - You.com operates a hosted remote MCP server for web search, content
extraction, and research tools
> - The server supports OAuth 2.1 with dynamic client registration, an
API key in a bearer header, and a keyless free profile at a separate
endpoint
> - This pull request adds You.com to the self-serve MCP research ledger
and generates its catalog entry with three connection methods: browser
sign-in, API key, and the keyless free profile
> - The benefit is that an operator can give agents live web search
through the normal connection governance, and the free profile needs no
account at all

## Linked Issues or Issue Description

No public issue exists for this provider. The problem description
follows the new-adapter issue template.

**Agent or provider**

You.com — web search and research tools over a hosted remote MCP server.

**Why this adapter is useful**

Agents that do research, monitoring, or fact-finding tasks need current
web results. You.com exposes web search (`you-search`), live page
extraction (`you-contents`), citation-backed research (`you-research`),
and finance research (`you-finance`) as MCP tools. Any Paperclip company
can connect it in a few clicks. The free profile offers `you-search`
without an account, so a new company can try agent web search at zero
cost and zero setup.

**How the agent is invoked**

Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`.
Three supported access paths, verified against the live server on
2026-09-16:

- OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate`
challenge with RFC 9728 protected-resource metadata and advertises a
dynamic client registration endpoint, so Paperclip's automatic DCR path
applies.
- API key. Sent as an `Authorization: Bearer` header per the provider's
official server manifest and docs. Keys come from you.com/platform and
unlock higher rate limits plus the full tool set.
- Keyless free profile at `https://api.you.com/mcp?profile=free`.
Provides a reduced, read-only tool set.

Official docs: https://you.com/docs/build-with-agents/mcp-server

**Are you willing to implement it?**

Yes. Implemented in this pull request.

**Additional context**

Research evidence collected 2026-09-16, from live protocol probes and
official provider sources only:

- Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with
`WWW-Authenticate: Bearer
resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource"
scope="Tools offline_access"`.
- RFC 9728 metadata lists one authorization server with scopes `Tools`
and `offline_access`.
- The authorization-server metadata (RFC 8414) publishes authorization,
token, and revocation endpoints, and advertises a
`registration_endpoint`, so DCR is available. No registration was
performed during research, per the runbook's non-registering preflight
rule.
- The keyless free profile answers `initialize` (server `You.com`,
version `4.0.1`), lists the tools `you-search` and `you-discover`, and
executed both tools successfully during the probe.
- The API-key placement matches the provider's official `server.json` in
the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer
<key>`.

## What Changed

- Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve
MCP research ledger in
`packages/shared/src/self-serve-mcp-research.json`, and refreshed the
ledger verification date.
- Added the You.com category (`ai`) and API-key header spec to
`scripts/ingest-app-definitions.mjs`.
- Added a You.com case to `specialMethodsFor` that emits three methods:
browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer
header), and keyless free profile (`mcp-free`, no auth).
- Regenerated `packages/shared/src/app-definitions/youcom.json` and the
generated registry via the ingestion script (`--definitions-only` mode;
no unrelated provider churn).
- Added the official You.com wordmark artwork (light and dark theme
variants, taken from the provider's docs site) under
`ui/public/brands/apps/`, with a manifest entry.
- Updated `packages/shared/src/app-definitions.test.ts`: ledger counts
(47 providers, 44 candidates), store count (48), verification date, and
assertions for the three You.com methods and their endpoints.

## Verification

- `node scripts/ingest-app-definitions.mjs --definitions-only` — passed.
Generated the new definition and registry import only; no other provider
JSON changed.
- `node scripts/check-app-brand-assets.mjs` — passed (71 identities).
- `node --test scripts/app-brand-validation.test.mjs` — passed.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests).
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-connection-removal.test.ts
ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server
suites that require embedded Postgres skipped on this machine by their
own environment gate; the gate is unrelated to this change.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed;
builds `@paperclipai/shared` with the new definition.
- `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592
skipped. Every failure is environmental on this container: the
embedded-Postgres suites refuse to start because the machine runs as
root, the native runtime suites need the Rust runner binary that this
container cannot build, and one media suite needs a native HEIC binary.
No failure touches the app-catalog, connection, branding, or
shared-package surface; those suites pass locally. CI is the
authoritative gate for the full suite.
- `pnpm --filter @paperclipai/server typecheck` — not completed: the
script's `prepare:runner-vendor` prelude builds the Rust runner, which
cannot build on this container. A direct `tsc --noEmit` reports only
pre-existing errors from the missing vendored runner types; no error
touches this change. No server code is changed.
- Live You.com proof on 2026-09-16 (keyless free profile, real network
calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR
endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned
results ✓, `you-discover` call returned results ✓.
- Live proof NOT run: an authenticated OAuth connect and an API-key call
against the full server. This environment has no You.com account or API
key. Per the runbook, this proof stays outstanding and must not be
assumed from the keyless probe. Both paths match the reviewed
`mcp_remote` patterns (DCR and bearer header) used by existing
providers.
- Browser e2e suites not run: opt-in per `AGENTS.md`, and this change
adds catalog data only, with no UI code.

## Risks

- Low risk. The change is catalog data plus generated output. It adds no
runtime code and touches no existing provider.
- The free-profile method is a fixed keyless endpoint. If You.com
changes or removes `?profile=free`, that method breaks and the entry
needs a ledger update. The OAuth and API-key methods do not depend on
it.
- The authenticated tool catalog is discovered live at connect time, so
provider-side tool changes appear through the normal catalog refresh and
quarantine flow, not through this definition.
- Rollback is a single revert; no migration and no state are involved.

## Model Used

- Provider: Zhipu AI, via OpenRouter
- Model: GLM-5.3 (`z-ai/glm-5.3`)
- Context window: 200K tokens
- Capabilities used: tool use (shell, file edits, live HTTP probes),
long-context repository reading
- The change was produced with AI assistance and reviewed by a human
before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-17 10:01:49 -07:00
5d9b20ccf0 fix(ui): keep task composer available while pause state loads (#13562)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task messages enter through the shared task composer.
> - The task page waits for a separate tree-control query to find active
pause holds.
> - The old loading guard disables the composer until that request
settles. A slow or stalled request blocks messages on both desktop and
mobile.
> - This PR allows sends while that query is pending. The server still
rejects paused board messages before saving a comment or waking an
agent.
> - Query failures and known pause holds still block the composer.
> - Regression tests cover pending sends, successful responses, query
failures, and late root or inherited pauses.

## Linked Issues or Issue Description

Fixes: #13561

Related: #13569 adds pending/error coverage for the same fix. This PR
now covers those cases and the late-pause transitions.

The report overstates two details: `isPending` clears after a successful
response, and the shared loading guard also affects mobile. The
reproduction holds the request pending. It does not prove why the
reporter's request remained unresolved.

## What Changed

- Preserve the original two-line fix that removes the pending-state
composer block.
- Add eight page regression cases using a real QueryClient and a
controlled API promise. Cover desktop and mobile, submission before the
query completes, successful resolution, rejected resolution, and late
root/inherited pause responses.
- Repair an existing server CI failure in a separate commit. Reuse the
shared unique-violation helper so Drizzle-wrapped duplicate inserts
become retryable document conflicts. Add a deterministic regression and
preserve unrelated database errors.
- Repair an existing Inbox test race in a separate commit. Wait for
workspace metadata, which resolves independently of the task list.

## Verification

- Red: restore the pre-fix `IssueDetail.tsx` and run the new `composer
tree control` cases. Both desktop and mobile pending-send cases fail
with `Checking task status…`. The other six cases pass.
- Green: restore the original PR fix. All 343 tests in IssueDetail,
TaskChatThread, and TaskChatComposer pass.
- The server comment/reopen and artifact-review suites pass (179 tests).
They include POST and PATCH pause checks that return 409 before any
comment, task mutation, or wakeup.
- The document error regression fails before the shared-helper fix and
passes after it. The document, artifact-review, and database-error
suites pass (31 tests). The handler matches only the issue-document key
constraint; revision and other constraint errors retain their original
identity.
- The Inbox suite passes (27 tests).
- `pnpm check:token-gates` passes.
- `pnpm -r typecheck` and `pnpm build` pass locally. The full CI matrix
passes on `5af1ed0b43269247aaba406cf4fd4d7fe1a22e75`: 54 successful
checks and two opt-in Storybook skips, including all general/serialized
tests, runner checks, release checks, and eight browser shards.
- One server shard initially hit an unrelated `EADDRINUSE` on test port
52000. Its single rerun passed without code changes.
- Greptile completed successfully on that exact commit with 5/5 and no
outstanding findings or review threads.
- The serial local `pnpm test:run` was started, then stopped after the
complete parallel CI matrix passed. It is not claimed as a completed
local full-suite run. The focused local suites above did finish
successfully.

For a manual reproduction, delay the task's `/tree-control-state`
response, open the task, and enter a message. Send should remain
available during the delay. Resolve the response with an active pause
hold and confirm that the pause takeover replaces the composer. Reject
the request and confirm that the error blocks sends.

## Risks

A user can attempt a send before the pause response arrives. The server
remains authoritative and returns 409 for a paused task before saving or
waking work. The known pause takeover and query-error block remain.
There are no schema, API, or styling changes.

The document change restores the existing conflict/retry behavior for
wrapped database errors. It does not retry unrelated database failures.
The Inbox change affects test synchronization only.

## Model Used

- Original fix: Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k
context, tool use and code editing, as reported by the author.
- Review, regression tests, and CI repairs: OpenAI GPT-6 (`gpt-6-astra`)
through Codex, with reasoning, tool use, and code execution. The session
does not expose its context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: austinpilz <austinpilz@users.noreply.github.com>
Co-authored-by: Dotta <bippadotta@protonmail.com>
2026-09-17 08:40:40 -05:00
Devin FoleyandPaperclip fae6980310 revert(apps): restore Google connector visibility (#13552)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - PR #13551 temporarily hid Google connectors.
> - We now want to restore their catalog visibility.
> - This PR reverts that change and restores the previous catalog
behavior.

## Linked Issues or Issue Description

Refs: #13551

Revert the temporary removal of Google connectors from the UI.

## What Changed

- Restore Gmail and eight Google Workspace entries to the catalog.
- Restore the matching branding flags and original catalog and service
tests.
- Remove the temporary-hiding documentation note.

This is an exact revert of commit
`cf1e873ab24277d55ffd3ab06074f77014dc4015`.

## Verification

- Passed: 507 catalog, UI, and connection service tests.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Full local Vitest was not repeated because the unchanged base has
confirmed macOS skill-cache permission failures. The full CI suites
passed.
- Passed: all GitHub CI gates; Greptile 5/5 on commit
`4e3dddef0ebfef1f99001e7735822ed4cba852ab`, with no review threads.
- Reviewer check: open Connectors and confirm that Gmail and Google
Workspace entries appear again.

## Risks

Low risk. This restores the previous catalog visibility and setup entry
points. Connector implementations and saved connection data are
retained.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:35:16 -07:00
Devin FoleyandPaperclip cf1e873ab2 fix(apps): temporarily hide Google connectors (#13551)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - We need to temporarily remove Google connectors from the UI.
> - The catalog already separates visibility from retained definitions.
> - This PR uses that setting so Google can return with a small change.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The Connectors catalog and its setup entry points.

**Current behavior**
The catalog shows Gmail and eight Google Workspace connectors.

**Proposed behavior**
Temporarily hide those nine entries. Keep their definitions and existing
connections.

**Reason and benefit**
Make the temporary UI removal easy to reverse.

**Breaking changes**
Fresh catalog setup no longer offers Google. Saved connections keep the
existing management and reconnect paths.

## What Changed

- Add the nine Google connector slugs to the existing hidden list.
- Match the branding manifest visibility flags.
- Update existing catalog and service tests. Keep backend Google
connection coverage and document how to restore visibility.

## Verification

- Passed: 507 targeted tests covering catalog definitions, URL matching,
setup routing, connector UI, branding, and the connection service.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Stopped the full local Vitest run after skill-cache permission
failures. Three failures in `company-skills-service.test.ts` also
reproduce on the unchanged base branch. The final connector service
suite passes all 319 tests.
- Greptile: 5/5 on the current commit, with no open review threads. CI
is retrying one unrelated preview-server readiness timeout. That test
file passes all seven tests locally.
- Reviewer check: open Connectors in a company with no Google
connections. Gmail and Google Workspace entries should be absent.
Existing saved connections remain manageable.

## Risks

Low risk. This uses the existing catalog visibility mechanism. No
connector implementation, credential, or database schema is removed.
Restoring visibility requires updating both the hidden list and branding
manifest.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:04:46 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:44:36 -05:00
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00
DottaandPaperclip 669bd0e7a9 fix(ui): show progress and verify subscriptions during onboarding (#13499)
Automatically verify detected Claude and Codex subscriptions during initial onboarding. Show connection progress, support retries, and ignore stale results after navigation. Prefer the personal default subscription while preserving the later create-agent chooser.

Allow Enter to advance from the agent name field. Add regression tests and production-component Storybook coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-15 17:39:25 -05:00
DottaandPaperclip 9cfa7fd2d1 fix(ui): share compact live and saved activity across runners (#13421)
## Thinking Path

> - Paperclip helps people manage AI agents and inspect their work.
> - Task threads show live activity and saved transcripts from several
runners.
> - Legacy runs used a separate activity renderer that expanded tool
cards as commands arrived.
> - A compact live row makes ongoing work easier to follow.
> - This PR shares the native runner activity group across live and
saved legacy turns.
> - Both paths now use the same labels, icon gutter, animation, and
history controls.

## Linked Issues or Issue Description

Related: Refs #13255, which introduced rolling native-runner activity
groups.

**What happened?**

During a legacy CLI run, each new command added an expanded tool card
under an activity heading. The original legacy parity story rendered a
completed turn, so it did not exercise this live path.

**Expected behavior**

Show one current activity line per commentary group. Roll that line
forward when a new activity starts. Keep tool icons aligned on the left,
use friendly labels, and show history only when expanded.

**Steps to reproduce**

1. Start a task with the Codex local adapter using the CLI engine.
2. Ask it to read two files in separate tool calls and run a test
command.
3. Watch the activity feed while the run is active, then expand its
history.

**Paperclip version or commit**

The live behavior was reproduced on
bdee5ebb21b2d09280e49c88bc329435da1b07c8. This update includes both live
and saved rendering fixes.

**Deployment mode**

Built from source with the local test-drive command and a real Codex CLI
agent.

## What Changed

- Use `TaskChatRunnerActivityGroup` for both saved legacy phases and
`TaskChatLiveTail`.
- Align the legacy Working spinner with the shared activity icon gutter.
- Cover streamed reasoning updates, successive commands, image labels,
hidden details, and explicit history expansion in tests.
- Feed raw legacy transcript events through the real adapter and live
renderer in Storybook. Add live, expanded, narrow, light, and completed
stories with the final reply preserved.

## Verification

- Passed 195 targeted tests covering the live tail, status pill, shared
activity group, native turn, and full task thread.
- Passed UI typechecking, token gates, UI build, and Storybook build
before submission.
- Observed a real Codex CLI run while it read separate files and ran
tests. The current activity stayed at one 32-pixel row without an
accumulated tool list; the spinner and activity icon centers aligned.
- Compared native and legacy Storybooks: identical rolling animation,
fixed height, persistent expansion, truncated long labels, and correct
light and completed states.
- Full repository `pnpm -r typecheck` and `pnpm build` passed after
replaying the branch on current master, including the Rust runner. The
full `pnpm test:run` suite is still running.

## Risks

- Legacy activity now starts collapsed. Users can expand each group and
each row to inspect the same transcript details.
- Live runtime request cards must retain their timeline positions.
Existing task-thread and native-runner regression tests cover this
boundary.
- No API, database, or transcript format changes.

## Model Used

- OpenAI GPT-6 through Codex for the live-path fix, tests, and browser
verification. Capabilities used: reasoning, tool use, local code
execution, and image inspection. The exact deployment ID and
context-window size are not exposed in this session.
- The original saved-turn change recorded OpenAI Codex with GPT-5.6; its
exact variant and context size were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 16:40:32 -05:00
DottaandPaperclip d49f168381 fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must publish generated files so users can inspect their
results after a sandbox stops.
> - Legacy sandbox bridges blocked attachment listing and could not
carry multipart binary uploads through the queue transport.
> - The native runner has a separate verified file registration path
that needs the same durable result.
> - This pull request repairs legacy binary transport, makes native
download receipts explicit, and reveals new outputs in the task
Artifacts tab.
> - Users can open generated files from either runner without a
transport flag change.

## Linked Issues or Issue Description

Refs #13355 for the existing native file publication path. Related
filename fixes: #2615 and #4788. Related sandbox persistence work:
#13376. This change repairs attachment delivery through the existing
API; it does not add workspace persistence.

**What happened?**

The upload helper first lists task attachments to avoid duplicates. Both
legacy bridge allowlists rejected that GET request with 403. A direct
multipart upload also failed: the queue bridge accepted only JSON,
excluded attachment uploads, and converted bytes to UTF-8 text. Enabling
HTTP/2 alone did not fix the missing listing route. These failures
occurred before attachment storage.

**Expected behavior**

Both runners can publish a workspace file, register its work product,
bind it to a response, and return a working download. The file stays
accessible after sandbox deletion. A new output opens the task Artifacts
tab. The agent receives accurate errors and decides how to retry or
report a failure.

**Steps to reproduce**

1. Run a legacy agent in Daytona with the duplex bridge disabled.
2. Invoke the bundled upload helper with Bash on a PNG or PDF.
3. Repeat with the duplex bridge enabled.
4. Register the same file through the native runner with generic API
tools disabled.
5. Retry registration, delete the sandbox, and compare the downloaded
bytes with the original file.

**Paperclip version or commit**

The failing baseline was `f2c5e54dc`. This branch is rebased onto
`6cfe4acff`.

**Deployment mode**

Source checkout with a local API and real isolated Daytona sandboxes.

## What Changed

- Allow authenticated attachment listing, upload, and content download
through both legacy bridge transports.
- Add optional base64 body encoding to queue envelopes. Preserve the
existing UTF-8 contract when the encoding field is absent. Decode binary
bodies before forwarding them.
- Preserve multipart headers. Bound raw bytes, encoded envelopes, and
in-flight reservations. Retain timeout and uncertain-write behavior.
- Preserve helper deduplication and return structured uncertain-write
failures. Document explicit Bash invocation in live skills.
- Add attachment IDs and content/download paths to native registration
receipts. Reuse verified local and remote file reads, attachment
storage, work-product registration, and response binding.
- Preserve Unicode upload filenames and provide a valid
Content-Disposition header.
- Open the task Artifacts tab when new stored outputs arrive, including
a closed desktop panel or mobile drawer. Deduplicate upload and
registration events by object ID. Preserve manual selection on
refetches, edits, and panel remounts.
- Remove task artifact filters, the company Artifacts footer link, and
the unassigned group heading and timestamp.

### Screenshot

![Generated images and a document in the task Artifacts
tab](https://pages.paperclip.ing/sandbox-file-delivery-2026-09-15/artifacts-tab.png)

This is the local display fixture. The image was generated separately
and published through the attachment and work-product APIs.

## Verification

- Post-rebase `pnpm -r typecheck` and `pnpm build` pass.
- The post-rebase local `pnpm test:run` passed 12,369 tests before one
existing conversation reset test timed out; all 33 tests in that suite
pass when rerun with isolated test configuration. The aggregate command
stopped before its remaining groups. GitHub runs the complete suite in
separate shards.
- All [GitHub verification
checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350)
pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks
and two configured skips. The native-session recovery assertion
initially raced its fire-and-forget Sentry report; all 13 tests pass
locally, and the same-commit CI rerun passes all 170 suites (3,079
tests).
- Live post-rebase Daytona: all three file-delivery tests pass. They
cover the real Bash helper with the queue bridge, the helper with
HTTP/2, and native `register_deliverable` with generic API tools
disabled.
- Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate
registration, response binding, authorization controls, and
byte-for-byte downloads after sandbox deletion.
- Local focused coverage includes transfer bounds, malformed encoding,
interrupted transfers, remote path containment, and native file
verification. The attachment route suite passes all 32 tests, including
an eight-case filename-header matrix for Unicode and special characters,
inline and forced downloads, and full and partial responses.
- Browser verification confirms image previews, persisted downloads,
automatic Artifacts selection, and preserved manual selection after
edits and reloads. Desktop/mobile component coverage passes. The latest
UI cleanup passes its 10 affected tests and token gates.
- Coverage limit: the Daytona tests call the real helper and native
registration path directly. They do not replay a complete model-led
image-generation task through the browser.

Live command (requires a configured Daytona credential):

```sh
PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts
```

## Risks

- Binary queue bodies use more memory because base64 adds encoding
overhead. Transfer and process limits must remain aligned.
- An interrupted write can have an unknown result. The bridge reports
this state and preserves stable retry identities.
- New artifacts intentionally change the active task tab. Existing
history and repeated updates must not take focus again.
- Transport flag defaults, server authorization, frozen skill snapshots,
and completion policies remain unchanged. No schema migration is
required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser testing. The runtime does not expose a more
specific model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 15:46:57 -05:00
DottaandPaperclip f4cdc7b231 fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The control plane prepares task workspaces before it starts an
agent.
> - Workspace preparation reads Git state so it can preserve edits and
exclude private files.
> - A failed scan was treated as a non-Git folder and lost its actual
failure code.
> - The resulting generic setup failure could not recover, even when the
cause was temporary.
> - This pull request keeps the cause and uses the existing bounded
retry schedule before provider startup.
> - Tasks can recover without human intervention, while permanent
failures and exhausted retries stop with useful guidance.

## Linked Issues or Issue Description

**What happened?**

A Git scan error during managed repository preparation became
`Configured repository folder is not a Git checkout`, followed by
generic `setup_failed`. The agent never started. Generic recovery could
not distinguish a temporary timeout from a bad workspace configuration.

**Expected behavior**

Keep the closed scan error code. Retry temporary timeouts and queue
saturation under the existing shared budget. Preserve edits, exclusions,
ownership, and pause gates. Stop permanent failures and exhausted
retries with a specific explanation. Do not replay historical generic
setup failures.

**Steps to reproduce**

1. Configure a task project with a local Git source that must be copied
into its managed repositories.
2. Make the ignored-file scan exceed its timeout before the agent
starts.
3. Before this fix, the snapshot returns null and the run ends as
non-retryable `setup_failed`.
4. Use the disposable browser fixture in
`tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and
test the full recovery path.

**Paperclip version or commit**

Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`.

**Deployment mode**

Built from source. The defect is in core workspace setup, not a specific
model provider.

Related work: Refs #13442 (managed repository preparation), Refs #11572
(bounded Git scheduler), Refs #12997 (separate adapter startup retry
work), Refs #13469 (separate terminal-workspace scan performance work).

## What Changed

- Return the non-Git fallback only for repository discovery. Propagate
failed scans of a confirmed repository.
- Replace full ignored status output with an ignored-only directory
listing. Preserve NUL-delimited paths and exclusions.
- Preserve typed, sanitized scan errors through workspace preparation
and persist pre-provider failure details.
- Retry only timeouts and queue saturation, using the existing durable
two-retry budget and issue gates. Prevent generic recovery from adding
another budget.
- Show workspace-specific failure copy and actionable exhausted-recovery
notices.
- Add red-green unit tests, real-database restart and retry-boundary
tests, and opt-in browser acceptance fixtures with real Git subprocess
timeouts.
- Document the recovery contract and browser verification procedure.

## Verification

- Red: injected scan failures returned null instead of rejecting; setup
lost the timeout code; task-thread and recovery notices had generic
copy.
- Green: 119 focused adapter/backend tests, 20 recovery-boundary tests,
and 136 task-thread tests.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed on the final production code.
- `pnpm check:token-gates` — passed.
- The initial local `pnpm test:run` overlapped source edits and was
interrupted after two late-added assertions saw pre-fix behavior; it is
not counted as a green full run. A fresh final-head run passed all 229
tests across the six affected adapter/backend/UI suites. The clean
latest-head CI full test matrix passed: all five general-server shards,
all five serialized-server shards, and all three general-workspace
shards.
- Latest-head CI also passed all three browser shards and their
aggregate gate, typecheck and release registry, build, runner
verification, canary dry run, policy, Docker context integrity, and
security gates. Greptile: 5/5, with the review thread resolved.
- Browser: created a task in a disposable instance. A real Git timeout
scheduled recovery, the next run completed through the run-scoped API
without manual Retry, and Done survived reload. The deterministic
process worker checked preserved source edits and excluded private
files; no model calls were made.
- `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec
playwright test --config
tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6
minutes). The persistent case made exactly three failed attempts, never
started the worker, showed the cause-specific notice, stayed stopped for
another scheduler tick, and retained Blocked after reload.
- Extra red-green coverage: 50 recovery tests passed after fixing an
exhausted-bootstrap classification that incorrectly implied unknown
provider actions. Missing or uncertain evidence still retains the safety
hold.
- Verified the documented Git executable override during repository
seeding.

## Risks

- A confirmed repository scan failure now fails closed instead of
falling back to directory sync. This prevents unfiltered copying but
makes previously hidden errors visible.
- Temporary host problems can create up to two additional setup
attempts, 30 seconds apart. Permanent scan errors do not auto-retry.
Generic recovery cannot reset this budget.
- The durable retry path still enforces ownership, pause, and work
eligibility. Integration tests cover restart, duplicate promotion,
pause, exhaustion, and non-retryable categories.
- No schema migration, new runtime setting, new retry budget, production
deployment, or historical task replay.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, repository
tools, shell execution, and browser testing. The exact deployment model
ID and context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 12:37:21 -05:00
Devin FoleyandPaperclip c0fda8fac5 fix(apps): restore action test picker scrolling and agent eligibility (#13414)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps action tests let operators use an agent's permissions.
> - The agent picker must scroll inside the test dialog.
> - Its body portal sits outside the dialog's scroll boundary and blocks
wheel input.
> - Admin permission bypasses also skip agent lifecycle checks.
> - This PR fixes scrolling and rejects agents that cannot receive
assignments.

## Linked Issues or Issue Description

**What happened?**

The Act as picker does not scroll with the mouse wheel inside an action
test dialog. Admins can also see terminated agents.

**Expected behavior**

The list scrolls normally. Terminated and pending-approval agents are
absent. Direct requests cannot test an action as one of those agents.

**Steps to reproduce**

1. Create enough agents to overflow the list. Terminate one agent.
2. Open a connected app's Permissions tab. Click Test on an action.
3. Open Act as and use the mouse wheel over the list.
4. Check whether the terminated agent appears as an admin.

**Paperclip version or commit**

Reproduced on master at f2c5e54dc. Rebased onto d351e08de.

**Deployment mode**

Local dev, built from source. This is a shared Apps bug and does not
require Railway credentials.

Related search result: #9918 added search to a separate secrets picker.
It does not cover this action test dialog. No duplicate action-test
picker PR was found.

## What Changed

- Keep the action tester's agent popover inside its dialog's scroll
boundary.
- Check company membership and the shared agent lifecycle policy before
assignment permission bypasses.
- Reject terminated and pending-approval agents in lists, previews, and
test calls.
- Reuse company-scoped rows during listing to avoid extra per-agent
queries.
- Add route tests for both admin modes and a browser wheel-scroll
regression.

## Verification

- 335 focused tool-access and TestPanel tests passed after rebase. After
the review cleanup, both admin regressions and writable-agent selection
passed again (3 tests).
- The browser regression failed before the fix because wheel input left
scrollTop at zero. It passed after the fix, including search and
selection. It executes no provider tools.
- Full typecheck, build, and token gates passed during implementation.
Token gates passed again after rebase.
- The full test run reported a failure in the GitHub installation
recovery chat test. That test passed in isolation. The full run was
stopped after the failure, so later groups were not completed.

Manual check: open an action's Test dialog, open Act as, scroll, search,
and select an agent. Terminated agents must be absent.

## Risks

The portal change affects only the picker inside the action test dialog.
The browser test covers scrolling and selection. Paused agents remain
eligible under existing assignment rules. There is no migration or
provider policy change.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
The exact serving model ID and context window were not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 21:18:12 -07:00
Nicky LeachandClaude Opus 5 b64469e403 feat(workspaces): add an operator default for isolated execution workspaces (#13444)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The execution workspace subsystem decides if a task run uses the
shared project checkout or an isolated per-task git worktree
> - The mode comes from the project policy, then the task settings. A
project that stores no policy always falls back to the shared checkout
> - An operator who wants every project to use isolated workspaces must
therefore edit each project one at a time, and must repeat this for each
new project
> - There is no instance-level control, so a fleet operator cannot set
this default at all
> - This pull request adds a managed experimental flag that moves the
default for projects that store no policy of their own
> - The benefit is that an operator sets the workspace default one time,
and every current and future project follows it

## Linked Issues or Issue Description

No public issue exists. The description follows the feature request
template.

**Subsystem affected**

Execution workspaces. The files are
`server/src/services/execution-workspace-policy.ts` and the run dispatch
path in `server/src/services/heartbeat.ts`.

**Problem or motivation**

`resolveExecutionWorkspaceMode` reads only the project policy, the task
settings, and a legacy field. Its last statement returns
`shared_workspace`. A project that stores no policy always gets the
shared project checkout.

An operator has no way to change this default for many projects at the
same time. The operator must edit each project, and must edit each new
project again later. Tasks in one project therefore share one checkout,
and they run one at a time when the environment driver makes the
scheduler serialize them.

**Proposed solution**

Add the managed experimental flag `enableIsolatedWorkspacesByDefault`.
When the flag is on, a project that stores no policy of its own resolves
as if it selected isolated workspaces. A project that stores a policy
keeps that policy.

The new helper substitutes a project policy. It does not move the last
statement of `resolveExecutionWorkspaceMode`. Two behaviors make this
necessary:

- A task that has no project must keep its current behavior. An isolated
workspace needs a repository to cut a worktree from.
`isUnrunnableWorktreeCombo` blocks an isolated task that has no
`projectId` and no `projectWorkspaceId`. A moved fallback would resolve
isolated for project-less tasks, such as agent chat, and stop them
before dispatch.
- The mode and the strategy must agree.
`buildExecutionWorkspaceAdapterConfig` supplies the default
`git_worktree` strategy only when one layer asserts workspace control. A
moved fallback would leave isolated mode with a `project_primary`
strategy.

**Alternatives considered**

- Change the last statement of `resolveExecutionWorkspaceMode` to
`isolated_workspace`. This is one line, but it changes the default for
every deployment. It is also not gated, so it would apply where isolated
workspaces are off.
- Write the policy to each project row with a script. This does not
cover new projects, and it does not cover new instances.
- Add an instance defaults section to the managed-config document. This
needs a new document key, new validation, and new delivery code. A
boolean flag reuses the delivery machinery that exists today.

**Roadmap alignment**

`ROADMAP.md` does not list execution workspace defaults. This change
adds an operator control to an existing capability. It does not add a
new capability.

**Additional context**

The flag is `tier: "managed"`. A cloud operator can therefore deliver it
with the managed-config machinery that exists today. No new delivery
code is needed.

## What Changed

- Add `enableIsolatedWorkspacesByDefault` to the feature catalog with
`tier: "managed"`. Both defaults are off.
- Add the flag to the experimental settings schema, the type, and both
branches of `normalizeExperimentalSettings`.
- Add `applyDefaultIsolatedExecutionWorkspacePolicy` to
`execution-workspace-policy.ts`. It substitutes `{ enabled: true,
defaultMode: "isolated_workspace" }` only when the flag is on, the task
has a project, and the project stores no policy.
- Apply the helper in the run dispatch path in `heartbeat.ts`, after the
existing `gateProjectExecutionWorkspacePolicy` call. The `hasProject`
argument reads the resolved project row, not the raw `projectId` of the
task.
- Gate the new flag behind `enableIsolatedWorkspaces` at the call site.
The new flag does nothing on its own.
- Add a toggle card to the instance experimental settings page. The card
shows only when isolated workspaces are on.
- Add eight tests for the new helper.

## Verification

Commands:

```
pnpm --filter @paperclipai/shared typecheck
pnpm --filter @paperclipai/ui typecheck
cd server && ../node_modules/.bin/tsc --noEmit
```

The server typecheck script also builds a Rust binary. I ran `tsc`
directly because this machine has no `cargo`. The server package reports
no type errors.

Tests:

```
./node_modules/.bin/vitest run \
  server/src/__tests__/execution-workspace-policy.test.ts \
  server/src/__tests__/instance-settings-service.test.ts \
  server/src/__tests__/instance-settings-cloud-defaults.test.ts \
  server/src/__tests__/instance-settings-managed-overlay.test.ts \
  server/src/__tests__/instance-settings-routes.test.ts \
  server/src/__tests__/managed-config.test.ts \
  server/src/__tests__/heartbeat-workspace-busy.test.ts \
  server/src/__tests__/heartbeat-workspace-session.test.ts \
  server/src/__tests__/heartbeat-workspace-ready-comment.test.ts \
  server/src/__tests__/execution-workspaces-service.test.ts \
  server/src/__tests__/issue-runtime-workspace-binding.test.ts \
  server/src/__tests__/run-trust-preset.test.ts \
  packages/shared/src/feature-catalog.test.ts \
  packages/shared/src/settings-visibility.test.ts \
  packages/shared/src/validators/instance.test.ts \
  ui/src/pages/InstanceExperimentalSettings.test.tsx \
  ui/src/components/Sidebar.test.tsx
```

All of these files pass. The new tests cover each of these cases:

- The helper substitutes an isolated policy for a project that stores
none.
- The helper changes nothing while the flag is off.
- The helper changes nothing for a task that has no project.
- The helper keeps a stored policy, including a policy with `enabled:
false`.
- The resolver returns `isolated_workspace` for an unpolicied project.
- An explicit task setting still wins over the operator default.
- The substituted policy produces the `git_worktree` strategy.
- A project-less task does not become an unrunnable worktree.

To confirm the behavior by hand:

1. Turn on Isolated Workspaces, then turn on Use Isolated Workspaces By
Default.
2. Open a project that has no execution workspace policy.
3. Start a task in that project.
4. The run gets its own worktree. Tasks in that project no longer wait
for each other.

## Risks

Low to medium. The details:

- The flag defaults to off, and it is inert unless
`enableIsolatedWorkspaces` is also on. An instance that does not turn on
both flags sees no change.
- A project that stores a policy keeps it. This includes a policy with
`enabled: false`, which the helper reads as a decision to stay on the
shared checkout.
- When an operator turns the flag on, the workspace configuration
fingerprint changes for projects that store no policy. Their next run
creates a new workspace. This is correct, because the mode did change,
but the first run after the change does more setup work.
- A task that is in flight when the flag changes resumes with a
different workspace path than the path its session remembers. An
operator should let current runs finish before turning the flag on.
- Isolated workspaces use more disk, because each task gets its own
worktree.

## Model Used

Claude Opus 5 (`claude-opus-5`) in Claude Code, with extended thinking
and tool use. The model read the repository, made the change, and ran
the typechecks and tests above.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 20:59:56 -07:00
DottaandPaperclip f912ecaacf fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents hire other agents and assign tasks to them.
> - Managed AI connections must follow those hires across legacy and
native runners.
> - Missing accounts should pause task execution and let the user
connect from the task.
> - Subscription contention must wait without asking for new
credentials.
> - This pull request fixes these paths and the native tool and Daytona
staging failures found during live tests.
> - The result is a working hire, subtask, and connection setup flow on
local and remote runners.

## Linked Issues or Issue Description

**What happened?**

A managed Claude or Codex agent could hire a teammate without a usable
AI binding. Cross-provider hiring could fail before the user had a
chance to connect the new provider. First-time task setup did not show
the existing AI credential form inline. A busy subscription could
request a new connection. Native API replies could stop the parent after
a hire had already committed. Fresh Daytona sandboxes could fail to
extract read-only skill directories created on macOS.

**Expected behavior**

Compatible hires inherit the managed connection choice. A hire for
another provider uses the responsible user's default. If that account is
missing, the hire succeeds and the task asks for a connection.
Completing setup in the task resumes work automatically. Explicit child
auth settings and existing unmanaged login paths keep precedence.
Shared-account access checks remain in force.

**Steps to reproduce**

1. Connect a Claude or Codex parent with a managed AI account.
2. Ask it to hire one agent of each provider and create a self-assigned
subtask.
3. Assign work to both hires without connecting the second provider
first.
4. Connect the missing provider from its task card.
5. Check that all tasks finish and same-provider work uses the original
account.
6. Repeat with native runners and fresh Daytona sandboxes. The opt-in
browser suite in `tests/hiring-ai-connections/README.md` performs these
steps.

**Paperclip version or commit**

The live failures were reproduced from `f2c5e54dc`. The branch is
rebased onto `5282cabde`.

**Deployment mode**

Isolated local development instance. Legacy CLI and native runners.
Local execution and ephemeral Daytona sandboxes.

Related work: Refs #13247 for managed AI connections. Refs #13268 for
legacy credential-reference inheritance, which this branch preserves.
Refs #13432 for a concurrent managed-inheritance fix. This PR also
covers cross-provider task setup, subscription waits, native API
replies, and Daytona extraction. It permits missing responsible-user
defaults at hire time; restricted shared selections still fail.

## What Changed

- Apply managed connection defaults to both agent creation routes.
Preserve explicit auth choices and legacy credential-reference
inheritance.
- Allow hires before their responsible user connects the provider. Keep
approval gates, company boundaries, and shared-account access checks.
- Reuse the production AI credential form inside the pending task card.
Resume the task after setup.
- Retry subscription lease contention without consuming the
provider-failure allowance or creating a connection request.
- Require task execution-lock ownership when scheduling, promoting, and
dispatching subscription retries. Recheck ownership under the issue row
lock.
- Rename the HTTP operation identity at the native tool boundary so it
cannot override the runner's operation identity.
- Delay directory permission restoration during Daytona extraction.
Preserve the final read-only modes.
- Add database-backed regressions, real browser acceptance tests, and
Storybook states. Document setup and run-log behavior.

## Verification

- Six real browser scenarios passed: both parent providers on legacy
local and legacy Daytona; native Codex locally; native Claude on
Daytona. Each scenario hires both providers, completes a self-subtask
and assigned work, and connects the missing provider inline with
automatic continuation.
- Successful runs verify the account, responsible user, runner mode, and
Daytona lease. All 18 test sandboxes were deleted.
- Live authentication used API keys. Subscription inheritance, lease
contention, and retry have integration coverage. Fresh subscription
OAuth sign-in was not automated.
- Red/green tests reproduced missing bindings, missing inline forms,
subscription contention, native API reply failure, and GNU tar
permission failure.
- Seven Storybook browser checks passed. They cover both providers,
method selection, narrow layout, completion, cancellation, and invalid
credentials.
- Full local suite coverage completed before rebase. Initial timing and
fixture startup failures passed unchanged on isolated reruns. The first
full command did not exit cleanly; the remaining workspace and
serialized groups were completed separately.
- After rebase, 107 hiring/auth/retry tests and 59 native API,
task-card, and Daytona tests passed. The full workspace typecheck,
production build, and token gates passed again. Storybook build passed
before rebase.

- Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and
four explicit cancellation-race cases passed. Eight cross-provider cases
cover stale auth keys on both creation routes and both runner types.
Server typecheck and build passed.

- Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks
passed. Two optional Storybook jobs were skipped. The full server,
workspace, browser, native runner, build, typecheck, and release checks
passed.
- Three unchanged tests initially failed on a busy port, a chat row-lock
race, and preview-server readiness. Each affected job passed after one
CI rerun. Isolated local checks also passed: 41 credential tests, the
chat-concurrency case, and 25 preview-runtime tests.
- Greptile reviewed the final commit at 5/5. Both review threads are
resolved. GitHub reports no merge conflicts.

## Risks

- A missing personal account now defers authentication to the first
task. Explicit incompatible bindings and restricted shared accounts
still fail at hire time.
- An inherited personal default uses the responsible user's existing
authorization to install access for the new agent. It never copies
credentials or another user's identity.
- Subscription contention retries after a delay and rechecks task
eligibility. It does not consume the provider-failure budget.
- Native hiring uses the existing managed API-tools opt-in. Remote
native runners require a matching Linux binary and provider pack, as
documented in the acceptance README.
- No schema changes. Live tests make paid provider calls and remain
opt-in.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository
inspection, code execution, browser automation, and API tools. The
context-window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 16:15:48 -05:00
DottaandPaperclip 5282cabde8 fix(ui): make every agent reachable from Chats (#13420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Experimental Agent Chat gives each person a persistent conversation
with each agent.
> - The sidebar only showed starred or previously visited agents.
> - A person could not start a conversation with an agent absent from
that list.
> - This pull request adds a searchable agent picker and a default
sidebar entry.
> - People can now reach every company agent while keeping their
frequent conversations close.

## Linked Issues or Issue Description

Refs #13283, which introduced the existing task-backed agent chat
surface. This change adds discovery to that surface after maintainer
review of the Storybook design.

**What happened?**

With Agent Chat enabled, a person with no starred or recent agents had
no agent shortcut. The sidebar also lacked a direct way to start a chat
with another agent.

**Expected behavior**

Show the first-created company agent by default. Let the person search
all company agents and open a persistent conversation.

**Steps to reproduce**

1. Enable Agent Chat in a company with agents.
2. Use an account with no starred agents or recent chats.
3. Look for a chat entry in the sidebar, or try to chat with an agent
absent from its shortcuts.

**Paperclip version or commit**

The existing behavior is present on master at 728f7185f.

## What Changed

- Add a compact picker that searches every company agent by name or
role. Include loading, retry, empty, and no-match states.
- Keep the first-created agent in Chats alongside stars and up to four
other recent conversations.
- Align the compose button with stars. Show it on hover or keyboard
focus, and keep it visible on touch devices.
- Use the existing company-aware chat route. Close the mobile sidebar on
selection and reset the picker when the company or user changes.
- Connect the approved Storybook scenarios to the production sidebar and
picker. Update product and design-guide documentation.

## Verification

- Ten focused tests pass for ordering, duplicate names, role search,
keyboard selection, search reset, retry states, full-roster navigation,
mobile dismissal, and company/user changes.
- The agent-chat sidebar end-to-end test passes against an isolated
PostgreSQL container. It verifies the first-use default, picker
navigation, lazy creation, stars, draft preservation, and reopening
existing history. The old four-row and See all link expectations now
match the approved behavior.
- Browser review of the production components in Storybook: open the
sidebar picker, search for an agent absent from the shortcuts, select
with Enter, and send a message through the existing composer.
- `pnpm build`, `pnpm -r typecheck`, `pnpm check:token-gates`, and `pnpm
build-storybook` pass on the rebased branch. The full local test run
reached 12,020 passing server tests, then stopped because an unrelated
environment-custom-images suite could not allocate PostgreSQL shared
memory on this Mac. All corresponding CI test shards pass on isolated
hosts.
- All CI checks pass on commit
`c324568324a6a8b65adfc6447ef19bcc710c61df`. Greptile reviewed that
commit at 5/5 with no open findings.
- Reviewer path: enable Agent Chat, open the Chats compose button,
search by name or role, and select an agent. Existing conversation
history must remain available.

## Risks

- Sidebar ordering changes: the earliest-created agent remains visible
even without a visit or star.
- The hover control remains available to keyboard and touch users.
- This change uses the existing feature gate, APIs, membership
preferences, and persistent conversation behavior. It adds no migration
or server contract.
- Storybook uses fixture responses. Its first-use landing page is a
review fixture and is not added as a product route.

## Model Used

OpenAI GPT-6 Astra (`gpt-6-astra`) through Codex, with reasoning, code
execution, and browser tools. The session does not expose its context
window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 13:19:13 -05:00
DottaandPaperclip 78ce96a48b fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner supplies each agent with instructions, assigned
skills, and tools.
> - Native Claude lost its skill snapshot before provider launch. It
also changed exact model IDs to aliases.
> - The read permission mode denied ordinary Paperclip reads because
this runner has no interactive approval handler.
> - This pull request restores the missing context and permits only
assigned tools that Paperclip defines as reads.
> - Other requests that need approval stop with a clear action for the
operator. They do not retry automatically.
> - Agents can complete read tasks. Users can see why a restricted task
stopped.

## Linked Issues or Issue Description

**What happened?**

Native Claude could not find assigned skills. The ACP layer could
replace an exact model ID with an alias and then fail identity
verification. The default `approve-reads` mode denied Paperclip read
tools. A denied write showed a generic transport failure.

The tool descriptions also called live operations “mock” operations. The
completion prompt said to call a completion tool once, although the
protocol can reject a claim and require a corrected call.

**Expected behavior**

Load assigned skills before launch. Keep the selected model ID. Allow
assigned Paperclip reads. Stop an operation that needs approval with
clear instructions when no approval handler exists.

**Steps to reproduce**

1. Configure a native Paperclip Runner agent with ACPX Claude and an
exact model ID.
2. Assign a skill and ask the agent to use it.
3. Select the read permission mode and ask the agent to read task
context and list documents.
4. Ask the agent to write a document. Check the task state and recovery
message.

**Paperclip version or commit**

Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`.

**Deployment mode**

Local native runner with real Claude, Rust runnerd, the Paperclip
server, and embedded PostgreSQL.

Related: #13196 fixed remote skill staging in the legacy `claude_local`
adapter. The native runner uses a separate path, which this pull request
fixes.

## What Changed

- Carry the runtime context through Rust and the ACPX sidecar. Load
assigned Claude skills after the provider lifetime lease is held.
Refresh the files on each open.
- Write the exact requested model ID into the isolated Claude settings.
- Grant exact MCP permissions for the intersection of assigned tools and
Paperclip's read catalog. Provider hints cannot grant access. Existing
task-control permissions stay in place.
- Stop requests that need an unavailable approval handler. Preserve the
typed error through the server. Show “Approval required” on the task and
require operator action without automatic retry.
- Label the setting “Allow Paperclip reads.” Remove “mock” from live
tool descriptions and regenerate the contracts.
- Change one completion-prompt sentence to require one accepted result.
Add regression tests and update the runner documentation.

## Verification

- Red/green regression tests cover read admission, unavailable approval
handling, server recovery, and the task error message.
- A real Claude read trial failed before the fix and completed after it:
[red
trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1),
[green
trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1).
- Full browser tests used real Claude and the production tool authority.
Reads completed. A write stopped with “Approval required.” No document
was created and no automatic retry was scheduled. All nine checks
passed: [events and
screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc).
- A separate full browser test assigned a skill, invoked it, and
completed with a marker absent from the task prompt: [skill
evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52).
- Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests,
27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm
build`. The server/UI red-green checks and `pnpm check:token-gates` also
passed. The full suite passed in CI, including all general and
serialized test shards, all three browser shards, and runner
`check:all`. The duplicate unsharded local full-suite run was stopped
after CI passed.

Braintrust links require project access. The traces contain
provider/runner events and app outcomes. They do not contain raw model
HTTP requests.

## Risks

- `approve-reads` now allows assigned Paperclip reads. Other operations
that need approval stop the turn. Users must review the operation and
change permissions before retrying.
- The Claude settings depend on the pinned ACP and SDK behavior.
Automated tests and real Claude trials cover this boundary.
- The native Codex path is unchanged. The ACPX Codex fallback shares the
clearer approval failure handling.
- Local execution was tested end to end. Remote execution was not run.
This change has no database migration.

## Model Used

OpenAI GPT-6 through Codex. The agent used reasoning, code execution,
and browser tools. The exact deployment ID and context-window size were
not exposed in the session. Live acceptance tests used
`claude-sonnet-5`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:56:24 -05:00
DottaandPaperclip 728f7185f6 feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Self-hosted boards need a way to show occasional product
announcements.
> - An app release should not be required to publish or withdraw a card.
> - Native card controls keep publishing consistent; the hero can use a
static image or isolated HTML/CSS animation.
> - This pull request renders a validated JSON feed with native
components.
> - It stores dismissals per account on each instance, so a closed card
stays closed across companies and browsers.
> - Named staging feeds let authors test content before production
publication.

## Linked Issues or Issue Description

**Subsystem affected**

Board application shell, announcement delivery, and user preferences.

**Problem or motivation**

Operators need a small, optional announcement card. Users need reliable
dismissal state. Authors need to test remote content without changing
the production feed.

**Proposed solution**

Add one non-modal AnnouncementWell. Fetch validated JSON and
content-addressed media through the instance server. Keep card controls
native, with optional sandboxed HTML/CSS animation in the hero. Use
stable announcement IDs for dismissal, an explicit empty manifest and
quiet 404 handling. Provide a staged publishing helper and isolated
test-drive guide.

**Alternatives considered**

Hosting the entire card as a page would move navigation and dismissal
into remote content. This change limits HTML to a scriptless, isolated
visual hero and keeps controls native. Browser-only storage would lose
dismissals across browsers, so the instance stores account preferences.

**Roadmap alignment**

ROADMAP.md has no overlapping announcement feature. A GitHub title
search found no related announcement pull requests. This work implements
a maintainer-requested feature.

## What Changed

- Add shared feed types, strict validation of every object, supported
routes, expiration and version checks.
- Add a board-only current-feed API, constrained media proxy, and
idempotent dismissal API. Store the first dismissal and its company
audit entry in one transaction.
- Cache upstream data for one hour. Use conditional requests, request
deduplication, response limits, public destination checks, and a
three-second deadline. Treat a remote 404 as an empty feed with a
fifteen-minute retry cooldown.
- Keep announcement visibility stable when focus moves to browser chrome
or another app pane; only tab visibility starts a return check.
- Add a responsive native announcement card. Respect onboarding, dialogs
and toast placement. Sync pending dismissals across tabs and retry after
reconnect or return.
- Add idempotent migrations for dismissals and validated publication
IDs, design-guide examples, static and animated Storybook examples, and
focused tests. The publication registry supports offline retries without
accepting caller-invented IDs.
- Add HTML/CSS animated heroes with static posters, automatic playback,
reduced-motion handling, strict DOMPurify validation, an empty iframe
sandbox and CSP that blocks scripts/network resources.
- Add validated staging publication, content-addressed assets, an empty
production manifest, preview fixtures, and authoring/operator
documentation.

## Verification

- The preceding implementation passed 98 targeted
shared/server/publisher/route/OpenAPI/UI tests and 127 tests including
the master rebase. The playback-control removal passes all 21
announcement UI tests, covering the rendered sandbox, fallback, reduced
motion, dismissal and slow/stale state lookups. The preceding
shared/server tests cover HTML validation and response sandbox headers.
- The playback-control removal passes UI typecheck, production UI build,
Storybook build and token gates locally. Browser verification confirms
the animated card has only its dismiss button and two links, with no
page errors. The full canonical CI matrix passed on current head
`00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two
optional Storybook deployment checks skipped. This run needed no
retries. Greptile reviewed this same head at 5/5 with no outstanding
findings.
- The local canonical general-server run passed 12,063 tests before
reporting embedded-PostgreSQL startup failures in an unrelated fixture.
All 31 tests in that fixture passed across isolated retries. The UI
group passed 6,219 tests and other workspace groups passed 3,201; two
CLI database-startup failures also passed individually. Serialized
server suites were verified by the full CI matrix rather than repeating
them locally. No source changes were needed for these environment
failures.
- The real S3/CloudFront staging manifest and both media asset headers
were verified. Production remains empty/unpublished. The guide
distinguishes the preview host's disabled edge cache from production
cache requirements.
- In the isolated test-drive, the animation visibly moves without
playback controls. A 390×844 browser viewport keeps the card above
navigation. Reduced motion makes no animation request. Both themes
render correctly and browser page errors are empty. Browser fault
injection verified that scripts cannot execute and CSS cannot make
network requests; a missing animation leaves its poster and controls.
- Refresh leaves the animated card visible. Closing it persists after
reload and the API returns null. Earlier live checks verified dismissal
across browsers, company-relative CTA navigation, modal
deferral/restoration, and new-ID eligibility after restarting the same
database.
- The deployed empty feed and a real remote 404 return HTTP 200 with
null from the board API, with a usable dashboard and no announcement
popup or browser warnings.
- Authoring documentation covers staging, animated HTML constraints,
test-drive, withdrawal, ID reuse and cache-refresh steps.

## Risks

- Animation supports self-contained visual HTML/CSS and inline SVG,
without JavaScript or external resources. A static image is required.
Older builds that do not recognize the optional animation field quietly
hide that unsupported feed.
- The default feed makes an outbound request from an instance when a
board is used. Operators can disable it. Requests contain no account
IDs, company data, cookies or interaction events.
- Feed publication and withdrawal can take about 65 minutes to reach
returning users because of CDN and instance caches. Expiration also
removes visible cards locally.
- Dismissals follow an account within one instance. No-login instances
share the existing local-board identity. Separate installations do not
share state.
- Both tables are additive. A unique key prevents duplicate dismissals;
the transaction prevents duplicate first-dismissal audit entries. The
publication registry retains only validated IDs. AGENTS.md and the
implementation spec document the required exception to company scope for
these instance-level records.
- Publication was limited to separate public staging prefixes on the
existing preview host. Production remains empty/unpublished. No AWS
policies or infrastructure were changed.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and
context-window size are not exposed in this session. Capabilities used:
reasoning, code editing, shell execution, tests, browser interaction,
and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:19:54 -05:00
DottaandPaperclip f1d57863d2 fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels.

Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-14 10:06:06 -05:00
Michael Nguyen 2d47a8058d fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly
approves.

## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connector screens need recognizable app artwork.
> - Several bundled marks have inconsistent artwork or dark-theme
behavior.
> - The shared logo component should retain the existing frame.
> - This change replaces selected artwork and fixes local theme
fallback.
> - Users see consistent connector icons across shared component
callers.

## Linked Issues or Issue Description

**Current behavior**

Some connector marks use outdated artwork or unsuitable theme variants.
A remote dark logo can override a canonical local mark that works in
both themes.

**Proposed behavior**

Use the selected bundled artwork with the existing gray rounded frame.
Use local light artwork in both themes unless a distinct local dark
variant exists.

**Subsystem affected**

Connector artwork, the shared AppLogo resolver, and its validation.

**Breaking changes**

No connector capability, permission, credential, or catalog activation
changes. No duplicate PR was found in the earlier search.

## What Changed

- Update 46 artwork files and only the app definitions whose logo paths
need to change.
- Keep a compact public manifest of identities, paths, visibility, and
aliases.
- Prefer local artwork in both themes and preserve the existing frame
and padding.
- Add locally runnable artwork safety checks and a light/dark Storybook
gallery; leave PR workflows unchanged.
- Document artwork conventions. Source research is kept outside the
public manifest.

## Verification

- Artwork check: 69 identities pass.
- Node artwork validation tests: 13 pass (rechecked September 14).
- Focused resolver, component, and catalog tests: 40 pass (rechecked
September 14).
- Maintainer decision: dedicated icon-validation CI is not required; the
workflow remains unchanged. Greptile acknowledged 5/5 with no remaining
code concerns on September 14.
- UI typecheck, token gates, and Storybook build pass.
- Earlier full build and repository typecheck passed. The broad local
test run was stopped after workspace-runtime dependency fixture failures
outside this change. Current GitHub CI remains the full-suite gate.
- Visual approval remains outstanding. Review the canonical icon
registry in both themes at 24–48px.

## Risks

The SVG checker rejects common active features; it is not a general
sanitizer for arbitrary uploads. Optical balance still requires human
review. Remote fallback for unknown brands retains existing behavior.
This change does not add an asset importer or new connector
capabilities.

## Model Used

OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and
browser tooling. Exact hosted model IDs and context window sizes were
not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Artwork comparison

Before and after for every affected identity. Images follow your GitHub
light/dark theme and are pinned to the base and PR commits. This
compares artwork; the existing gray rounded product frame and padding
are unchanged.

| Connector | Before | After |
|---|:---:|:---:|
| AgentMail | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg"
alt="AgentMail" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"
alt="AgentMail" width="48" height="48"></picture> |
| Airtable | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"
alt="Airtable" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"
alt="Airtable" width="48" height="48"></picture> |
| Asana | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"
alt="Asana" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"
alt="Asana" width="48" height="48"></picture> |
| ClickHouse | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"
alt="ClickHouse" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg"
alt="ClickHouse" width="48" height="48"></picture> |
| Cloudflare | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"
alt="Cloudflare" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"
alt="Cloudflare" width="48" height="48"></picture> |
| Cloudinary | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg"
alt="Cloudinary" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg"
alt="Cloudinary" width="48" height="48"></picture> |
| Discord | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"
alt="Discord" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg"
alt="Discord" width="48" height="48"></picture> |
| GitHub | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg"
alt="GitHub" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg"
alt="GitHub" width="48" height="48"></picture> |
| Gmail | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"
alt="Gmail" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"
alt="Gmail" width="48" height="48"></picture> |
| Google Calendar | <picture><source media="(prefers-color-scheme:
dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"
alt="Google Calendar" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"
alt="Google Calendar" width="48" height="48"></picture> |
| Google Chat | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"
alt="Google Chat" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"
alt="Google Chat" width="48" height="48"></picture> |
| Google Docs | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"
alt="Google Docs" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"
alt="Google Docs" width="48" height="48"></picture> |
| Google Drive | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"
alt="Google Drive" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"
alt="Google Drive" width="48" height="48"></picture> |
| Google People | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"
alt="Google People" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"
alt="Google People" width="48" height="48"></picture> |
| Google Sheets | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"
alt="Google Sheets" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"
alt="Google Sheets" width="48" height="48"></picture> |
| Google Slides | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"
alt="Google Slides" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"
alt="Google Slides" width="48" height="48"></picture> |
| Google Workspace Search | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"
alt="Google Workspace Search" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"
alt="Google Workspace Search" width="48" height="48"></picture> |
| Hugging Face | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"
alt="Hugging Face" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"
alt="Hugging Face" width="48" height="48"></picture> |
| Jam.dev (library only) | New library entry | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg"
alt="Jam.dev" width="48" height="48"></picture> |
| Jira | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg"
alt="Jira" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"
alt="Jira" width="48" height="48"></picture> |
| Linear | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"
alt="Linear" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg"
alt="Linear" width="48" height="48"></picture> |
| Manufact | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"
alt="Manufact" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg"
alt="Manufact" width="48" height="48"></picture> |
| Microsoft Teams | <picture><source media="(prefers-color-scheme:
dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"
alt="Microsoft Teams" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"
alt="Microsoft Teams" width="48" height="48"></picture> |
| Miro | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg"
alt="Miro" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"
alt="Miro" width="48" height="48"></picture> |
| Mixpanel | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"
alt="Mixpanel" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg"
alt="Mixpanel" width="48" height="48"></picture> |
| Netlify | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"
alt="Netlify" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg"
alt="Netlify" width="48" height="48"></picture> |
| Notion | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg"
alt="Notion" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"
alt="Notion" width="48" height="48"></picture> |
| PagerDuty | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"
alt="PagerDuty" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"
alt="PagerDuty" width="48" height="48"></picture> |
| PostHog | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg"
alt="PostHog" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg"
alt="PostHog" width="48" height="48"></picture> |
| Postman | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"
alt="Postman" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"
alt="Postman" width="48" height="48"></picture> |
| Shopify | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"
alt="Shopify" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"
alt="Shopify" width="48" height="48"></picture> |
| Slack | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"
alt="Slack" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"
alt="Slack" width="48" height="48"></picture> |
| Stripe | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"
alt="Stripe" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"
alt="Stripe" width="48" height="48"></picture> |
| Supabase | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"
alt="Supabase" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"
alt="Supabase" width="48" height="48"></picture> |
| Telegram | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"
alt="Telegram" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"
alt="Telegram" width="48" height="48"></picture> |
| Todoist | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"
alt="Todoist" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"
alt="Todoist" width="48" height="48"></picture> |
| Wix | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"
alt="Wix" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg"
alt="Wix" width="48" height="48"></picture> |
| Zapier | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"
alt="Zapier" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"
alt="Zapier" width="48" height="48"></picture> |
2026-09-14 05:05:27 -10:00
DottaandPaperclip 3052686b1e fix(ui): hide task chat attribution badge (#13251)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task chat shows messages from agents and users.
> - An agent message can contain an internal on-behalf-of user value.
> - The new task view showed this value as a badge next to the agent
name.
> - This badge added unwanted identity text to the task chat.
> - This pull request removes the badge from the new task view.
> - The benefit is a simpler agent identity row in the task chat.

## Linked Issues or Issue Description

**What happened?**

The new task view shows a `for <user>` badge next to an agent name when
a message has an on-behalf-of user value.

**Expected behavior**

The task chat shows the agent name without the on-behalf-of badge.

**Steps to reproduce**

1. Open the new task view.
2. Show an agent message that has an on-behalf-of user value.
3. See the attribution badge next to the agent name.

**Paperclip version or commit**

The issue occurs on `master` at commit `d0b7ba419`.

**Deployment mode**

Local dev and all other modes that use the web UI.

## What Changed

- Removed the attribution chip from the task chat agent identity.
- Added a regression test that supplies an on-behalf-of value and
confirms that the badge is absent.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/task-chat/TaskChatBubble.test.tsx` passed with 21 tests.
- `pnpm check:token-gates` passed.
- `pnpm --filter @paperclipai/ui typecheck` passed.

## Risks

- Low risk. The change only removes one badge from the new task chat
view.
- The message data remains unchanged.

> This is a small UI fix. It does not add roadmap scope.

## Model Used

- OpenAI Codex, GPT-5.4, with reasoning and tool use. The provider did
not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g.
`docs/no-internal-issue-references`, `fix/sandbox-secret-resolution`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 09:58:44 -05:00
DottaandPaperclip ef84f363b7 fix(ui): remove execution workspace access card (#13406)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Execution workspace pages show the tools and state for one isolated
workspace.
> - The workspace access card repeats runtime state that the page
already shows.
> - The card also adds open and repair actions that are not needed on
this page.
> - This pull request removes the complete workspace access card from
execution workspace detail pages.
> - The benefit is a smaller page that keeps attention on workspace
controls and work results.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The execution workspace detail page shows a workspace access status
card.

**Current behavior**

The page can show Ready, preparation, degraded, failed, or Workspace is
not running states. The card can also show Open workspace and repair
actions.

**Proposed behavior**

The page does not show the workspace access status card or its page-only
actions.

**Reason and benefit**

The card adds status and controls that are not needed on this page. Its
removal makes the isolated workspace page simpler.

**Breaking changes**

The workspace access card and its actions are no longer available from
the execution workspace detail page. Workspace runtime controls remain
available.

## What Changed

- Removed the workspace access status card from execution workspace
detail pages.
- Removed the page-only login handoff and repair mutations that served
the card.
- Added a regression test for the removed card and text.

## Verification

- `pnpm exec vitest run ui/src/pages/ExecutionWorkspaceDetail.test.tsx`
passes with 8 tests.
- `pnpm typecheck` passes.
- `pnpm check:token-gates` passes.
- `pnpm build` passes.
- `pnpm test:run` found unrelated failures in
`server/src/__tests__/workspace-runtime.test.ts` during the local broad
run. The focused changed-area tests pass.
- The complete GitHub Actions matrix passes on the latest run, including
build, typecheck, server tests, runner verification, and e2e tests.

## Risks

- Low risk. The change removes one UI card and the page-only code that
supported it.
- Users must use the remaining workspace runtime controls instead of
this card.

> This change is focused UI cleanup. It does not add a roadmap feature.

## Model Used

- OpenAI Codex based on GPT-5, with reasoning, tool use, and code
execution. The exact context window is not exposed in this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 09:47:10 -05:00