Commit Graph
594 Commits
Author SHA1 Message Date
DottaandPaperclip 8725d6ce09 fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions.

Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:35:33 -05:00
Devin FoleyandPaperclip e3d8fb0876 feat: attest standard production images at full source commits (#13797)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators deploy its standard production container on several CPU
architectures.
> - Downstream image builders need to identify the exact source of their
base image.
> - A short commit tag does not provide signed source evidence.
> - This pull request adds a full commit tag and signed image digest for
canonical master pushes.
> - Consumers can verify the source and compose from the immutable
digest.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Publication of the standard multi-platform production image.

**Current behavior**

The Docker workflow publishes short commit tags and channel tags. It
does not provide a signed standard-image contract tied to the complete
master commit.

**Proposed behavior**

Canonical master pushes also publish `sha-<full-commit>` and attest the
exact index digest after platform validation and an immutable-image
orphan-reaping check. The signer certificate binds the source
repository, commit, workflow and ref. Existing tags and the separate
cloud producer remain available.

**Reason and benefit**

Downstream builders can prove the source of a standard base without
adding their dependencies or repository details to the public workflow.
No matching open issue or duplicate PR was found.

## What Changed

- Add the canonical full-SHA tag without changing existing tag mappings.
- Validate amd64 and arm64 descriptors and hash the exact registry
response bytes and require its digest header to match.
- Verify the immutable image and sign it with GitHub artifact
attestations.
- Run the contract tests in trusted PR verification and document the
consumer contract.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/standard-image-contract.test.mjs`: 37 passed.
- A read-only check against an existing published index returned its
exact expected digest.
- `actionlint -shellcheck='' .github/workflows/docker.yml
.github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports
only existing `ls` and word-splitting warnings.
- `pnpm build`: passed locally with Cargo available.
- `pnpm -r typecheck`: passed locally.
- Full local Vitest was attempted: 8,260 passed, 14 failed, with 34
failing suites. The failures were missing embedded-PostgreSQL library
aliases in this fresh install and existing macOS runtime-skill-cache
rename errors. Native aliases are now restored. Rerunning the 33
affected database suites produced 577 passes and two unrelated AgentMail
skill-root lookup failures (32 suites passed). The four directly failing
database tests also pass independently. This is not a claim that the
full local suite passed.
- All final-head CI checks pass. One unrelated routine-route mock
assertion passed on the single-shard retry; its 15 tests also pass
locally. Greptile is 5/5 on this exact head, with no unresolved threads.
- Actual signing requires a canonical master push. This draft PR does
not publish trusted provenance.

## Risks

The new attestation step requires OIDC and attestation write permissions
in the merge job. Signing failure leaves the image available but without
the new admission proof. Consumers must fail closed when proof is
missing. Existing release tags, the legacy producer, and image retention
remain unchanged. No database or application behavior changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell
execution and test tools. The session does not expose a more specific
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 07:33:30 -07:00
Devin FoleyandPaperclip 8326e33ada Fix oversized sandbox process launch payloads (#13793)
## Thinking Path

> - Paperclip manages agent work and preserves context across retries.
> - Sandbox ACP runs encode their command and environment into one
launch value.
> - A long continuation can make that encoded value exceed Linux's exec
limit.
> - The launch shell then exits before the agent can initialize.
> - This PR transfers large command envelopes through a private
temporary file.
> - The agent receives the complete environment and can start normally.

## Linked Issues or Issue Description

Refs #13777.

**Bug description**

Sandbox tasks with long retry context fail during ACP initialization
with exit 127. The streamed bridge also drops the shell error that
explains the failure.

**Steps to reproduce**

Launch the streamed sandbox process bridge on Linux with one valid
110,000-byte environment value. Its base64 command envelope exceeds the
limit on one exec argument or environment string. Daytona reports
`argument list too long: env`, then exit 127.

**Expected behavior**

The command envelope must not make a valid child environment too large
to launch. A shell startup failure must retain its diagnostic in the run
log.

## What Changed

- Keep envelopes up to 64 KiB on the existing launch path. Upload larger
envelopes in bounded chunks inside a mode-0700 session directory. Set
the final payload file to mode 0600.
- Read the file without following symlinks and delete it before spawning
the child. Remove incomplete uploads on failure. Both streamed and
polled bridges use the same envelope.
- Preserve stderr when the launch shell fails before the wrapper emits a
terminal event. Emit a fixed terminal error and shutdown acknowledgement
if a payload cannot be read or parsed, without exposing its contents.
- Cover large environments, file permissions, payload deletion,
interrupted uploads, missing or malformed payloads, and startup
diagnostics. Document the transfer and cleanup behavior.

## Verification

- Both large-envelope regressions fail before the fix and pass after it.
- Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass:
370 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- A gated live Daytona probe on the current sandbox image reproduced
exit 127 with the old bridge. The fixed bridge launched the same command
successfully. A second probe completed real Claude ACP initialization
with a 110,000-byte context value. It did not run an agent task. All
temporary sandboxes were deleted.
- `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner
step and stop because this machine has no `cargo` executable.
- Full CI passes on `ea816cf58d`: [run
35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754).
All 53 checks pass; two optional checks are skipped. The local
full-suite run was stopped after equivalent CI suites passed; it has no
final local result.
- Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads.
The branch is mergeable.

## Risks

Large envelopes require extra upload calls during startup. The temporary
data stays inside the private session directory and is removed before
child startup or during failure cleanup. Individual child environment
values still obey the operating system's native limits. No migration or
configuration change is required.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 20:54:07 -07:00
Devin FoleyandPaperclip c496acb570 Return client errors for known OAuth reconnect states (#13794)
Map missing OAuth refresh credentials and terminal reauthorization to HTTP 422. Preserve reconnect instructions and reporting of unexpected provider failures.

All 363 focused tests and server typecheck pass. Required CI and review checks passed; Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:50:01 -07:00
Devin FoleyandPaperclip a7d3b17a97 Explain disabled Slack MCP app access during discovery
Recognize Slack's exact disabled-app response and return actionable setup
instructions from catalog and health routes. Bound response parsing and keep
unknown upstream errors reportable without exposing provider settings links.

Verified 361 focused tests, server typecheck, and authenticated discovery.
The three route regressions fail before this change and pass afterward.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:25:06 -07:00
Devin FoleyandPaperclip ded156a904 Keep interaction continuations scoped to the target task
Separate an interaction's producer from explicit resume history. Do not
import another task's comments or results. Filter newly captured foreign
origins and recover older inherited origins only when the saved producer
context and comment row prove their source. Keep missing context and
company boundaries fail-closed.

Verified 46 continuation tests, 201 related recovery tests, and server
typecheck. Added 20 database regressions for provenance and scope guards.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-21 20:06:05 -07:00
Devin FoleyandPaperclip 878734a061 fix(tools): treat OAuth sign-in challenges as client errors (#13786)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connected apps can require OAuth sign-in before they list their
tools.
> - Remote discovery recognizes this condition as `oauth_challenge`.
> - Discovery and catalog refresh currently return it as HTTP 502.
> - The server error handler reports that response as a crash.
> - This pull request returns HTTP 422 for the explicit sign-in
challenge.
> - The operator keeps the sign-in instructions, while unexpected
upstream failures remain reportable.

## Linked Issues or Issue Description

**What happened?**

Connecting a remote MCP app that answers with a recognized OAuth
challenge returns HTTP 502. Reading an empty catalog or explicitly
refreshing it does the same. The server error handler then sends the
expected sign-in condition to error monitoring.

**Expected behavior**

A known sign-in requirement returns HTTP 422 with the existing
`oauth_challenge` code, message, and setup/reconnect links. An
unexplained upstream HTTP 400 or an unavailable upstream service still
returns 502 and reaches error monitoring.

**Steps to reproduce**

1. Configure a remote MCP app that returns HTTP 401 with a Bearer
challenge.
2. Connect the app, read its empty catalog, or request a catalog
refresh.
3. Observe HTTP 502 and a server error report before this change.

**Paperclip version or commit**

Reproduced on `6de50ba594b15efaa3eae6ed869cd39b3a436456`.

**Deployment mode**

Server with remote MCP connections. The regression coverage uses local
PostgreSQL and mocked upstream HTTP responses.

Related: #9750 addresses MCP initialization and session recovery. It
does not change the classification of this recognized sign-in condition.
Targeted searches found no duplicate classification PR.

## What Changed

- Return 422 for `oauth_challenge` from discovery and from catalog
health-error normalization.
- Preserve the existing structured error and remediation links.
- Test automatic empty-catalog reads and explicit refreshes. Verify that
OAuth challenges produce no Sentry capture and that upstream 400/503
failures still do.
- Update the direct-connect and blocked-redirect expectations and
document the monitoring behavior.

## Verification

- Before the fix, three sign-in route regressions fail with 502 instead
of 422; all four upstream-error controls pass.
- After the fix, all 339 tool-access and error-handler tests pass,
including authorization and redirect protections.
- These suites ran against disposable Homebrew PostgreSQL 16.14 through
the existing test-constructor seam. The temporary setup and config
remain outside the repository. CI uses the ordinary embedded PostgreSQL
setup.
- Direct server `tsc --noEmit` passes.
- Full build and recursive typecheck were attempted; the Runner Rust
step cannot run because `cargo` is absent on this machine.
- Full `pnpm test:run`: 8,211 passed, 14 failed, 4,759 skipped. The 36
failed files match the existing embedded PostgreSQL startup/cleanup and
macOS runtime-cache `EACCES` limitations. The changed database-backed
service suite passed separately with local PostgreSQL.
- Greptile: 5/5 with no unresolved comments. Seven CI workers received a
simultaneous shutdown signal; the failed jobs are being retried through
the normal workflow. Other completed checks passed.

## Risks

Low risk. Clients now receive 422 instead of 502 for the explicit
`oauth_challenge` condition. The code, message, and remediation links
remain available. No permissions, credential handling, OAuth discovery
rules, retry policy, or schema change. Other upstream failures retain
their existing behavior.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (339 service and
error-handler tests; full workspace limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 00:34:45 +00:00
Devin FoleyandPaperclip 6de50ba594 fix(sentry): carry the deployment environment to the browser (#13784)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can enable Sentry for the server and the signed-in
browser.
> - The server SDK reads `SENTRY_ENVIRONMENT` from the process
environment.
> - The browser receives its DSN through the session response, but
receives no environment.
> - A browser in staging therefore reports errors under the SDK's
production default.
> - This pull request passes the configured environment through the
existing session and monitoring gate.
> - Browser errors then identify the deployment environment while
preserving the existing privacy settings.

## Linked Issues or Issue Description

**What happened?**

With `SENTRY_ENVIRONMENT=staging`, browser exceptions are tagged
`production`. This can send errors to the wrong environment's alerts and
makes deployment follow-up unreliable.

**Expected behavior**

The browser uses the server's configured Sentry environment. A reused
image works in either staging or production. A signed-out browser still
sends no events.

**Steps to reproduce**

1. Configure a frontend Sentry DSN and `SENTRY_ENVIRONMENT=staging`.
2. Sign in and capture a browser exception.
3. Inspect the event environment. Before this change, it is
`production`.

**Paperclip version or commit**

Reproduced on `a3749aac4680a901fa0fe1cc898907887abc9908` with the real
browser SDK and a local test transport.

**Deployment mode**

Authenticated server and browser with optional Sentry monitoring
enabled.

No duplicate environment-attribution issue or pull request was found in
the targeted GitHub search.

## What Changed

- Add `sentryEnvironment` to the authenticated session response and
shared schema. The optional field supports a newer browser reading an
older server response.
- Pass the environment to the browser SDK. An environment change
restarts the client through its existing serialized lifecycle.
- Cover environment attribution with a real SDK event, session
authorization, unchanged-session refetches, environment changes, and
legacy responses.
- Document configuration and compatibility. Keep the loaded bundle's
release identity and existing privacy filters.

## Verification

- The regression test emits `production` for a requested staging
environment before the fix.
- Focused route, schema, browser lifecycle and real-SDK tests: 69 pass.
- UI and shared-package typechecks, direct server `tsc --noEmit`, and
token gates pass.
- Full `pnpm build` and `pnpm -r typecheck` were attempted. Both stop at
the Runner Rust step because `cargo` is absent on this machine.
- Complete UI suite: 6,540 tests pass in 626 files.
- Full `pnpm test:run`: 8,210 passed, 14 failed, 4,753 skipped; 36 files
fail due to embedded PostgreSQL startup/cleanup and macOS runtime-cache
`EACCES`. These match the existing local baseline; none touch the
changed behavior.
- Greptile: 5/5, no unresolved review threads. Linux CI has passed
Build, Typecheck + Release Registry, and the completed test jobs so far.
Remaining jobs are running or queued: the AWS runner provisioner is
retrying EC2 CreateFleet `InternalError` responses. Full results will be
recorded before merge.

## Risks

Low risk. This adds one optional session field and changes Sentry
attribution only. No migration or new monitoring opt-in is introduced.
Missing settings keep the browser SDK default. Agent and unauthenticated
requests still receive 401 without monitoring settings. Existing loaded
browser bundles keep their old behavior until refreshed.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, code
editing, and test execution. The session does not expose an exact model
snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused and full UI suites
pass; full-root environment failures documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 23:40:46 +00:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 57fd8b70d2 feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack connections let a team talk to those agents in Slack.
> - Agents now have a saved avatar, but Slack setup did not offer that
image.
> - A matching avatar helps a team recognize its agent.
> - This pull request adds an optional avatar step and a download in
connector Settings.
> - Users download a PNG and upload it directly in Slack with clear
instructions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Slack connector onboarding and its Settings page.

**Current behavior**
Setup does not offer the assigned agent's avatar or explain how to
upload it in Slack.

**Proposed behavior**
After Slack connection verification, users can download a 512 × 512 PNG
of their agent's saved avatar. They can upload it in Slack, confirm, or
skip. Settings keeps the download and upload instructions available
after onboarding.

**Reason and benefit**
The same avatar helps people recognize the agent across Paperclip and
Slack. Users who skip the optional step can return to it in Settings.

**Breaking changes**
None. No schema, authentication, Slack scope, or provider API change.
Completed connections keep their existing completion state. Searched
existing Slack avatar and Cliptoon PRs; no matching implementation was
found.

## What Changed

- Add an optional avatar step before personal Slack account linking.
Keep the numbered sidebar and shared footer.
- Resolve the selected agent's saved appearance for the preview and PNG
download.
- Add the same download and expandable upload instructions to connector
Settings.
- Remember uploaded or skipped per company and endpoint in browser
storage. Treat uploaded as user confirmation, not provider verification.
- Reject failed or non-PNG download responses and allow retry.
- Reuse the production avatar components in onboarding and Settings
stories.
- Test wizard progression, resume, Settings, download recovery, storage
isolation, and terminated assigned agents.
- Exercise real PNG downloads in the Slack browser flow and keep default
app names consistent with app creation.
- Fetch the assigned agent directly so its saved avatar remains
available after termination.

## Verification

- Focused chat suites: 48 passed; the two affected suites passed again
after the final naming fix (32 tests).
- Slack browser E2E passed through setup, avatar download, account
linking, and Settings download. PNG signature and 512 × 512 dimensions
verified.
- UI token gates passed.
- Browser: downloaded the real 512 × 512 PNG; checked confirmation,
return, mobile layout, and Settings instructions.
- Full workspace typecheck, application build, and production Storybook
build passed.
- All latest-head CI checks passed (54 passed, 2 skipped), including all
browser, chat, general, and serialized test groups. The unrelated Sentry
test failed once and passed on the single CI rerun; its suite also
passed locally.
- Local full-suite attempt encountered a rapid Slack callback ordering
failure under concurrent build load; that test passed in isolation, and
all three chat shards passed in CI. The remaining local run was not used
as the merge gate.
- Review the Connections / Slack / Add avatar and Avatar in Settings
stories.

## Risks

- Slack upload is manual. Confirmation does not claim to verify the
Slack icon.
- Optional step progress is browser-local. Clearing storage or changing
browsers can show it again. Setup still works when storage is
unavailable.
- The existing avatar API remains the image source. Download failures
show a retry message.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 09:38:48 -05:00
Devin FoleyandPaperclip c65fc9e3c8 fix: recover authentication and browser connection failures (#13724)
fix: recover authentication and browser connection failures

Include connect timeouts in the bounded retry policy for idempotent actor
synchronization. Handle WebSocket constructor failures through existing
reconnect paths and preserve HTTP polling while realtime is unavailable.
Refresh visible company queries until the socket recovers and clear all
fallback timers on hiding or unmount.

Verify 172 focused tests, server/UI typechecks, UI build, and design token
gates. Full workspace build/typecheck require the unavailable Rust toolchain;
the full test run is tracked separately.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 12:43:25 -07:00
DottaandPaperclip 9f30eb10dd fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path

> - Paperclip manages AI agents and keeps their work attached to tasks.
> - Chat connectors carry user messages and agent replies between a
provider and those tasks.
> - Each extra startup and context reset delays a reply.
> - Managed account metadata was lost during adapter decoding, so
compatible follow-ups started fresh.
> - This branch fixes the reset and measures the remaining preparation,
execution, and delivery costs.
> - The changes must preserve account isolation, authorization, durable
output, and recovery ownership.

## Linked Issues or Issue Description

Refs #13699. The related service lifecycle work in #13410 and #13408 is
separate; this branch focuses on task-bound chat response latency.

**What happened?**

Managed AI follow-ups started new provider sessions even after their
configuration fingerprint stayed stable. The Codex codec removes unknown
fields. The resume check then read the removed credential identity and
treated it as a credential change.

**Expected behavior**

Compatible follow-ups resume the correct provider session. Changes to
credentials, responsible users, permissions, or task configuration
retain their reset behavior.

**Steps to reproduce**

1. Use a Slack connector with a managed AI connection.
2. Send a message, then send a same-thread follow-up.
3. Inspect the configuration reset reason and the provider session
identity.

**Paperclip version or commit**

Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`.

**Deployment mode**

Cloud staging with a native Codex runner.

## What Changed

- Read saved credential identity before adapter decoding discards it.
- Remove the internal credential identity from adapter-facing session
params.
- Test the real Codex codec and missing, changed, or unmanaged identity
cases.
- Preserve configured warm Codex runners and flush refreshed credentials
after every turn.
- Fence detached or closing session handles from successor credential
ownership.
- Stage current Codex launch credentials after restoring durable session
history, uploading launch assets only once.
- Reuse a runner binary already in the retained sandbox only when its
SHA-256 matches the controller-owned artifact; still verify required
capabilities before launch.
- Lock the task before the run when saving results, preventing deadlocks
with task updates.
- Scope reusable projectless sandboxes to the company, environment,
task, agent, and runtime configuration; verify Daytona sentinels for
that scope.
- Admit a new authorized chat message after a fully committed failed run
and verified process cleanup.
- Send compact deltas for verified plain-text Slack continuations. Match
the actual prior run and current comment identity/body; exclude edited
historical comments and prior agent output, preserve genuine brief edits
and the full bootstrap fallback.
- Keep attachments, omitted input, questions, approvals, recovery, and
other providers on their existing framing.
- Document managed session compatibility, credential lifecycle, and
compact continuation boundaries.

## Verification

- Workspace/session coverage: 156 tests passed.
- Native session and credential ownership coverage: 390 tests passed,
including exact artifact reuse, mismatches, failed probes, timeouts, and
explicit artifact overrides.
- Explicit continuation and durable chat authorization coverage: 172
tests passed.
- Session resume and launch preparation coverage: 416 tests passed.
- Result persistence coverage: 15 tests passed. The new concurrency test
reproduced a PostgreSQL deadlock before the lock-order fix.
- Environment lifecycle coverage: 92 tests passed, including projectless
reuse and task/agent isolation at both selection and atomic handoff.
- Daytona plugin coverage: 237 tests passed; 6 gated tests skipped.
Standalone plugin build passed.
- Compact Slack continuation and native resume coverage: 69 tests
passed, including full-bootstrap retention, matching message
authors/bodies, current-delivery selection, rejection of duplicate
identities and historical comments, brief edits, and attachment/recovery
fallbacks.
- Final frozen-head `pnpm test:run` on repository-supported Node 26: 668
suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped.
The sole failure was a local `socket hang up` in
`issue-recovery-actions.test.ts`, not an authorization assertion
mismatch. All 57 tests in that suite passed three fresh reruns, and the
suite passed latest-head CI. The full local invocation is therefore not
claimed green.
- An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup
failure; that 11-test catalog suite passes on Node 26 and in CI. No test
behavior or timeout was relaxed.
- Full local typecheck and build passed. Latest-head CI is green;
Greptile is 5/5 with no unresolved review threads.
- Two real Slack baseline replies took 25.1 and 24.6 seconds
(24.9-second mean). Three same-thread signed probes on this head took
23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample
and a modest wall-clock improvement, not a large or statistically
established speedup.
- In that same thread, uncached provider input fell from 8,514 tokens
before compact input to 694–765 tokens afterward. The current delivery
uses a 362-character delta; the full 19–21k-character bootstrap remains
available for failed resume. Verified runner artifact preparation fell
from about 1.2 seconds to 0.6 seconds.
- A fresh thread created a separate task, sandbox, and provider session
with full bootstrap (24.4 seconds). Its follow-up reused its own
sandbox/session and compact input (28.3 seconds, including 16 seconds of
model execution). Model variability and process startup remain
substantial.
- A signed duplicate webhook produced exactly one user comment, one
successful run, and one final Slack reply. Slack's API independently
confirmed the actual replies and a public task URL without an internal
or pool hostname.
- Earlier signed probes verified recovery after a failed run and reuse
across a server deployment. The final idle test observed Daytona report
the sandbox as stopped, then delivered a new reply in 19.9 seconds using
the same sandbox/provider-session identity and compact input. Slack’s
API confirmed that reply.
- Live probes use signed synthetic inbound webhooks and real outbound
Slack delivery, read back through Slack’s API. The final browser recheck
found the Mac locked and the Slack tab blocked by another extension, so
this is not claimed as full UI E2E proof.
- This is a review branch. Do not merge until the maintainer reviews it.

## Risks

- Incorrect session reuse could mix account or task context. Missing or
changed identities continue to reset, and existing authorization checks
remain in place.
- Warm mode remains opt-in. Remote warm mode requires a reusable sandbox
lease. Retained processes keep credentials until they close, so idle
expiry and ownership fences are required.
- A fresh user message may continue after a committed provider failure.
Approval, current authorization, process termination, and prior-result
checks remain required.
- Projectless sandbox reuse is task- and agent-scoped. Missing or
mismatched ownership cannot replace an existing lease; existing
workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox
per task/agent, so provider auto-stop and deletion policies still
determine idle compute and storage costs. Fleet defaults are unchanged.
- Compact prompts apply only after proven resume and a matching
prior-run delta. Missing or specialized context falls back to full
input; fresh sessions always receive the full bootstrap.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-20 13:12:44 -05:00
Devin FoleyandPaperclip 600e552d7b fix: attribute Sentry errors to the loaded source release (#13719)
Attribute optional server and browser Sentry events to their source build.
Use validated build commits for Docker and source/npm artifacts, preserve
explicit server release overrides, and keep cached browser bundles tied
to the commit they loaded.

Verify 127 focused tests, server/UI typechecks, Docker and source build
stamps, all 53 CI checks, and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-20 08:08:33 -07:00
DottaandPaperclip 2a99de80ec fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.

## Linked Issues or Issue Description

Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.

**What happened?**

An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.

**Expected behavior**

Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.

**Steps to reproduce**

1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.

**Paperclip version or commit**

Reproduced on the source at b193077582.

**Deployment mode**

Cloud staging with a native runner.

## What Changed

- Supply an exact public task URL in fresh and resumed external-chat
context.
- Return a nullable `url` on task records in `get_task_context` and
`search_tasks`, including child tasks.
- Resolve the current claimed Cloud origin at use time. Preserve the
existing safe-URL filter and omit unsafe links.
- Normalize only the exact credential-home paths created by the managed
AI runtime when comparing session configuration. Do not change the
actual execution environment.
- Add regression tests and update the connector playbook.

## Verification

- 255 focused tests passed for public URLs, prompt context, AI
connections, and session configuration.
- 19 native task-tool tests passed, including public URL lookup, origin
changes, and unsafe URL rejection.
- Full workspace typecheck and build passed locally. The full local test
run is still in progress; all CI test lanes passed at the final head.
- All 54 CI checks passed, with two optional jobs skipped. Greptile
scored 5/5 with no review comments.
- After merge, deploy the exact merged commit to the authorized staging
stack. Check a fresh Slack mention, a same-thread reply, the returned
task link, and the measured session reuse and response time.

## Risks

- Existing saved fingerprints can reset once after this change. Future
runs should reuse a compatible session.
- The fingerprint change applies to managed AI connections. Tests
preserve resets for real account, credential, model, permission, and
environment changes.
- A public task URL still requires normal Paperclip access. This change
grants no access and keeps internal URL filtering.
- This change does not remove all runner startup or checkpoint latency.
Staging measurements are still required.
- No schema or migration changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, tool use, and test
execution. The exact serving model ID and context-window size were not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 20:40:05 -05:00
DottaandPaperclip b193077582 fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack connections turn messages into governed agent runs and return
replies to Slack.
> - Remote runs must replace an incompatible sandbox runner with the
controller's packaged binary.
> - The fallback used a package-relative path that does not match the
vendored server layout.
> - Slack callback health also compared the internal proxy address with
the public callback address.
> - Master now contains the runner fallback fix; this pull request fixes
callback health and extends missing-runner regression coverage.

## Linked Issues or Issue Description

Refs #13677, #13680, #13691, and #13686.

**What happened?**

A packaged controller stopped a Slack-triggered remote run with
`runner_remote_artifact_unavailable` when the sandbox runner needed
replacement. Working Slack callbacks also showed a stale URL warning
behind the Cloud gateway.

**Expected behavior**

The controller stages its packaged runner when needed. Callback health
uses the observed public address and still detects real address changes.

**Steps to reproduce**

1. Run a packaged server with a sandbox that has an older runner or no
runner.
2. Send a Slack mention to an agent that uses that sandbox.
3. Route signed Slack callbacks through a claimed Cloud gateway that
rewrites the upstream host.
4. Check the run and the Slack callback health panel.

**Paperclip version or commit**

Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The
fix branch also includes #13691.

**Deployment mode**

Packaged server with a Cloud gateway and a Daytona sandbox.

## What Changed

- Extend the controller-owned runner fallback tests with a missing
sandbox binary case; retain the fix now merged in #13686.
- Prefer dedicated gateway diagnostic headers that survive provider
rewrites of standard forwarded headers. Use validated host hints only
for callback-health evidence on claimed Cloud instances after provider
acceptance. Preserve request bodies, routing, authentication, and
configured callback URLs.
- Cover current, stale, and missing sandbox runners, all Slack callback
surfaces, rejected callbacks, malformed proxy hints, real host and port
changes, and the existing self-hosted behavior.
- Document the artifact lookup and callback-health boundaries.

## Verification

- Native session executor and binary resolver suites: 379 tests passed
after merging current master (`45c99a0d0`).
- Targeted callback integration suite with disposable PostgreSQL: 4
tests passed before rebase.
- Full `pnpm -r typecheck` and `pnpm build` passed after merging current
master. The full local suite passed 12,668 tests; one suite failed to
start its disposable PostgreSQL. Rerunning that suite alone passed all
31 tests.
- The built server resolver selected the executable under
`server/dist/vendor/paperclip-runner/bin/`.
- Final callback regression: all 4 targeted integration tests pass,
covering provider header rewrites, default ports, uppercase/trailing-dot
hosts, and ignored self-hosted hints.
- Live staging proof of the runner fix: a previously failed Slack thread
recovered, a new mention received its requested response, and the
account-connect command succeeded. Unsigned callbacks returned 401.
- Live browser and Slack acceptance passed: generated callback URLs,
account linking, a new mention, an interactive question and answer, and
all three callback-health indicators. The connector was activated
through the onboarding UI.
- All 54 latest-head checks pass, two optional checks are skipped, and
Greptile is 5/5 with no unresolved threads.

## Risks

- Runner fallback must select a binary for the remote platform. This
preserves explicit remote artifact overrides and the existing capability
checks.
- Proxy headers are not identity proof. They are used only for
diagnostics on claimed Cloud instances after the provider accepts the
request. Self-hosted instances ignore them. Wrong public hosts and ports
still warn.
- No database migration or new public API contract.

## Model Used

OpenAI GPT-6 via Codex. Used reasoning, repository tools, code
execution, and browser/native-app testing. Exact model variant and
context window were not exposed by the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 17:32:51 -05:00
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00
DottaandPaperclip 04546c82d5 fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner can execute a task inside a Daytona sandbox.
> - The sandbox can keep running when the Paperclip controller restarts.
> - Recovery treated sandbox process IDs as local process IDs and
selected the wrong recovery path.
> - Live verification also found races between startup, shutdown, and
queued task cleanup.
> - This pull request verifies the existing remote owner and orders
those transitions.
> - Users can continue the same task and provider session after a
controller restart.

## Linked Issues or Issue Description

**What happened?**

The Daytona `recover-controller` cases failed with
`runner_state_identity_mismatch`. Remote process IDs can be absent on
the controller or collide with unrelated local processes. Recovery then
looked for remote state in the local runner directory. Later turns could
also start before the previous executor released its sandbox resources.

**Expected behavior**

Reconnect to the original sandbox and authenticated runner. Preserve the
task, provider session, and queued comments. Reject a replacement
sandbox or mismatched identity. Do not start another provider during
reattachment.

**Steps to reproduce**

Run the `everyday-workflows` `recover-controller` case for
`runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a
Python tool, requests a revision, restarts the controller during
execution, and queues another revision. It then downloads and tests the
final ZIP.

Related: #13682 is the preceding operational fix. #13291 addresses
legacy sandbox conversation recovery, a different execution path. #13666
includes broader run-capacity work; this change guards cleanup of an
existing native task executor.

## What Changed

- Add remote runner recovery without interpreting sandbox PIDs on the
controller.
- Verify the original provider lease, remote workspace, durable state,
process marker, and authenticated PRP authority before adoption.
- Compare the process marker with live Linux boot identity and start
ticks to reject PID reuse. Read virtual proc files through the
guaranteed Node runtime; unavailable proof blocks adoption without
blocking a fresh launch.
- Make the E2E supervisor own the actual server process so forced
restart cannot leave a late database closer behind.
- Scope the chat delivery lease test to its own fixture instead of
draining other tests’ pending deliveries.
- Preserve provider-attempt counts and recorded evidence during
reattachment.
- Serialize an idle-session checkpoint with admission of the next native
turn.
- Wait for an in-progress startup to acknowledge restart detachment.
Fail after a bounded deadline if it cannot.
- Keep a queued comment waiting until the previous native task executor
releases its resources. Allow unrelated tasks to continue.
- Update the Daytona image's resolved lock digest to match current
dependency manifests.
- Add classifier, ownership, process, startup, checkpoint, and
queued-admission regression tests. Document recovery behavior.

## Verification

- 415 focused tests passed across native execution, restart recovery,
workspace synchronization, queued admission, and real-process restart
tests. The final Node-based fingerprint change passed all 375
native-session tests.
- Runner harness unit tests: 394 passed. Chat integration shard 2: 335
passed after fixture isolation.
- The exact fingerprint command succeeded twice in a disposable Daytona
sandbox and returned the same identity; the sandbox was deleted.
- 11 real-process restart integration tests passed, including absent and
colliding remote PIDs.
- Repository typecheck and final build passed. Broad local checks found
machine-dependent database startup and timing failures; focused retries
passed. The final-revision PR pipeline is green. One unrelated browser
shard hit a five-second blank-page timeout on the first run and passed
its targeted retry.
- Final-revision local headed browser E2E:
`everyday-workflows.runner-acpx-claude.daytona.recover-controller`
passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser
inspection confirmed Done, all three ZIPs, and delivery of the queued
follow-up. All three runs succeeded using the same provider session. The
harness downloaded and independently tested the final artifact.
- Final-revision Daytona campaign:
https://github.com/paperclipai/paperclip/actions/runs/35463999611 —
**Codex passed first attempt (4.8 minutes); ACPX Claude passed first
attempt (6.1 minutes)**. Campaign aggregation/publication is finishing;
both test jobs succeeded.
- Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**,
no open findings.
- Staging browser verification is pending selection of a disposable
staging instance and removal of a Chrome extension UI block.

## Risks

- Recovery now depends on the original sandbox remaining available. A
replacement or mismatched identity still blocks adoption.
- Shutdown waits up to 30 seconds for a native startup to reach a safe
detach point. An unfinished startup returns a clear failure instead of a
false detach receipt.
- Queued native work on the same task waits for cleanup. Unrelated tasks
remain eligible.
- The image digest update rebuilds the Daytona runtime image. No
database migration or public API change is included.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code execution, and
browser tools. The runtime does not expose the exact deployed model ID
or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:13:34 -05:00
Nicky LeachandClaude Opus 5 aeef493f4a chore(db): keep only the newest 5 drizzle snapshots and stop shipping them (#13687)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - `@paperclipai/db` owns the Drizzle schema and the migration history
> - Drizzle writes a full copy of the schema as a snapshot for each
generated migration. Each snapshot is now about 1.3 MB.
> - The `meta/` folder is about 103 MB. That is about half of each
checkout and each worktree. The build also copies it into `dist`, so the
published `@paperclipai/db` package is 112.8 MB unpacked.
> - `drizzle-kit generate` reads only the newest snapshot. The runtime
migrator reads only the `.sql` files and `_journal.json`.
> - This pull request keeps the newest 5 snapshots and removes snapshots
from `dist`.
> - The benefit is a checkout that is about 100 MB smaller, and a
published package that is about 1.5 MB instead of 113 MB.

## Linked Issues or Issue Description

Refs #11240, #11254, #12333 (earlier snapshot work: diff collapse,
binary diffs, drift repair)

**What existing behavior does this improve?**
The size of the Drizzle migration snapshots in the repository and in the
published `@paperclipai/db` package.

**Subsystem affected**
`packages/db`: migrations and the build.

**Current behavior**
`packages/db/src/migrations/meta/` holds 141 snapshots (102.8 MB). The
size grows faster than the number of migrations, because each snapshot
is a full copy of the schema. `build` runs `cp -r src/migrations
dist/migrations`. `@paperclipai/db@2026.916.0` contains 147 snapshot
files. It is 112.8 MB unpacked and 5.3 MB as a tarball.

**Proposed behavior**
Keep the newest 5 snapshots. `generate` deletes older snapshots after it
runs. `dist` gets only `*.sql` and `meta/_journal.json`.

**Reason and benefit**
- In drizzle-kit 0.31.10, `generate` sorts `meta/*` and diffs against
the last snapshot only (`bin.cjs`, `preparePrevSnapshot`). The [generate
docs](https://orm.drizzle.team/docs/drizzle-kit-generate) also say it
compares against "the most recent" snapshot.
- Gaps in the snapshot history already work. 138 of the 280 migrations
never had a snapshot, because they were written by hand before
`doc/DATABASE.md` required `generate`.
- We keep 5 snapshots instead of 1. This lets a developer undo the
latest generated migration, and it keeps the `prevId` chain for recent
branches.
- Git keeps the history cheaply. The 631 snapshot versions use only 2.2
MB of the pack, because git stores each version as a delta of the
previous one. The cost is in the checked-out files, not the clone
download. Therefore this change does not use git-lfs and does not
rewrite history. Old snapshots stay available with `git show
<rev>:<path>`.

## What Changed

- `packages/db/package.json`: a new `prune:snapshots` script keeps the
newest 5 `*_snapshot.json` files. It is `ls | sort -r | tail -n +6 |
xargs rm -f`, which works with the BSD tools on macOS and the GNU tools
on Linux. `generate` runs this script after `drizzle-kit generate`.
- `packages/db/package.json`: `build` copies only `src/migrations/*.sql`
and `meta/_journal.json` into `dist/migrations`.
- Deleted 137 older snapshots. `0277`–`0281` remain.
- `chat-identity-migration-reconciliation.test.ts`: removed the walk
over snapshots `0254`–`0268`. Those files do not change after merge, and
the walk would fail after pruning. The journal-order assertions in the
same test remain. `migration-snapshot-drift.test.ts` still makes sure
that the newest snapshot matches the schema.
- `doc/DATABASE.md`: documented the retention rule.
- The snapshots were already marked `linguist-generated=true -diff
-merge` by `packages/db/.gitattributes` (#11240). No change there. `git
check-attr` confirms it.

## Verification

- `pnpm --filter @paperclipai/db exec vitest run`: 43 files, 160 tests
pass.
- `pnpm --filter @paperclipai/db build`: `dist/migrations` contains 280
`.sql` files and `meta/_journal.json`. It is 1.5 MB, compared with about
105 MB before.
- `prune:snapshots` was run on macOS (BSD) and in `debian:stable-slim`
(GNU findutils 4.10). With 7 fixture snapshots, it keeps the newest 5.
When 5 or fewer are present, it deletes nothing and exits 0 on both.

## Risks

- Low risk. Runtime migration does not read snapshots. The only commands
that read older snapshots are `drizzle-kit check` and `drizzle-kit
drop`. No script or CI job calls them, and they still have the newest 5
snapshots.
- A branch that is open now can still add its own snapshot. If the
branch conflicts, the rule is the same as today: renumber the migration
and run `generate` again.
- When we upgrade to drizzle-kit v1 (folder per migration), check
whether its new cross-branch "commutativity" checks need a longer
snapshot history.

## Model Used

- Claude Opus 5 (`claude-opus-5`) in Claude Code, with tool use (shell,
file edits, web fetch).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 10:40:38 -07:00
DottaandPaperclip da257c3069 Warn when routine webhook URLs may not be publicly reachable (#13684)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Routines can start that work when another app sends a webhook.
> - Local and private URLs often cannot receive events from public
services.
> - HTTPS alone does not make a Tailscale address public.
> - This pull request explains these limits during setup and editing.
> - Users can still finish setup for senders on their own network.

## Linked Issues or Issue Description

Refs #13637. Webhook setup needs a clear warning when the generated URL
appears local, private, or unencrypted. The warning must explain how to
make the endpoint reachable without blocking private-network use.

## What Changed

- Add a shared warning banner to the Connect, Check connection, and Edit
webhook views.
- Distinguish localhost, private network addresses and domains, HTTP,
and Tailscale hostnames.
- Explain the difference between Tailscale Serve and Funnel. Link to the
Paperclip HTTPS guide.
- Add five full-page Storybook examples, design guide examples, and
documentation.
- Add URL classification tests and a regression test that finishes setup
despite the warning.

## Verification

- Passed 41 focused URL and trigger-flow tests.
- Passed workspace typecheck, workspace build, token gates, and
Storybook build.
- Browser-tested the Tailscale story through Check connection, Finish
setup, and Edit webhook. The warning stays visible and does not block
setup.
- Open Product / Routines / Webhooks stories 11–15 to review the warning
states.
- All 54 PR checks passed, including the full test matrix and eight
browser shards; two optional Storybook jobs were skipped by workflow
policy.
- The server supervisor readiness test timed out once in CI, then passed
on rerun and locally (6 tests).
- The duplicate local full-suite run was stopped after the complete CI
matrix passed.
- Greptile: 5/5 on the current commit, with no unresolved review
comments.

## Risks

- URL checks are hints. They do not test DNS, firewall rules, or actual
reachability.
- A Tailscale hostname can serve either private Serve traffic or public
Funnel traffic. The warning explains this uncertainty and permits both.
- No API, schema, authentication, or webhook delivery behavior changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 12:24:06 -05:00
Devin FoleyandPaperclip b70641f23f feat(plugins): support image catalogs and persistent application overlays (#13646)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins extend the application without adding each integration to
Core.
> - A downstream image needs a way to supply prebuilt plugins.
> - Some plugin UI must stay mounted as users move between pages.
> - This change adds an image catalog and a persistent application slot.
> - Operators can upgrade or remove these plugins through their image
and configuration.

## Linked Issues or Issue Description

**Subsystem affected**

Plugin packaging, activation and application UI.

**Problem or motivation**

The built-in plugin catalog is fixed in Core source. Downstream images
cannot add entries through an explicit catalog. Existing page slots also
cannot preserve a small application overlay across route changes.

**Proposed solution**

Read a bounded catalog of prebuilt plugins from the image. Verify its
files before importing manifests. Use the existing managed selection and
plugin lifecycle. Add an `appShellOverlay` slot with account and company
cleanup.

**Alternatives considered**

A downstream fork adds merge work. Script injection provides no plugin
lifecycle. A separate runtime download system adds a second distribution
channel.

**Roadmap alignment**

This extends the existing plugin system. Related PR #9006 covers runtime
install replication; this change covers immutable image contents. PR
#12555 covers CLI scaffolding. Neither provides this catalog or
application slot. The maintainer requested this work directly.

## What Changed

- Validate catalog identities, confined paths, package versions and
bundle hashes before importing code.
- Apply image selection to persisted plugin installs, including removal
and rollback. Adopt the verified image path from existing npm/local
installs and bind runtime worker/UI entrypoints to verified package
declarations.
- Mount application overlays in both UI shells. Preserve route state and
clear it on account, company and onboarding changes.
- Restrict service-worker offline storage/fallback to hashed public
assets in a separate cache namespace; exclude application HTML and
extension/API data, including after worker restart.
- Document the packaging contract, trust model and rollback
requirements.

## Verification

- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Affected server/UI typechecks and builds, plus token
gates, passed again after rebasing onto current master; the 124 focused
tests also passed after rebase.
- Latest focused verification: 124 tests in nine files passed for
catalog/reconciliation/loader, overlay lifecycle, Layout and
service-worker policy. The broader UI/shared/SDK run passed 7,204 tests
in 690 files with canonical `TMPDIR`.
- Real disposable Core/PostgreSQL: catalog install, selection removal,
0.1.0→0.1.1→0.1.0, same-version npm/legacy-path adoption, and
preservation of disabled status passed. Added permissions entered
`upgrade_pending`, withheld UI across restart, and activated only after
explicit operator enable.
- Real Chromium: desktop/mobile layout, route draft retention and
Escape/focus passed with mocked extension responses. A persistent
browser restart retained public hashed-asset offline fallback while
refusing seeded legacy/current private entries and legacy HTML.
- Full `pnpm test:run`: 12,539 passed; 17 failed across six existing
files, stopping later phases. macOS read-only directory renames fail in
runtime-skill-cache and company-skills-service; email tests require an
absent local AgentMail fixture. Native runner/comment-redaction passed
in isolation after temporary Rust setup; agent-conversations also passed
in isolation. No unrelated source was changed to hide failures.
- After rebase, two unchanged chat timing tests failed in CI and passed
locally in isolation. Their CI shard passed on its single retry. All
other current-head CI jobs passed on the initial run; review is 5/5 with
no unresolved threads.
- No live deployment or external plugin service was used.

## Risks

- Plugins are trusted code. The catalog detects packaging errors; it
does not authenticate an untrusted image builder.
- Invalid catalogs fail startup. Images must contain the catalog and
bundles together, with stable directories.
- A host older than this contract lacks the activation guard. Disable
added plugins and remove their configuration keys before reverting to
it.
- Offline navigation now returns 503 instead of replaying cached
application HTML. Only public build assets have offline fallback.
- Rolling back an unapproved permission change retains the approval
gate; review the current manifest and explicitly enable it. A reduced
permission set cannot establish prior approval or prior enabled status.
- Plugin data migrations need their own rollback policy. This change
retains installed records and does not reverse migrations.

## Model Used

- OpenAI GPT-6 (Codex), model ID `gpt-6`, with repository inspection,
code execution and browser verification. The runtime does not expose an
exact context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites; broad
macOS server-run exceptions are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (fresh run on 488b3754ae; chat
shard passed its single retry)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(fresh review on 488b3754ae)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:18:10 -07:00
DottaandPaperclip f589660ec0 feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Routines turn scheduled work and external events into tasks for an
assigned agent.
> - Webhook setup was disabled, and actor authentication rejected valid
webhook bearer keys.
> - Operators need to connect and test a sending app before events can
start work.
> - This pull request adds a guided setup with durable connection tests
that cannot dispatch a task.
> - It keeps trigger management, execution tasks, and activity within
the routine.
> - The benefit is a webhook that can be configured, verified, and
operated from one place.

## Linked Issues or Issue Description

Fixes #11937.

Related: #13216 adds provider-specific Sentry support. This PR addresses
general routine setup and ingress. #6841 addresses legacy secret
bindings; this PR retains the existing secret service.

**Current behavior**

Webhook creation is disabled. Bearer deliveries can fail in agent
authentication before the routine checks its key. Setup has no safe
connection test. Runs and Activity send the operator away from the
routine.

**Proposed behavior**

Choose a schedule or a webhook. Follow the setup steps, copy credentials
or complete agent instructions, and test delivery without creating work.
Finish setup to allow future events to start tasks. Edit or remove
compact trigger cards, undo removal, and inspect tasks and activity
inside the routine.

**Reason and benefit**

An operator can verify credentials and delivery before enabling
automatic work. Durable setup state survives refreshes and restarts.
Retry receipts prevent an old test event from starting work after
activation.

## What Changed

- Add a production trigger wizard using reusable Slack setup navigation
and footer components.
- Add schedule and webhook choices, one-time credentials, agent
instructions, and live connection feedback.
- Persist pending setup, test delivery receipts, connection status, and
reversible trigger removal.
- Keep setup checks free of routine runs, tasks, and agent wakeups.
Preserve delivery idempotency after activation.
- Add compact trigger cards, inline editing, key rotation, pause
controls, removal, and Undo.
- Keep Runs and Activity in the routine. Use the shared task list and
compact activity rows.
- Permit only exact public delivery POSTs through actor authentication.
Retain webhook authentication, JSON-object validation, and log
redaction.
- Add production-backed Storybook states and focused server, database,
and UI coverage.
- Document signing modes, setup checks, retries, rotation, HTTPS
ingress, and navigation.

## Verification

- Full workspace typecheck, build, and token gates passed on the rebased
branch. Storybook also builds.
- Focused routine, middleware, logging, shared wizard, and UI coverage
passes on the rebased branch: 195 tests across 14 files. The migration
passed on a fresh PostgreSQL database and on two repeated applications.
- Browser testing used the real app, database, and a deterministic
process worker through Tailscale HTTPS and the current Cloud proxy code.
- Verified rejected keys, safe setup deliveries, persisted state after
restart, activation, retry deduplication, key rotation, schedule
editing, removal, and Undo.
- Fresh bearer and GitHub-signed deliveries created tasks that the
worker checked out and completed. Runs and Activity stayed within the
routine.
- Current Cloud ingress tests passed. Public delivery POSTs passed
through without a browser session; management routes remained gated.
- All 54 current-head PR checks pass, including general and serialized
tests, all eight browser E2E shards, typecheck, build, runner checks,
security checks, and the canary dry run. Two optional Storybook jobs are
skipped by workflow conditions.
- Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review
threads. The stale connection-status finding is fixed and covered by a
regression test.
- No production deployment was performed.

## Risks

- Migration 0281 adds three trigger columns and a test-receipt table. It
is additive and safe to reapply. Apply it before running the new server.
Existing triggers remain live by default.
- Requests without delivery IDs are new events after activation. Senders
must reuse an event's delivery ID for retries.
- Completed webhooks keep normal dispatch behavior. Their management
connection check can start work; the UI states this.
- Removing a trigger archives it. Undo restores the URL and credentials.
Permanent deletion remains available through the existing API.
- Public ingress must remain restricted to the delivery POST route. The
tenant verifies credentials. Cloud sleeping-stack behavior is unchanged.
- Shared setup components also serve Slack. Existing setup contracts and
navigation tests cover that integration.
- Senders must use application/json with an object. Other media types
receive 415.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell
execution, and browser testing. The exact deployment model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 10:49:05 -05:00
DottaandPaperclip 36dbb7ed1c fix: harden agent chat runner tools and recovery (#13678)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat turns discussion into plans, tasks, reviews, and hires.
> - These workflows need reliable tool results and task context on the
native runner.
> - Live Claude and Codex tests exposed lost retry requests, invalid
project inputs, and a child startup crash.
> - Recovery also exposed a misleading retry action and missing child
task context.
> - This pull request fixes those paths and adds regression coverage.
> - Agents can continue the original request and operators can inspect a
stopped run.

## Linked Issues or Issue Description

**What happened?**

A failed Agent Chat retry could lose the user's question. Project
creation accepted unsupported icons in its tool schema. Codex could stop
when a helper's MCP startup event arrived before its thread lineage. A
stopped task offered Retry even when the server required execution
reconciliation. Resumed agents could miss existing delegated tasks.
Hiring and review instructions did not describe the native runner's
available tools and source requirements.

**Expected behavior**

Retries retain the selected request. Tool schemas match the API. Child
startup information does not gain authority over the parent or stop it.
Recovery actions match the server's requirements. Task context exposes
existing child work. Handoffs contain the material the assignee needs.

**Steps to reproduce**

1. Enable experimental Agent Chat in an isolated development instance.
2. Configure native Codex and ACPX Claude agents on Paperclip Runner.
3. Ask for a plan, revise it, approve task creation, and request a hire
and status report.
4. Retry a failed chat turn and check that it answers the original
request.
5. Start a Codex helper before its thread lineage arrives.
6. Resume a delegated task and inspect its existing children and saved
output.

**Paperclip version or commit**

The live failures were found at `f2c5e54dc`. This branch is rebased onto
`86b7ee992`.

**Deployment mode**

Isolated local development instance with native Codex and ACPX Claude.
No database migration or default permission change.

Related work: Refs #13284 for Agent Chat. Refs #13438 for the
server-side API receipt fix, which this branch preserves. The transport
also accepts the earlier HTTP receipt format. Refs #13655 for the
current Codex continuation and helper lineage handling, which this
branch also preserves.

## What Changed

- Preserve failed Agent Chat wake-comment IDs and session generation
from the authorized source run. Reject pre-reset retries.
- Wrap API receipts with the correct semantic call identity. Test
current and earlier receipt formats through real HTTP and runnerd.
- Classify early child MCP startup notifications as information. Keep
foreign completion and result events rejected.
- Constrain project icons on both tool surfaces and regenerate the
protocol contracts.
- Include bounded, company-scoped visible direct child tasks in task
context. Filter hidden tasks before applying the limit.
- Replace the rejected Retry action with Inspect run for native
continuation reconciliation.
- Update hiring, review handoff, status reporting, and development
guidance.

## Verification

- Live tests covered Claude and Codex questions, plan revisions,
approval, task creation, hiring, status, chat reset, failures, and
recovery.
- The recovered task produced its saved checklist and example. A later
follow-up read the existing child tasks and document without creating
more work.
- Full build, repository type checks, token gates, 142 focused tests,
188 runner TypeScript tests, and the Rust notification/descendant
regressions passed after rebase. The separate local full-suite run was
stopped after the complete CI suite passed.
- Review fixes passed the updated route, tool-authority, and icon
regression tests plus server type checking.
- Required commands: `pnpm build`, `pnpm -r typecheck`,
`PAPERCLIP_IN_WORKTREE=false pnpm test:run`, and `pnpm
check:token-gates`.
- At `4ce8047b0`, all 55 applicable GitHub checks pass (two Storybook
checks are intentionally skipped), including the complete
general/serialized test matrix, runner tests, browser tests, build, type
checks, Docker checks, and canary dry run.
- Fresh Greptile review is 5/5 on `4ce8047b0`; all three findings were
fixed with regressions and there are no unresolved review threads.
- Two initial CI service-startup timeouts passed unchanged in local
reproductions and in the latest CI run.

## Risks

- The new event classification is limited to MCP startup information. It
does not authorize foreign task completion, results, or tool requests.
- Task context returns at most 100 direct child tasks and reports
truncation. It excludes hidden tasks and other companies. This improves
delegation context but does not enforce semantic duplicate detection.
- Native reconciliation still requires an operator to inspect and record
prior outcomes. The new link does not replace the recovery API.
- API tools remain opt-in. Claude permission choices remain explicit. No
default permission, schema, or workflow changes.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository editing, code
execution, API tools, and browser testing. The exact deployment
identifier and context-window size are not exposed in this session. Live
acceptance agents used OpenAI `gpt-5.6-sol` and Anthropic
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:11:04 -05:00
DottaandPaperclip 9335b7db10 fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance.

Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale.

Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-19 08:51:18 -05:00
86b7ee992c feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents have a persistent visual identity (#13171): a ClipLab
character in one of 17 palettes, rendered as cached PNGs in lists and as
a live character in larger placements.
> - The onboarding wizard is where a person meets that identity first,
and it showed a stock ClipLab expression on the previous engine while
the rest of the app would show a different character on a newer one.
> - The wizard's steps also cut from one screen to the next, so the arc
read as separate pages rather than one walk.
> - This pull request puts one character on one engine everywhere, gives
the wizard's hero the studio's sleepy → wink → idle sequence on Review,
and hands the steps over inside one presence.
> - The benefit is that what wakes on Review is exactly what the agent
looks like on the dashboard afterwards, and the walk to it reads as one
screen changing.

## Linked Issues or Issue Description

Refs #13171, now merged into master. This PR contains the onboarding and
ClipLab update on top of that foundation. Original feature work by
@tonio-alucema; merge preparation preserves the original commits.

**Problem or motivation**

The onboarding hero and the app's avatars were two different characters
on two different ClipLab engines. Steps 1 → 4 of the wizard cut between
screens, and the wizard mounted cold when a cloud-managed workspace
arrived from Cloud's naming screen.

**Proposed solution**

Vendor ClipLab v0.2.0 as the shared engine and render one
studio-exported character from it in every palette, for every pose and
size. Play the export's one-shot wake on Review with the palette fading
in over the gray dormant loop. Hand steps over inside one presence so
the footer slides instead of jumping, and play the arrival half of that
hand-off when the wizard opens directly on the agent step.

**Alternatives considered**

Exporting mp4/webm loops per size: no cursor following, no clean alpha,
and the palette "colour in" is a runtime blend. Minting a `cap-v2`
character version: nothing had shipped `cap-v1`, so the artwork is
regenerated in place instead of migrated. Keeping the separately
vendored runtime bundle for the hero: two engines and two characters in
one app.

## What Changed

- `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0
(`987b6db0`) with the Paperclip adaptations replayed (optional graphics
backend for the Node SVG snapshot path, supersampled live textures,
character framing, deterministic SVG id prefixes); new upstream
`particles.ts`.
- `packages/shared/src/cliplab/character.ts`: the studio export,
mirrored from `ui/src/assets/cliplab/onboarding.character.json` by
`scripts/sync-cliplab-character.mjs` (drift caught by
`check:token-gates`). `characterDefinition` builds every palette from
it; the resting portrait is its idle beat.
- `OnboardingCharacter`: gray `sleepy` loop through the agent and
connect steps; on Review the one-shot sleepy → wink → idle plays on two
lock-step canvases while the palette fades in, then the `idle` loop.
Body-follows the pointer, page-scoped. 160px in the wizard.
- `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one
`AnimatePresence` (departing content fades and gives its room back;
arriving content opens its room then fills); the hero has a room that
opens on the walk into the agent step; opening directly on the agent
step plays the arrival half; the self-hosted naming step uses the arc's
label and field.
- Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`,
`ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`).
- Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent
arc` walkable from the naming step plus `Arrive From Cloud`; the
companies fixture answers the wizard's create call with a company.
- Uses the shared runtime for onboarding; `doc/agent-personas.md`
documents the shared character.
- Releases both onboarding canvases after partial startup or transition
failure. Registers each canvas before seeking so synchronous render
errors can release it. Six component tests cover these failures and
palette changes before or during wake.
- Refreshes both sleeping canvases when the palette changes, including a
palette change in the same render as wake.
- Moves choreography values into the CSS token layer and preserves the
shared motion catalog drift check across the imported stylesheet.
- Repairs the static Storybook avatar route and uses accessible heading
names/current button labels in the wizard play functions.
- Closes the lazy avatar worker pool during application shutdown.

## Verification

- Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm
build-storybook`, and `pnpm check:token-gates` pass. The final UI
typecheck and 123 focused onboarding, lifecycle, and token catalog tests
pass. All 55 checks on final head
`b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full
sharded test suite, runner verification, and all eight browser shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)).
The duplicate monolithic local `pnpm test:run` was stopped after CI
completed; it is not claimed as a separate completed local run.
- Chromium walkthrough: palette change, wake, return to sleep, WebGL
failure fallback, Review step hand-offs and cloud arrival pass with
normal and reduced motion; no browser errors. The signoff happy-path
browser test also passes against a disposable instance.
- The final CI run confirms the catalog fix and a passing signoff
browser shard. The earlier signoff failure was a heartbeat-run
availability timeout; the focused local reproduction and final CI passed
without signoff code changes.
- Original author verification:
- `pnpm check:token-gates` (includes the new character sync check);
shared, server avatar/persona (17) and UI onboarding/persona (137)
suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs
through the worker pipeline.
- Storybook: `Agents / Personas` Sizes, Expressions and Palettes render
the studio character at every size and pose; `Onboarding / Character →
Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc`
walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the
arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by
~560ms, footer travel continuous).
- The original author walked the agent → connect → review flow and wake
after a real sign-in on staging.
- Not done here: the Linux Storybook visual baselines
(`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for
the new engine, hero size and naming-step changes.

## Risks

- Every avatar's pixels change (new engine, new character) under the
unchanged `cap-v1` name. Stacks that rendered avatars on the previous
engine keep those PNGs in their cache
(`generated-agent-avatars/cap-v1/...`, served immutable) until cleared;
only the two pinned staging stacks ever did.
- The one-shot handoff to the idle loop is timed from the sequence's
authored duration (the engine reports completion by continuing into idle
itself); presentation only, nothing in the wizard's state waits on it.
- Reduced motion skips the wake and the hand-offs; jsdom is treated the
same way, so the wizard tests see the next step's content immediately.
- The committed export differs from the studio by one animation (Loop
off, leading idle step removed); a re-export without that fix would play
a 5.6s idle before the wake.

## Model Used

Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in
Claude Code, with shell, browser, and file tools. The original context
window was not recorded.

Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex,
with reasoning, shell execution, file editing, GitHub CLI, and automated
tests. The session does not expose an exact runtime model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 08:30:14 -05:00
DottaandPaperclip c9e8677979 docs: document chat connector UX and make the runbook self-contained (#13675)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Connections let agents work with external services.
> - Contributors use the connection runbook to add and review providers.
> - The runbook depended on private issue references and did not capture
the chat setup UX conventions.
> - This pull request adds a chat connector UX guide and puts the
missing requirements in the runbook.
> - Contributors can apply the guidance without access to the internal
issue tracker or a personal skill installation.

## Linked Issues or Issue Description

**Issue type**

Missing documentation and unclear contributor instructions.

**Where is the issue?**

`doc/connections/CONNECTOR-PLAYBOOK.md` and the chat connector setup
guidance.

**What's wrong?**

The runbook sent contributors to private issues for validation, OAuth
ownership rules, and catalog review. The Slack setup work also
established useful UX rules that other providers should share.

**Suggested fix**

Add a companion UX document. Link it from the runbook. Include
validation and review requirements directly in the public documentation.
Use portable company and instance examples.

Related implementation:
https://github.com/paperclipai/paperclip/pull/13638. A search of related
PRs found no duplicate documentation change.

## What Changed

- Add `CHAT-CONNECTOR-UX.md` with setup, credential, identity, footer,
test, and management conventions.
- Include adaptation examples for Discord, Telegram, and email
providers.
- Link the companion from the runbook introduction, contents, and UX
section.
- Replace private issue references with inline architecture boundaries,
risk classification, validation evidence, and per-tool review
requirements.
- Replace personal deployment examples with sample company and instance
addresses.
- Mark the recorded Notion provider observations as a dated snapshot.

## Verification

- Passed `git diff origin/master --check`.
- Passed local validation of all 35 relative links and heading anchors
in the two documents.
- Passed code-fence and private-reference scans.
- Reviewed the guide against the Slack setup decisions and the existing
connection docs.
- Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. They
could not complete because this fresh worktree has no installed
dependencies (`@types/node` and the CLI `tsx` entry point are missing).
No application code changed.
- GitHub CI passed on commit `070fefc8b2513cb9469bae76e66940c33c1465ea`,
including typecheck, build, general and serialized tests, runner
verification, browser tests, and canary dry run.
- Security checks passed. Greptile scored the current commit 5/5 with no
findings or unresolved review threads.

## Risks

Low risk. This changes two documentation files only. Provider APIs and
screens can change. The guide requires authors to verify provider
capabilities, and the Notion example identifies its observation date.
The documentation does not assert new runtime support.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and context
window are not exposed in this session. Used reasoning, repository
inspection, shell execution, and documentation editing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal issue
id or instance-derived details
- [x] I have run the applicable documentation checks locally and they
pass; application checks were attempted and the environment limitation
is recorded above
- [x] I have added or updated tests where applicable (documentation
validation; no runtime changes)
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 08:16:56 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
Devin FoleyandPaperclip 685d4faba3 Fix PostgreSQL recovery after a transaction connection closes (#13643)
Reject queued and late work from disconnected transaction and reservation
scopes. Keep closed reservations out of the open pool, and clear old
connection buffers and responses so new requests can reconnect safely.

Twelve real-PostgreSQL regression cases cover crash prevention, recovery,
and transaction isolation in both ESM and CommonJS. Database checks and
all PR CI checks pass. Greptile: 5/5, no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 16:27:28 -07:00
DottaandPaperclip 924f07be8c feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connections let people start and continue that work from Slack.
> - Setup mixed app creation, credentials, URL verification, account
linking, and testing on the same screens.
> - People also needed a safe way to link their own Slack identity after
the first operator finished setup.
> - This pull request gives each step a clear place and keeps membership
approval separate from identity linking.
> - It also makes connection details easier to use and fixes misleading
callback health behind HTTPS proxies.

## Linked Issues or Issue Description

**Subsystem affected**
Cross-cutting: chat routes and services, shared contracts, and the Apps
board UI.

**Problem or motivation**
Slack onboarding made users find settings without enough guidance. A
second user needed operator help to link their account. Activity stopped
at 100 records, and TLS termination could mark working callbacks as
stale.

**Proposed solution**
Use six setup steps with editable app names, a generated manifest,
credential guidance, URL verification, account linking, and an optional
message test. Send each Slack user a private, expiring confirmation
link. Require company membership or an approved access request before
linking. Add cursor pagination and tolerate the internal HTTP hop in
callback diagnostics.

**Roadmap alignment**
This improves the existing connected-app surface and supports CEO Chat
without changing the task-and-comments model. The maintainer requested
and reviewed the flow during a live Slack test drive.

**Additional context**
Related work: #7, #3349, #13000, and #13620. Those cover broader chat
capabilities, older webhook paths, or plugins. This PR improves the
existing native connector's setup and account-linking flow. HTTPS
documentation was published separately in
paperclipai/paperclip-docs#128.

## What Changed

- Split Slack onboarding into six clickable sidebar steps. Keep
secondary and primary actions on one row.
- Generate the Slack creation link and read-only manifest from editable
app, bot, and command names. Add credential prefix validation and direct
instructions.
- Add live account-link status and an optional mention-based message
test.
- Add private, single-use Slack account invitations and membership
access requests. Retain cloud authentication/bootstrap checks and
enforce the chat rollout flag in all identity APIs. Default new Slack
connections to linked users only.
- Put Settings, Access, Conversations, and Activity in the sidebar.
Simplify conversation rows and remove active header badges.
- Add 25-item activity pages, stable timestamp/ID cursors, and replay
safety across pages. Preserve the legacy array API for clients without
pagination parameters.
- Fix false callback warnings when HTTPS terminates at a proxy. Keep
host, port, and path drift detection.
- Document the setup flow, pagination, callback diagnostics, and shared
wizard footer rule.

## Verification

- Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed: focused Slack callback and pagination integration tests; UI
clipboard, wizard, pagination, and activity tests; OpenAPI route tests.
The final access-gate fix also passes 27 focused tests covering cloud
authentication/bootstrap, nonmember invitations, token validity, and the
server-enforced rollout flag.
- Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11
provider browser scenarios (including mobile light/dark navigation).
After rebase, the identity route, sidebar, and 25 clipboard tests pass.
- The full local `pnpm test:run` was attempted. The first run found 14
Slack fixtures that needed explicit guest access; those are fixed and
the complete chat suite passes. Unrelated embedded PostgreSQL
startup/resource failures and timeouts prevented a clean full local run.
All CI checks pass on `2d858b036`, including the full chat, server,
workspace, build, typecheck, and browser suites.
- Live test drive: Slack app creation, credential setup, URL
verification, private account confirmation, mention messages, and thread
replies. Verified the callback warning clears for the existing proxied
connection.
- Review: create a Slack connection, follow the six steps, link a second
user's account, and browse older activity with Next and Previous.

## Risks

- Identity invitations carry a temporary capability. Tokens are hashed,
expire after 15 minutes, work once, and require explicit confirmation by
a company member. Access requests do not grant membership.
- New Slack connections reject unlinked people by default. Existing
connection settings remain intact.
- Activity is a live ledger. Updated action rows can move forward in
time. Older pages do not poll.
- Proxy tolerance affects health display only. Slack signature checks
and proxy authentication settings remain unchanged.
- No database migration or package-lock changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose an
exact model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; full
local-run limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-18 17:23:53 -05:00
Devin FoleyandPaperclip 4b8dc416da Fix default isolation for projects without workspace configuration (#13636)
Require a company-scoped configured workspace before applying the operator default for Git worktree isolation. Projects that only have a plain managed directory retain their existing behavior. Explicit isolation requests still require a valid checkout.

Add policy and heartbeat integration regressions and document the default. The regression fails before the fix. An isolated checkout passes 541 relevant tests, and the server TypeScript check passes. Local repo-wide typecheck and build require the missing Rust toolchain; all CI lanes passed, and Greptile reviewed the refreshed head at 5/5 with no comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 14:49:45 -07:00
scotttong 3f1d897a7c copy(connectors): say "organization" in the connector setup flow (#13589) 2026-09-17 19:37:14 -07:00
Devin FoleyandPaperclip 5442f2d869 fix: repair managed Git launchers in sandbox projects (#13588)
## Thinking Path

> - Paperclip runs agents in local and remote execution environments.
> - Managed GitHub launchers select credentials for each Git operation.
> - Remote launchers are written inside the project checkout as
extensionless CommonJS scripts.
> - An ES module project makes Node interpret those launchers as ESM, so
they crash before credential resolution.
> - When the launcher can start, empty identity variables also override
valid repository and command-line Git configuration.
> - This change gives the launchers their own CommonJS scope and clears
empty identity overrides while preserving managed credential isolation.

## Linked Issues or Issue Description

**What happened?**

In a repository with `"type": "module"`, the managed `git` and `gh`
launchers fail immediately with `ReferenceError: require is not defined
in ES module scope`. The launchers use CommonJS but inherited the
enclosing project's module type.

Sandbox agents also report empty `GIT_AUTHOR_NAME` and
`GIT_COMMITTER_NAME` variables and try to unset them for each command.
With no managed identity available, the Git launcher recreated those
empty values. `git commit` failed with `fatal: empty ident name`, even
with explicit `user.name` and `user.email` configuration.

**Expected behavior**

Managed `git` and `gh` start in both ES module and CommonJS projects.
Local commits with an explicitly configured identity work without manual
environment cleanup. Managed credentials and captured identity continue
to take precedence. Missing identity does not silently select the host
user's details.

**Steps to reproduce**

1. Create a sandbox project whose `package.json` contains `"type":
"module"`.
2. Stage the managed GitHub launchers and run `git --version` or `gh
--version`. Before this fix, the launcher fails at its first
`require()`.
3. In a CommonJS project with no available managed identity, configure
repository `user.name` and `user.email`, or supply them with `git -c`.
4. Run `git commit --allow-empty -m test`. Before this fix, both
identity configuration forms fail with empty identity.

**Paperclip version or commit**

Reproduced from master commit `165b10bd9`.

**Deployment mode**

Sandbox execution. The shared launcher is also used for managed local
and SSH execution.

Related work: #13094 introduced the local-operation fallback; #13053
changes launcher discovery on Windows. Neither fixes empty identity
overrides. Related identity work in #8945 and #8946 configures worktree
authorship and does not remove these environment overrides.

## What Changed

- Stage `package.json` with `"type": "commonjs"` in the launcher
directory before the Node scripts. Keep the project's package
configuration unchanged.
- Leave inherited author and committer variables unset in the real Git
process. When credentials are absent, require explicit Git identity
configuration with `user.useConfigOnly`.
- Clear empty identity merge overrides in staged shell profiles after
environment merging. Preserve nonempty captured identity values.
- Exercise real Git commits with repository and command-line identity,
broker failures, and managed-user switching. Verify startup in ES module
and CommonJS projects, shell cleanup, and captured identity
preservation.
- Document launcher module scope and local identity behavior in the
execution GitHub identity contract.

## Verification

- Confirmed both new local-commit regression cases fail before the fix
with `fatal: empty ident name`.
- Confirmed the new ES module project regression fails before the fix
with `require is not defined in ES module scope`.
- Focused launcher and shell tests: 28 passed.
- `pnpm exec vitest run --project @paperclipai/adapter-utils --exclude
'**/dist/**'`: 1,216 passed, 11 skipped across 58 files.
- `pnpm --filter @paperclipai/adapter-utils typecheck` and `pnpm
--filter @paperclipai/adapter-utils build`: passed.
- `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the
unchanged native runner because Cargo is not installed on this machine.
- Full `pnpm test:run`: started locally; stopped the duplicate run after
the complete CI suite passed. No local full-suite success is claimed.
- CI on `99ea8050e`: all 53 checks passed (2 skipped), including full
tests, typecheck, build, native runner checks, and browser checks.
- Greptile reviewed `99ea8050e`: 5/5 with no findings or unresolved
comments. GitHub reports no merge conflicts with master.
- No live sandbox or GitHub push probe performed.

## Risks

- The new package scope is confined to the run-specific launcher
directory. It does not change the project's module type, launcher names,
or credential selection.
- Without a managed identity, an explicitly configured repository author
can now create local commits. GitHub access remains subject to the
existing credential broker. Global/system Git configuration, ambient
credentials, and SSH identity remain isolated.
- Managed identity still wins over repository settings. Missing local
identity still fails instead of guessing host details.
- New or resumed executions must stage the updated launcher and shell
profiles. Existing processes retain their prior files and environment
until refreshed. No database migration or sandbox image rebuild is
required.
- Revert this change to restore the prior behavior.

## Model Used

- OpenAI GPT-6 via Codex, with code inspection, implementation, and
local test execution. The hosted model variant and context window were
not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (the affected adapter-utils
package)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 18:40:22 -07:00
DottaandPaperclip 84fe89906d fix: complete native agent review handoffs (#13581)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Native execution uses durable runs, issue locks, wake requests, and
typed tool authority
> - A child can finish with a native agent review request while its
original assignee stays responsible for the work
> - The reviewer then needs a bounded execution path that can inspect
the child, record one decision, and finish safely
> - Before this change, assignee-only gates rejected the reviewer or
left the parent waiting after the child review ended
> - This pull request adds typed reviewer admission, scoped reviewer
tools, durable wake and recovery handling, and parent continuation
evidence
> - The benefit is that native review handoffs complete without changing
child ownership or granting broad mutation access

## Linked Issues or Issue Description

Refs: #13314
Refs: #13574

**What happened?**

A native child run could report `needs_review` for an agent reviewer.
The reviewer wake then failed assignee and execution-lock checks. The
child remained in review and the parent remained waiting.

**Expected behavior**

The named reviewer should receive one durable wake. The reviewer should
inspect the child and resolve the exact review card. The child assignee
should stay unchanged. The parent should receive the recorded review
outcome after the child reaches its terminal state.

**Steps to reproduce**

1. Run a native task with a different named agent reviewer.
2. Keep the child assigned to its original worker.
3. Let the worker finish with a native completion review request.
4. Start the durable reviewer wake.
5. Resolve the review and finish the reviewer run.
6. Observe the child and parent state.

**Paperclip version or commit**

Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification
source: `eea171aae`.

**Deployment mode**

Built from source.

**Installation method**

Built from source (pnpm build).

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Access context**

Both.

**Database mode**

Embedded PostgreSQL in the isolated live test fixtures.

## What Changed

- Add server-validated native review assignment facts.
- Admit only the exact company, issue, source run, decision, revision,
addressee, and resolver policy.
- Give reviewer runs a narrow set of Paperclip read and resolve tools.
File and shell access follow the configured agent and environment
policy, so reviewers can run tests.
- Separate server-owned reviewer instructions from untrusted persisted
review data. Escape the data boundary; retain server-enforced
authorization.
- Keep the child assignee unchanged. Atomically claim the reviewer run,
wake request, and issue execution lock. A competing lock prevents
provider startup.
- Require the exact running reviewer session and current issue lock to
resolve its assigned card. Reject missing, unrelated, or terminal
reviewer runs.
- Add durable reviewer wake, lock, stale-card, and abandoned-run
recovery handling.
- Prevent duplicate native wake dispatches during deferred admission and
recovery.
- Carry accepted or rejected child review outcomes into parent task
context and continuation evidence.
- Add focused server, runner, and native protocol coverage.
- Preserve upstream continuation rules. Add child review decisions as
separate evidence, while keeping real human answers in their own field.
- Return actionable completion validation feedback to both providers.
Permit a corrected completion after rejection. Keep strict terminal
acknowledgment validation.
- Apply exclusive shared-workspace locks to sandbox environments. Local
and SSH folders can run concurrently, including when old settings
request serialization.
- Repair test timing, native event parsing, and the review artifact
assertion. Allow a valid reject, correct, and accept review sequence.
Check the accepted card against its reviewer run and decision. Keep
polling within the existing deadline when review acceptance precedes the
parent wake projection; report a specific missing-continuation error at
timeout.
- Apply the ACPX pending-call limit to reserved finish/block calls, with
capacity-release and cancellation tests.

## Verification

- `pnpm build`: passed on `eea171aae`.
- `pnpm -r typecheck`: passed on `eea171aae`.
- `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on
`b31ad9ab8`; runner E2E typecheck also passed.
- `pnpm check:token-gates`: passed.
- Focused DB review, reviewer authority, and prompt-boundary checks: 31
tests passed. They cover invalid reviewer runs, competing locks, atomic
admission, duplicate claims, and valid resolution.
- Heartbeat, workspace, and recovery checks: 30 tests passed.
- ACPX sidecar suite: 27 tests passed. Moving the capacity guard back
below reserved handling makes both new regression cases fail.
- Four focused live continuation checks passed on their first attempt at
`f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and
question tool guidance (12/12 each). These cases do not use the reviewer
prompt path changed afterward.
- Fresh Codex and Claude review-handoff checks passed all 29 native
checks each on their first attempt at `eea171aae`. Both runs received
the expected fixed prompt and completed cleanup. Only the six selected
live flows were tested; no full paid provider catalog run.
- The final commit only extracts the existing test-harness timeout
diagnostic into a shared helper and adds positive and negative coverage.
Removing the accepted-review guard makes two regression assertions fail;
restoring it passes all six timeout tests. Production runtime code,
prompts, deadlines, and grading criteria are unchanged by this final
commit.
- Deadline regressions: a valid continuation delayed 20 seconds succeeds
within its 30-second unit-test deadline; an absent wake returns a
specific candidate-failure diagnostic at that same deadline. Both
assertions failed before the fix. Production E2E deadlines remain
unchanged.
- Historical native failures remain recorded: Docker availability
failures; a valid reject/correct/accept sequence that the first-card
grader misread; and a test that rejected the gap between accepted child
review and parent wake projection. No failed result was regraded. The
latest tests use a protected reference to the pinned Docker image and
the unchanged artifact oracle and time limits.
- Full repository verification runs in GitHub CI. Local verification
uses the focused suites above, full build, and full typecheck. An
unchanged Codex shutdown timing test failed once in CI, passed in
isolation, and its full shard passed on the final commit without changes
to that test or its causal code path. The original failure is retained
in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no
outstanding actionable findings. All review threads are resolved. All
current-head CI gates passed, including the isolated native runner
Docker build (55 successful checks; two skipped by the workflow).

## Risks

- Reviewer admission depends on exact persisted decision and interaction
bindings. A stale or changed card is rejected.
- Paperclip control-plane tools are limited to inspection and review
resolution. This is not a filesystem permission boundary; provider file
and shell access retain the configured policy.
- Deferred wake recovery changes dispatch receipt coalescing. A
scheduler regression could delay a continuation if the receipt state is
wrong.
- Parent review outcomes are evidence for the model. They do not grant
tool authority or change issue ownership.
- This change does not address legacy lease-hold handoff behavior.

> Roadmap review: native execution, review gates, and durable recovery
are existing roadmap capabilities. This PR completes a narrow
reliability path for those capabilities.

## Model Used

OpenAI `gpt-6-astra` with reasoning, tool use, and code execution.
OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and
journal work. Context window size is not exposed by this session. Live
test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the
PR authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 15:52:19 -05:00
Nicky LeachandPaperclip e926b13017 fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path

> - Paperclip must stop remote execution after losing its controller.
> - A bridge can remain blocked while the sandbox still incurs costs or
performs actions.
> - Waiting forever for that bridge prevents provider termination.
> - A temporary provider outage can also exhaust cleanup attempts
permanently.
> - This pull request bounds bridge drain and persists cleanup retries
with backoff.
> - Cleanup ends only after provider confirmation, without a user
accepting uncertain side effects.

## Linked Issues or Issue Description

Builds on merged #13285. Related #13254 added exact provider termination
receipts; merged #13272 adds explicit user retry. Merged #13352 stops
active sandbox startup before waiting for setup. This PR preserves that
immediate cancellation path and extends bounded teardown to ordinary
release and destroy. Cleanup continues automatically after repeated
provider failures. Refs #12953 for provider failures blocking execution.

**What happened?**
Daytona release waits for in-flight bridge activity before stop/delete.
A dead bridge can prevent that wait from finishing. The host also stops
cleanup after five failed attempts.

**Expected behavior**
Provider termination proceeds after a bounded bridge drain. Cleanup
retries survive service restarts and provider outages.

**Steps to reproduce**
Start a sandbox command whose bridge promise never resolves, then
release its lease. Separately, persist a pending-cleanup lease with five
failed attempts and recover the provider.

**Deployment mode**
Hosted Paperclip with a Daytona provider; rebased onto master at
`728f7185f` on September 14.

## What Changed

- Bound bridge drain and provider lifecycle calls. Prefer stop for
reusable sandboxes, with delete fallback.
- Persist cleanup attempt identity, renewable in-flight deadline, and
cooldown. Fence completion writes against superseded attempts.
- Preserve scoped explicit Retry and its activity log. Explicit Retry
can skip cooldown, but cannot take over a live cleanup attempt.
- Continue cleanup after five failures with slower retries and an
operator warning.
- Exclude leases in cooldown before paging so they do not starve due
work.
- Add hung-bridge, restart, provider-recovery, and concurrent-cleanup
regressions.

## Verification

- Rebased onto master at `728f7185f`. The outstanding diff contains only
cleanup changes; the merged controller-ownership prerequisite is
excluded.
- Daytona plugin suite: 160 passed, including immediate startup
cancellation, graceful release, hung activity, and teardown regressions.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31
passed. Two added integration cases verify explicit Retry during
cooldown and while another cleanup owns the lease. They also verify run
scoping and the activity log.
- Targeted cleanup and cancellation cases in
`environment-runtime.test.ts`: 20 passed.
- Earlier live disposable Daytona test: provider stop ended background
work, resume preserved files without restarting the old process, a new
command succeeded, and the sandbox was deleted. This verifies provider
behavior; it was not repeated for this rebase.
- Latest-head CI and automated review are pending. Broad local tests,
typecheck, and build were not rerun for this focused rebase; CI supplies
those checks.

## Risks

- Timing out bridge drain permits provider termination; it never
supplies a stop receipt.
- A crashed cleanup attempt remains protected for 15 minutes, then
becomes eligible again. Repeated failures retry every 30 minutes after
escalation.
- The existing counter saturates at the escalation threshold; the new
attempt identity and deadline prevent overlapping claims.
- No schema, UI, telemetry, lockfile, or workflow change.

## Model Used

OpenAI GPT-6 through Codex, using reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 10:01:57 -07:00
Devin FoleyandPaperclip 165b10bd98 fix: enable GitHub Actions MCP toolset (#13553)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The GitHub connector lets agents use repository tools through MCP.
> - GitHub excludes Actions from its default MCP toolsets.
> - Approval of Actions permissions therefore does not make workflow
tools appear in Paperclip.
> - This pull request adds Actions to the requested toolsets for
discovery and execution.
> - Users can refresh existing connections and use workflow tools under
the existing access rules.

## Linked Issues or Issue Description

**What happened?**

GitHub Actions tools remain absent after the GitHub App receives Actions
read/write access and the user refreshes actions in Paperclip. Paperclip
does not request the Actions MCP toolset.

**Expected behavior**

Authorized GitHub connections expose workflow tools, including
`actions_run_trigger` with `method: "run_workflow"`, so agents can
dispatch an existing release workflow.

**Steps to reproduce**

1. Connect GitHub to Paperclip with access to a repository that has a
dispatchable workflow.
2. Grant the GitHub App Actions read/write permission and approve the
installation update.
3. Refresh the connection's actions in Paperclip.
4. Observe that the workflow tools are absent.

**Paperclip version or commit**

Base commit: `fae698031`.

**Deployment mode**

Hosted instance with a managed GitHub connection. The same missing
header affects PAT connections.

No matching public issue or pull request was found in the duplicate
search.

## What Changed

- Send `X-MCP-Toolsets: default,actions` through the shared GitHub MCP
header helper. This covers discovery, refresh, and execution for
existing and new managed or PAT connections, including legacy rows
identified through `transportConfig`.
- Test catalog refresh, tool risk classification, and workflow dispatch
through a mock MCP server.
- Document tool names, workflow arguments, required GitHub permissions,
and the refresh step.

## Verification

- Passed both affected test suites: `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/tool-gateway.test.ts` (391 tests).
- Passed `pnpm check:token-gates` and `git diff --check`.
- Live provider check: the default catalog returned 45 tools.
`default,actions` returned 49 tools, with no tools removed. The four
added tools were `actions_get`, `actions_list`, `actions_run_trigger`,
and `get_job_logs`.
- Live `actions_get` / `get_workflow` call succeeded. No workflow was
dispatched during live verification.
- Passed `pnpm -r typecheck` and `pnpm build` with the existing Rust
toolchain added to PATH.
- Rechecked server typecheck and build after the legacy-connection fix;
both passed.
- The full local test run has reported three skills-cache failures in
`company-skills-service.test.ts`. All three reproduce on the untouched
base commit (`fae698031`) on this macOS host: runtime-cache directory
renames fail with `EACCES`. The full run remains in progress.
- Greptile: 5/5 on `7d391e3c7`, with no unresolved review threads.
- After deployment, use **Refresh actions** on an existing GitHub
connection and verify the workflow tools appear.

## Risks

- Refreshed GitHub catalogs expose more tools. Existing access,
approval, and quarantine rules still apply. `actions_run_trigger` keeps
GitHub's destructive classification because it also supports
cancellation and log deletion.
- GitHub still enforces token and installation permissions. Dispatch
requires Actions write permission and a workflow with
`workflow_dispatch`.
- No database migration or saved connection edit is required.

## Model Used

OpenAI GPT-6 through Codex, with code execution and tool use. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 17:10:34 -07:00
Devin FoleyandPaperclip fae6980310 revert(apps): restore Google connector visibility (#13552)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - PR #13551 temporarily hid Google connectors.
> - We now want to restore their catalog visibility.
> - This PR reverts that change and restores the previous catalog
behavior.

## Linked Issues or Issue Description

Refs: #13551

Revert the temporary removal of Google connectors from the UI.

## What Changed

- Restore Gmail and eight Google Workspace entries to the catalog.
- Restore the matching branding flags and original catalog and service
tests.
- Remove the temporary-hiding documentation note.

This is an exact revert of commit
`cf1e873ab24277d55ffd3ab06074f77014dc4015`.

## Verification

- Passed: 507 catalog, UI, and connection service tests.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Full local Vitest was not repeated because the unchanged base has
confirmed macOS skill-cache permission failures. The full CI suites
passed.
- Passed: all GitHub CI gates; Greptile 5/5 on commit
`4e3dddef0ebfef1f99001e7735822ed4cba852ab`, with no review threads.
- Reviewer check: open Connectors and confirm that Gmail and Google
Workspace entries appear again.

## Risks

Low risk. This restores the previous catalog visibility and setup entry
points. Connector implementations and saved connection data are
retained.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:35:16 -07:00
Devin FoleyandPaperclip cf1e873ab2 fix(apps): temporarily hide Google connectors (#13551)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - We need to temporarily remove Google connectors from the UI.
> - The catalog already separates visibility from retained definitions.
> - This PR uses that setting so Google can return with a small change.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The Connectors catalog and its setup entry points.

**Current behavior**
The catalog shows Gmail and eight Google Workspace connectors.

**Proposed behavior**
Temporarily hide those nine entries. Keep their definitions and existing
connections.

**Reason and benefit**
Make the temporary UI removal easy to reverse.

**Breaking changes**
Fresh catalog setup no longer offers Google. Saved connections keep the
existing management and reconnect paths.

## What Changed

- Add the nine Google connector slugs to the existing hidden list.
- Match the branding manifest visibility flags.
- Update existing catalog and service tests. Keep backend Google
connection coverage and document how to restore visibility.

## Verification

- Passed: 507 targeted tests covering catalog definitions, URL matching,
setup routing, connector UI, branding, and the connection service.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Stopped the full local Vitest run after skill-cache permission
failures. Three failures in `company-skills-service.test.ts` also
reproduce on the unchanged base branch. The final connector service
suite passes all 319 tests.
- Greptile: 5/5 on the current commit, with no open review threads. CI
is retrying one unrelated preview-server readiness timeout. That test
file passes all seven tests locally.
- Reviewer check: open Connectors in a company with no Google
connections. Gmail and Google Workspace entries should be absent.
Existing saved connections remain manageable.

## Risks

Low risk. This uses the existing catalog visibility mechanism. No
connector implementation, credential, or database schema is removed.
Restoring visibility requires updating both the hidden list and branding
manifest.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:04:46 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:44:36 -05:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00
DottaandPaperclip 18989a9e73 docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its evaluations test both the Runner and complete product workflows.
> - The guides and run histories are in separate places.
> - The shared Evalbook viewer can make the test boundary unclear.
> - This pull request names the two families and adds a guide, authoring
skills, and a public hub.
> - Contributors can choose the correct test and inspect its history.

## Linked Issues or Issue Description

**Issue type**

Missing documentation.

**Where is the issue?**

Runner and Product E2E evaluation guides, case-authoring procedures, and
public result navigation.

**What's wrong?**

There is no single entry point. A report format can be mistaken for an
execution boundary. There are no dedicated case-authoring skills for
these two families.

**Suggested fix**

Add a guide and three skills. Link both existing histories from a public
hub. Keep existing campaign URLs and grading unchanged.

## What Changed

- Add `doc/evals.md` and links from existing guides.
- Add the `paperclip-evals`, `add-runner-eval`, and
`add-product-e2e-eval` skills. Install copies in
`~/paperclipai/.agents/skills`.
- Add a static hub builder that reads the existing public history feeds.
- Show a dated snapshot for each family. Label partial campaigns and
preserve measurement dates across report refreshes.
- Document publication and refresh commands for
https://pages.paperclip.ing/evals/.

## Verification

- Seven Python summary tests pass: `python3 -m unittest discover -s
scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR
does not modify package scripts.
- All three skills pass the skill-creator `quick_validate.py` check with
`/usr/bin/python3`.
- Build tested with saved history fixtures and the live public feeds.
- Desktop and mobile browser checks pass. The mobile page has no
horizontal overflow.
- Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200,
no page errors, all eight links return HTTP 200, no mobile overflow.
- Independent skill exercises found the existing Notion-decline case and
a direct Runner permission-denial case. Roster validation with an
explicit run ID passes.
- Missing refresh measurement date: regression fails before the fix and
passes after it.
- `git diff --check` passes.
- No paid evals were run for this documentation and reporting change.
The preceding head passed typecheck, build, server/workspace tests,
runner verification, browser E2E, and the canary dry run. Checks for the
latest commit are pending. Local repository-wide typecheck, test, and
build were not repeated because no product code changed.

## Risks

The hub is a dated static snapshot. It can lag behind the linked
histories until an operator refreshes it. A changed history schema stops
the build. Existing archives and grades are not modified. The published
guide link is pinned to the reviewed commit so branch deletion cannot
break it. Later builds can use master.

## Model Used

OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna
for documentation and independent skill checks. Both used repository
tools and code execution. Context window sizes are not exposed by this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (latest commit pending)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(preceding head was 5/5; latest commit pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 08:41:06 -05:00
a8d32e5e61 feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent work runs in sandboxes that provider plugins supply
> - Operators can choose a provider to run agent work
> - CreateOS adds another provider with workspace-preserving pause and
resume
> - This pull request adds a CreateOS provider plugin
> - The benefit is that operators can preserve a workspace between runs
without keeping its compute active

## Linked Issues or Issue Description

Refs #13203 and the earlier closed #13096.

This continues the CreateOS contribution from @bhautikchudasama and
@ashwaq06. The branch preserves the original implementation commit.
Thank you to both contributors.

When squash-merging, preserve the original author's credit in the squash
commit body:

```text
Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com>
```

The original fork rejects maintainer pushes. This branch includes the
merge-conflict resolution and review fixes. The request is described
below using `adapter_request.yml`.

**Agent or provider**

CreateOS sandbox API (https://api.sb.createos.sh).

**Why this adapter is useful**

CreateOS can pause a sandbox and resume it by ID. The workspace survives
the pause. This adds a reusable-lease option to the existing sandbox
provider system.

**How the agent is invoked**

Build and install the local plugin as described in its README. Open
Instance Settings, then Environments. Select the `createos` driver.
Supply an API key and shape. The driver then supplies sandbox leases for
agent runs.

**Are you willing to implement it?**

Yes. This pull request is the implementation.

## What Changed

- Adds the `createos` sandbox provider under
`packages/plugins/sandbox-providers/createos`.
- Calls the CreateOS HTTP API directly. The package adds no vendor SDK.
- Implements the environment lifecycle hooks, incremental process
output, and binary workspace sync.
- Registers the optional bundled provider and its trusted host
credential fallback. The fallback is limited to the official API origin;
custom endpoints require an explicit key.
- Lists the package in the release manifest with `publishFromCi: false`
until its first npm publish is bootstrapped.
- Waits through delayed pause/resume state updates without duplicate
action requests.
- Cancels queued API requests promptly while preserving request spacing.
- Uses direct CLI invocation in the setup guide so paths and IDs are
passed without an extra shell expansion.
- Includes current master and retains its existing Git-subfolder
containment fix.

## Demo

Fresh setup and a run against a CreateOS sandbox.


https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110


https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084

## Verification

All 25 jobs in [CI run
34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542)
passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including
typecheck, build, native runner verification, server and workspace
tests, browser tests, and the canary release dry run. Greptile reviewed
the same commit at 5/5 with no unresolved review threads.

GitHub reports no merge conflicts. The remaining merge gate is
code-owner approval for the new `package.json`, as required by
`.github/CODEOWNERS` and the `master` ruleset. Reviewers have been
requested automatically.

Local checks passed:

- Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke
skipped), and `pnpm build`.
- Host: focused credential and bundled-plugin tests (17 passed), plus
CLI invocation safety (39 passed).
- Release: package manifest check and release policy tests (18 passed).

The full local `pnpm test:run` attempt caught the README command issue;
its focused rerun now passes. The full local run stopped after its
general-server group: 7,804 tests passed, with unrelated embedded
PostgreSQL startup failures and 10 failures in unchanged
runtime-skill-cache tests (`EACCES` on directory rename on macOS). It
did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm
build` reach the runner package and stop because this machine has no
Rust/Cargo installation. The corresponding CI checks passed on
provisioned runners, as linked above.

The live CreateOS smoke requires explicit provider credentials and was
not run during this review. It is available with `CREATEOS_LIVE_TEST=1
pnpm test` in the provider directory. The author supplied the demo links
above.

## Risks

The provider is opt-in and is not installed by default. It is available
through a local-path install or explicit image inclusion. npm
publication remains disabled until a maintainer bootstraps the package
and enables publishing.

Sandbox creation has no idempotency key. An ambiguous create response
can leave a resource that requires provider-account inspection. Process
tracking is in memory; durable lease recovery belongs to the host. The
provider does not advertise guaranteed expiry, interactive login,
snapshots, duplex channels, or ingress. Live native-runner qualification
remains outside this PR's tested claims.

## Model Used

Original provider implementation: human-authored by @bhautikchudasama,
as reported in #13203. The original description reports Claude Opus 5
assistance.

Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review,
editing, and tool execution. The precise runtime model variant and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full-suite
environment limits documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:09:43 -07:00
DottaandPaperclip 9adeabb590 fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let people give agents access to external tools.
> - A personal connection needs the current user's authorization.
> - A new MCP URL must be probed before Paperclip can discover its
sign-in method.
> - Requiring a personal grant before that probe prevents sign-in from
starting.
> - This pull request permits the initial probe for a creator-owned
draft with no credentials.
> - People can complete personal setup while later requests retain
authorization checks.

## Linked Issues or Issue Description

**What happened?**

Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and
"This connection needs the current user's authorization". Paperclip
checked for a personal grant before contacting the provider. The health
wrapper also changed the expected authorization error into a server
error.

**Expected behavior**

Discover OAuth and start browser sign-in. Create a personal grant after
consent. For a public endpoint, discover its tools and create the empty
personal grant after a successful probe. Keep missing authorization on
later health checks as HTTP 422.

**Steps to reproduce**

1. Add an unknown remote MCP URL with no saved credentials.
2. Select "Just me".
3. Check the link. Before this fix, the request fails before sign-in or
tool discovery.

**Paperclip version or commit**

The three original regressions fail against `6cfe4acff` with the service
fix removed and pass with it restored.

**Deployment mode**

The defect was reported in production and reproduced in local server
tests with isolated PostgreSQL.

Related work: Refs #11831. Refs #11144. Searches found no duplicate fix.

## What Changed

- Allow an initial credential-free probe only for the creating user's
personal draft with unknown authentication and no supplied credentials.
- Leave OAuth grant creation to the callback. Create an empty personal
grant only after a public probe succeeds.
- Preserve the personal grant for URLs that already contain a
credential.
- Make empty personal grant creation conflict-safe without overwriting a
concurrent grant or duplicating its creation audit.
- Commit the empty grant and audit atomically. Retain the established
public draft identity after a later catalog failure, so a failed retry
cannot remove a successful retry's grant.
- Run the following catalog/default-profile step in its own transaction,
so failures discard partial catalog, profile, binding, and audit changes
without deleting the established identity.
- Verify archived personal connections retain their owner: another user
is rejected before probing, while the original owner can resume setup.
- Preserve `user_authorization_required` and HTTP 422 in health
failures.
- Cover the connect and OAuth callback routes, real loopback HTTP,
credential-bearing URLs, and later health checks.
- Document personal setup and the test fixtures.

## Verification

- Red/green: the original three tests failed with the exact reported
error before the fix and passed after it.
- The credential-bearing personal URL regression also failed before its
guard was added.
- Focused suite: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/generic-mcp-connection.test.ts
src/__tests__/tool-access-service.test.ts` passed all 385 tests across
the two suites. An earlier run had a socket hang-up in an existing
agent-permissions test; the unchanged suite passed on rerun.
- The concurrent rollback regression failed before its fix because the
successful retry's grant was deleted. It now verifies the grant and
draft survive and a later normal health check succeeds.
- Database fault injection during profile-entry insertion reproduced
partial catalog writes before the transaction fix. The regression now
verifies unchanged catalog rows, no partial profile/bindings, a retained
grant, and successful retry.
- CI's first serialized-server shard 3 attempt failed an existing
peer-agent mutation test (the real run-context guard ran despite the
test's mock). The test passed in isolation and all 108 tests in that
suite passed unchanged locally. The single failed-shard rerun passed
without code changes.
- Final-commit CI: all 32 applicable checks passed on
`639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected
tests, the full test matrix, browser suite, build, typecheck, release
checks, and security checks. The two Storybook-only checks were not
applicable and skipped. Greptile is 5/5 with all review threads
resolved. [Successful CI
run](https://github.com/paperclipai/paperclip/actions/runs/35027478353).
- `pnpm -r typecheck` passed.
- `pnpm smoke:mcp-fixtures -- --require-paperclip` passed.
- `pnpm build` passed.
- Full local `pnpm test:run` was attempted: its general-server group
finished with 12,372 passed, 4 failed, and 70 skipped tests. The run
started before review edits; its two MCP failures used the old cached
service (including an insert without the new conflict clause). All 385
focused tests pass on the final code. The other failures were existing
workspace-cleanup and runtime-port tests; their unchanged suites passed
on rerun (66 passed, and 25 passed/3 skipped). The local command stopped
before later groups. The final-commit CI matrix is the full-suite merge
gate; this local run is not claimed as green.

## Risks

- The initial probe must not become a general authorization bypass. It
is restricted to the creating user's draft. Normal health checks retain
authorization enforcement.
- Public endpoints get a personal grant with no secrets only after they
answer successfully. Credential-bearing URLs keep their existing grant.
- The concurrent-probe regression seeds the catalog and default profile
to isolate grant creation. Existing first-time catalog/profile creation
races are outside this change; this does not claim to make the entire
setup flow concurrency-safe.
- No database migration or UI change is required. OAuth tests use a
simulated provider; the public endpoint test uses real loopback HTTP.

## Model Used

OpenAI Codex, GPT-6-based assistant for regression tests and PR
preparation; a GPT-5-based Codex assistant assisted with the initial
implementation. Exact runtime model IDs and context-window sizes are not
exposed in this session. Both used reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:11:59 -05:00