Commit Graph
1831 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip e18ed02a56 Validate heartbeat run IDs before database lookups (#13657)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators inspect heartbeat runs, logs, and provider traces through
the API.
> - These routes use UUID database keys.
> - A malformed path value such as `undefined` reaches the database and
causes a server error.
> - This pull request validates run IDs before those lookups.
> - Valid requests keep the existing company and permission checks.

## Linked Issues or Issue Description

Related: Refs #8135. That proposal guards actor run headers and activity
writes; this fix covers heartbeat run path parameters.

**What happened?**

A request such as `GET /api/heartbeat-runs/undefined` passes a non-UUID
value to a UUID lookup and returns a server error. Run logs and the
other heartbeat run endpoints have the same unchecked path input.

**Expected behavior**

Reject malformed run IDs with HTTP 400 before any run lookup. Preserve
valid run reads, company isolation, and the existing board and
instance-admin checks.

**Steps to reproduce**

1. Request a heartbeat run endpoint with `undefined`, `null`, another
malformed ID, or a UUID with surrounding whitespace.
2. Observe the database UUID error.
3. Run the route regressions before and after this change.

**Paperclip version or commit**

Confirmed on master at `6d0342868`.

**Deployment mode**

The defect affects API deployments backed by PostgreSQL. Regression
tests exercise the actual Express routes and authorization code with
stubbed services. The historical request does not identify the
originating client, so this change does not alter a guessed UI caller.

## What Changed

- Share strict run-ID validation across the 12 heartbeat run endpoints
in the agent router.
- Keep existing board and instance-admin gates ahead of validation. Keep
valid-run company and telemetry checks intact.
- Reject surrounding whitespace, which the shared UUID helper accepts
but PostgreSQL rejects.
- Encode the UUID constraint and document the 400 response in OpenAPI.
Test the generated parameter pattern on all 12 endpoints.
- Cover malformed IDs on every affected endpoint, uppercase UUIDs,
missing and cross-company runs, and permission precedence. Use
UUID-shaped run fixtures in existing route tests.

## Verification

- Before the fix: four malformed-ID regression cases fail; three
access-control cases pass.
- Focused agent route, permission, cross-company, and OpenAPI suites:
154 tests passed, including the final uppercase-UUID case.
- Direct server typecheck passed: `pnpm --filter @paperclipai/server
exec tsc --noEmit`.
- The full local test attempt is still running; the complete Linux test
suite passed in CI. Local results will be recorded when it finishes.
- Complete Linux CI passed on the exact head: 53 checks passed, two
non-applicable checks skipped. One untouched preview-service readiness
test failed initially; its three targeted cases passed locally and the
failed-jobs-only CI rerun passed. Greptile scored the final head 5/5
with no unresolved comments.
- Full local `pnpm -r typecheck` and `pnpm build` reach the native
runner step and stop because `cargo` is absent. The complete Linux CI
checks passed.

## Risks

Low risk: this changes malformed route inputs to HTTP 400. Valid UUID
requests keep their existing lookup and authorization paths. There is no
migration, dependency, provider operation, or configuration change. The
separate activity router and actor run headers are outside this change.

## Model Used

OpenAI GPT-6 through Codex, with code editing, shell execution, and test
tools. The exact context window size was not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full-check limits
are recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-18 21:36:09 -07:00
Devin FoleyandPaperclip 6d03428682 Skip task-only connector reads for agent chat views (#13654)
Agent chat views reuse the task surface with synthetic chat-prefixed IDs.
Skip their task-only email and external chat-binding queries, and reject
invalid UUIDs after existing authentication checks on both read routes.
Preserve normal task reads and company isolation.

Verified failing regressions before the fix, all 6,377 UI tests, route and
OpenAPI regressions, server/UI TypeScript checks, all Linux PR CI gates,
and Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 20:14:03 -07:00
Devin FoleyandPaperclip d54b750111 Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset
time when the runtime reduced provider metadata to a generic category error.
Inspect terminal metadata in memory and retain only safe recovery labels and
a parsed reset timestamp. Preserve the existing handling of other limits.

Verified real child processes on both pinned ACPX runtimes, adapter and
server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate
a pre-existing chat regression from unrelated fixtures’ retry work.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 19:42:28 -07:00
DottaandPaperclip 924f07be8c feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connections let people start and continue that work from Slack.
> - Setup mixed app creation, credentials, URL verification, account
linking, and testing on the same screens.
> - People also needed a safe way to link their own Slack identity after
the first operator finished setup.
> - This pull request gives each step a clear place and keeps membership
approval separate from identity linking.
> - It also makes connection details easier to use and fixes misleading
callback health behind HTTPS proxies.

## Linked Issues or Issue Description

**Subsystem affected**
Cross-cutting: chat routes and services, shared contracts, and the Apps
board UI.

**Problem or motivation**
Slack onboarding made users find settings without enough guidance. A
second user needed operator help to link their account. Activity stopped
at 100 records, and TLS termination could mark working callbacks as
stale.

**Proposed solution**
Use six setup steps with editable app names, a generated manifest,
credential guidance, URL verification, account linking, and an optional
message test. Send each Slack user a private, expiring confirmation
link. Require company membership or an approved access request before
linking. Add cursor pagination and tolerate the internal HTTP hop in
callback diagnostics.

**Roadmap alignment**
This improves the existing connected-app surface and supports CEO Chat
without changing the task-and-comments model. The maintainer requested
and reviewed the flow during a live Slack test drive.

**Additional context**
Related work: #7, #3349, #13000, and #13620. Those cover broader chat
capabilities, older webhook paths, or plugins. This PR improves the
existing native connector's setup and account-linking flow. HTTPS
documentation was published separately in
paperclipai/paperclip-docs#128.

## What Changed

- Split Slack onboarding into six clickable sidebar steps. Keep
secondary and primary actions on one row.
- Generate the Slack creation link and read-only manifest from editable
app, bot, and command names. Add credential prefix validation and direct
instructions.
- Add live account-link status and an optional mention-based message
test.
- Add private, single-use Slack account invitations and membership
access requests. Retain cloud authentication/bootstrap checks and
enforce the chat rollout flag in all identity APIs. Default new Slack
connections to linked users only.
- Put Settings, Access, Conversations, and Activity in the sidebar.
Simplify conversation rows and remove active header badges.
- Add 25-item activity pages, stable timestamp/ID cursors, and replay
safety across pages. Preserve the legacy array API for clients without
pagination parameters.
- Fix false callback warnings when HTTPS terminates at a proxy. Keep
host, port, and path drift detection.
- Document the setup flow, pagination, callback diagnostics, and shared
wizard footer rule.

## Verification

- Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed: focused Slack callback and pagination integration tests; UI
clipboard, wizard, pagination, and activity tests; OpenAPI route tests.
The final access-gate fix also passes 27 focused tests covering cloud
authentication/bootstrap, nonmember invitations, token validity, and the
server-enforced rollout flag.
- Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11
provider browser scenarios (including mobile light/dark navigation).
After rebase, the identity route, sidebar, and 25 clipboard tests pass.
- The full local `pnpm test:run` was attempted. The first run found 14
Slack fixtures that needed explicit guest access; those are fixed and
the complete chat suite passes. Unrelated embedded PostgreSQL
startup/resource failures and timeouts prevented a clean full local run.
All CI checks pass on `2d858b036`, including the full chat, server,
workspace, build, typecheck, and browser suites.
- Live test drive: Slack app creation, credential setup, URL
verification, private account confirmation, mention messages, and thread
replies. Verified the callback warning clears for the existing proxied
connection.
- Review: create a Slack connection, follow the six steps, link a second
user's account, and browse older activity with Next and Previous.

## Risks

- Identity invitations carry a temporary capability. Tokens are hashed,
expire after 15 minutes, work once, and require explicit confirmation by
a company member. Access requests do not grant membership.
- New Slack connections reject unlinked people by default. Existing
connection settings remain intact.
- Activity is a live ledger. Updated action rows can move forward in
time. Older pages do not poll.
- Proxy tolerance affects health display only. Slack signature checks
and proxy authentication settings remain unchanged.
- No database migration or package-lock changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose an
exact model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; full
local-run limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-18 17:23:53 -05:00
Devin FoleyandPaperclip 4b8dc416da Fix default isolation for projects without workspace configuration (#13636)
Require a company-scoped configured workspace before applying the operator default for Git worktree isolation. Projects that only have a plain managed directory retain their existing behavior. Explicit isolation requests still require a valid checkout.

Add policy and heartbeat integration regressions and document the default. The regression fails before the fix. An isolated checkout passes 541 relevant tests, and the server TypeScript check passes. Local repo-wide typecheck and build require the missing Rust toolchain; all CI lanes passed, and Greptile reviewed the refreshed head at 5/5 with no comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 14:49:45 -07:00
Devin FoleyandBender de9414dd8b fix: return empty read instead of past-EOF range when log reader is caught up (#13592)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server streams agent run logs to the UI. It reads the log in
byte ranges from a local file, or from an S3 mirror after a pod restart.
> - The range math in `readS3Range()` clamps the range end up to the
range start. A fully caught-up reader then asks S3 for the range
`bytes=total-total`.
> - S3 rejects a range that starts at the end of the object. It returns
a 416 `InvalidRange` error. The API turns this into a 500 error, and the
log poller repeats it.
> - This pull request removes the clamp. A caught-up or past-EOF reader
now gets an empty read, and the server does not send an invalid range to
S3.
> - The benefit is that log polling after a pod restart does not cause
repeated 500 errors.

## Linked Issues or Issue Description

No public GitHub issue exists for this bug. The description follows the
bug report template.

**What happened?**

The run-log API returns a 500 error when a client polls a run log that
lives on the S3 mirror and the client is fully caught up (`offset ===
total`). The cause is in `server/src/services/run-log-store.ts`. The
function `readS3Range()` computes `end = Math.max(start, Math.min(start
+ limitBytes - 1, total - 1))`. When `offset === total`, the
`Math.max(start, …)` clamp forces `end` up to `start`. The `start > end`
empty-read guard does not operate, and the code sends `Range:
bytes=total-total` to S3. S3 rejects a range that starts at or past the
end of the object with a 416 `InvalidRange` error. The error monitor
records this error many times, only in the staging environment, because
only the S3 fallback path is sensitive to it. The local-file path has
the same math, but Node file streams accept past-EOF reads. The function
`readFileRange()` in
`server/src/services/workspace-operation-log-store.ts` has the same
latent math.

**Expected behavior**

A caught-up reader gets an empty read: `{ content: "", nextOffset:
undefined }`. The server does not send an invalid range request to S3.
The poller sees no contract change.

**Steps to reproduce**

1. Start a run and let it write a run log.
2. Let the log upload to the S3 mirror, and remove the local file (this
occurs when the pod restarts).
3. Poll the run-log read endpoint until the client offset is equal to
the log size.
4. Poll one more time. The server sends `bytes=total-total` to S3, S3
returns 416 `InvalidRange`, and the API returns a 500 error.

**Relevant logs or output**

```
InvalidRange: Invalid range
    at readS3Range (server/src/services/run-log-store.ts)
```

## What Changed

- `server/src/services/run-log-store.ts` — remove the up-clamp in
`readLocalRange()` and `readS3Range()`. A caught-up or past-EOF reader
gets an empty read.
- `server/src/services/workspace-operation-log-store.ts` — apply the
same fix to the shared math in `readFileRange()`.
- `server/src/services/run-log-store.test.ts` — the in-memory S3 mock
now rejects past-EOF ranges with `InvalidRange`, the same as real S3.
Add two regression tests for caught-up readers on the S3 path and on the
local path.

## Verification

- Run `pnpm vitest run server/src/services/run-log-store.test.ts`. All
17 tests pass.
- Revert only the source fix, and the new regression test fails with the
exact caught-up scenario. This shows the test covers the bug.
- Run the suites that use the workspace operation log store
(`workspace-runtime-control-recovery`,
`workspace-operations-reconciliation`). All 15 tests pass.
- Run `tsc --noEmit` on the server package. It reports no errors.

## Risks

- Low risk. The change only affects the empty and caught-up boundary of
range reads. Normal in-range reads give byte-identical results.
- Behavior change: a read with `limitBytes <= 0` now returns an empty
chunk instead of one byte. No caller passes a non-positive limit (the
default is 256000).
- Caught-up local reads keep the `nextOffset: undefined` semantics, so
pollers see no contract change.

## Model Used

- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), with extended
thinking and tool use, run through the Claude Agent SDK.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
2026-09-18 07:21:11 -07:00
DottaandPaperclip 84fe89906d fix: complete native agent review handoffs (#13581)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Native execution uses durable runs, issue locks, wake requests, and
typed tool authority
> - A child can finish with a native agent review request while its
original assignee stays responsible for the work
> - The reviewer then needs a bounded execution path that can inspect
the child, record one decision, and finish safely
> - Before this change, assignee-only gates rejected the reviewer or
left the parent waiting after the child review ended
> - This pull request adds typed reviewer admission, scoped reviewer
tools, durable wake and recovery handling, and parent continuation
evidence
> - The benefit is that native review handoffs complete without changing
child ownership or granting broad mutation access

## Linked Issues or Issue Description

Refs: #13314
Refs: #13574

**What happened?**

A native child run could report `needs_review` for an agent reviewer.
The reviewer wake then failed assignee and execution-lock checks. The
child remained in review and the parent remained waiting.

**Expected behavior**

The named reviewer should receive one durable wake. The reviewer should
inspect the child and resolve the exact review card. The child assignee
should stay unchanged. The parent should receive the recorded review
outcome after the child reaches its terminal state.

**Steps to reproduce**

1. Run a native task with a different named agent reviewer.
2. Keep the child assigned to its original worker.
3. Let the worker finish with a native completion review request.
4. Start the durable reviewer wake.
5. Resolve the review and finish the reviewer run.
6. Observe the child and parent state.

**Paperclip version or commit**

Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification
source: `eea171aae`.

**Deployment mode**

Built from source.

**Installation method**

Built from source (pnpm build).

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Access context**

Both.

**Database mode**

Embedded PostgreSQL in the isolated live test fixtures.

## What Changed

- Add server-validated native review assignment facts.
- Admit only the exact company, issue, source run, decision, revision,
addressee, and resolver policy.
- Give reviewer runs a narrow set of Paperclip read and resolve tools.
File and shell access follow the configured agent and environment
policy, so reviewers can run tests.
- Separate server-owned reviewer instructions from untrusted persisted
review data. Escape the data boundary; retain server-enforced
authorization.
- Keep the child assignee unchanged. Atomically claim the reviewer run,
wake request, and issue execution lock. A competing lock prevents
provider startup.
- Require the exact running reviewer session and current issue lock to
resolve its assigned card. Reject missing, unrelated, or terminal
reviewer runs.
- Add durable reviewer wake, lock, stale-card, and abandoned-run
recovery handling.
- Prevent duplicate native wake dispatches during deferred admission and
recovery.
- Carry accepted or rejected child review outcomes into parent task
context and continuation evidence.
- Add focused server, runner, and native protocol coverage.
- Preserve upstream continuation rules. Add child review decisions as
separate evidence, while keeping real human answers in their own field.
- Return actionable completion validation feedback to both providers.
Permit a corrected completion after rejection. Keep strict terminal
acknowledgment validation.
- Apply exclusive shared-workspace locks to sandbox environments. Local
and SSH folders can run concurrently, including when old settings
request serialization.
- Repair test timing, native event parsing, and the review artifact
assertion. Allow a valid reject, correct, and accept review sequence.
Check the accepted card against its reviewer run and decision. Keep
polling within the existing deadline when review acceptance precedes the
parent wake projection; report a specific missing-continuation error at
timeout.
- Apply the ACPX pending-call limit to reserved finish/block calls, with
capacity-release and cancellation tests.

## Verification

- `pnpm build`: passed on `eea171aae`.
- `pnpm -r typecheck`: passed on `eea171aae`.
- `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on
`b31ad9ab8`; runner E2E typecheck also passed.
- `pnpm check:token-gates`: passed.
- Focused DB review, reviewer authority, and prompt-boundary checks: 31
tests passed. They cover invalid reviewer runs, competing locks, atomic
admission, duplicate claims, and valid resolution.
- Heartbeat, workspace, and recovery checks: 30 tests passed.
- ACPX sidecar suite: 27 tests passed. Moving the capacity guard back
below reserved handling makes both new regression cases fail.
- Four focused live continuation checks passed on their first attempt at
`f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and
question tool guidance (12/12 each). These cases do not use the reviewer
prompt path changed afterward.
- Fresh Codex and Claude review-handoff checks passed all 29 native
checks each on their first attempt at `eea171aae`. Both runs received
the expected fixed prompt and completed cleanup. Only the six selected
live flows were tested; no full paid provider catalog run.
- The final commit only extracts the existing test-harness timeout
diagnostic into a shared helper and adds positive and negative coverage.
Removing the accepted-review guard makes two regression assertions fail;
restoring it passes all six timeout tests. Production runtime code,
prompts, deadlines, and grading criteria are unchanged by this final
commit.
- Deadline regressions: a valid continuation delayed 20 seconds succeeds
within its 30-second unit-test deadline; an absent wake returns a
specific candidate-failure diagnostic at that same deadline. Both
assertions failed before the fix. Production E2E deadlines remain
unchanged.
- Historical native failures remain recorded: Docker availability
failures; a valid reject/correct/accept sequence that the first-card
grader misread; and a test that rejected the gap between accepted child
review and parent wake projection. No failed result was regraded. The
latest tests use a protected reference to the pinned Docker image and
the unchanged artifact oracle and time limits.
- Full repository verification runs in GitHub CI. Local verification
uses the focused suites above, full build, and full typecheck. An
unchanged Codex shutdown timing test failed once in CI, passed in
isolation, and its full shard passed on the final commit without changes
to that test or its causal code path. The original failure is retained
in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no
outstanding actionable findings. All review threads are resolved. All
current-head CI gates passed, including the isolated native runner
Docker build (55 successful checks; two skipped by the workflow).

## Risks

- Reviewer admission depends on exact persisted decision and interaction
bindings. A stale or changed card is rejected.
- Paperclip control-plane tools are limited to inspection and review
resolution. This is not a filesystem permission boundary; provider file
and shell access retain the configured policy.
- Deferred wake recovery changes dispatch receipt coalescing. A
scheduler regression could delay a continuation if the receipt state is
wrong.
- Parent review outcomes are evidence for the model. They do not grant
tool authority or change issue ownership.
- This change does not address legacy lease-hold handoff behavior.

> Roadmap review: native execution, review gates, and durable recovery
are existing roadmap capabilities. This PR completes a narrow
reliability path for those capabilities.

## Model Used

OpenAI `gpt-6-astra` with reasoning, tool use, and code execution.
OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and
journal work. Context window size is not exposed by this session. Live
test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the
PR authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 15:52:19 -05:00
Nicky LeachandPaperclip e926b13017 fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path

> - Paperclip must stop remote execution after losing its controller.
> - A bridge can remain blocked while the sandbox still incurs costs or
performs actions.
> - Waiting forever for that bridge prevents provider termination.
> - A temporary provider outage can also exhaust cleanup attempts
permanently.
> - This pull request bounds bridge drain and persists cleanup retries
with backoff.
> - Cleanup ends only after provider confirmation, without a user
accepting uncertain side effects.

## Linked Issues or Issue Description

Builds on merged #13285. Related #13254 added exact provider termination
receipts; merged #13272 adds explicit user retry. Merged #13352 stops
active sandbox startup before waiting for setup. This PR preserves that
immediate cancellation path and extends bounded teardown to ordinary
release and destroy. Cleanup continues automatically after repeated
provider failures. Refs #12953 for provider failures blocking execution.

**What happened?**
Daytona release waits for in-flight bridge activity before stop/delete.
A dead bridge can prevent that wait from finishing. The host also stops
cleanup after five failed attempts.

**Expected behavior**
Provider termination proceeds after a bounded bridge drain. Cleanup
retries survive service restarts and provider outages.

**Steps to reproduce**
Start a sandbox command whose bridge promise never resolves, then
release its lease. Separately, persist a pending-cleanup lease with five
failed attempts and recover the provider.

**Deployment mode**
Hosted Paperclip with a Daytona provider; rebased onto master at
`728f7185f` on September 14.

## What Changed

- Bound bridge drain and provider lifecycle calls. Prefer stop for
reusable sandboxes, with delete fallback.
- Persist cleanup attempt identity, renewable in-flight deadline, and
cooldown. Fence completion writes against superseded attempts.
- Preserve scoped explicit Retry and its activity log. Explicit Retry
can skip cooldown, but cannot take over a live cleanup attempt.
- Continue cleanup after five failures with slower retries and an
operator warning.
- Exclude leases in cooldown before paging so they do not starve due
work.
- Add hung-bridge, restart, provider-recovery, and concurrent-cleanup
regressions.

## Verification

- Rebased onto master at `728f7185f`. The outstanding diff contains only
cleanup changes; the merged controller-ownership prerequisite is
excluded.
- Daytona plugin suite: 160 passed, including immediate startup
cancellation, graceful release, hung activity, and teardown regressions.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31
passed. Two added integration cases verify explicit Retry during
cooldown and while another cleanup owns the lease. They also verify run
scoping and the activity log.
- Targeted cleanup and cancellation cases in
`environment-runtime.test.ts`: 20 passed.
- Earlier live disposable Daytona test: provider stop ended background
work, resume preserved files without restarting the old process, a new
command succeeded, and the sandbox was deleted. This verifies provider
behavior; it was not repeated for this rebase.
- Latest-head CI and automated review are pending. Broad local tests,
typecheck, and build were not rerun for this focused rebase; CI supplies
those checks.

## Risks

- Timing out bridge drain permits provider termination; it never
supplies a stop receipt.
- A crashed cleanup attempt remains protected for 15 minutes, then
becomes eligible again. Repeated failures retry every 30 minutes after
escalation.
- The existing counter saturates at the escalation threshold; the new
attempt identity and deadline prevent overlapping claims.
- No schema, UI, telemetry, lockfile, or workflow change.

## Model Used

OpenAI GPT-6 through Codex, using reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 10:01:57 -07:00
mouse-value-add fcdb3f2499 feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents that do research work need live information from the web
> - Paperclip reaches external systems through governed, catalog-based
MCP connections
> - The Apps catalog is data-driven: a researched provider with a hosted
remote MCP server becomes a connectable app with no runtime code change
> - You.com operates a hosted remote MCP server for web search, content
extraction, and research tools
> - The server supports OAuth 2.1 with dynamic client registration, an
API key in a bearer header, and a keyless free profile at a separate
endpoint
> - This pull request adds You.com to the self-serve MCP research ledger
and generates its catalog entry with three connection methods: browser
sign-in, API key, and the keyless free profile
> - The benefit is that an operator can give agents live web search
through the normal connection governance, and the free profile needs no
account at all

## Linked Issues or Issue Description

No public issue exists for this provider. The problem description
follows the new-adapter issue template.

**Agent or provider**

You.com — web search and research tools over a hosted remote MCP server.

**Why this adapter is useful**

Agents that do research, monitoring, or fact-finding tasks need current
web results. You.com exposes web search (`you-search`), live page
extraction (`you-contents`), citation-backed research (`you-research`),
and finance research (`you-finance`) as MCP tools. Any Paperclip company
can connect it in a few clicks. The free profile offers `you-search`
without an account, so a new company can try agent web search at zero
cost and zero setup.

**How the agent is invoked**

Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`.
Three supported access paths, verified against the live server on
2026-09-16:

- OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate`
challenge with RFC 9728 protected-resource metadata and advertises a
dynamic client registration endpoint, so Paperclip's automatic DCR path
applies.
- API key. Sent as an `Authorization: Bearer` header per the provider's
official server manifest and docs. Keys come from you.com/platform and
unlock higher rate limits plus the full tool set.
- Keyless free profile at `https://api.you.com/mcp?profile=free`.
Provides a reduced, read-only tool set.

Official docs: https://you.com/docs/build-with-agents/mcp-server

**Are you willing to implement it?**

Yes. Implemented in this pull request.

**Additional context**

Research evidence collected 2026-09-16, from live protocol probes and
official provider sources only:

- Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with
`WWW-Authenticate: Bearer
resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource"
scope="Tools offline_access"`.
- RFC 9728 metadata lists one authorization server with scopes `Tools`
and `offline_access`.
- The authorization-server metadata (RFC 8414) publishes authorization,
token, and revocation endpoints, and advertises a
`registration_endpoint`, so DCR is available. No registration was
performed during research, per the runbook's non-registering preflight
rule.
- The keyless free profile answers `initialize` (server `You.com`,
version `4.0.1`), lists the tools `you-search` and `you-discover`, and
executed both tools successfully during the probe.
- The API-key placement matches the provider's official `server.json` in
the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer
<key>`.

## What Changed

- Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve
MCP research ledger in
`packages/shared/src/self-serve-mcp-research.json`, and refreshed the
ledger verification date.
- Added the You.com category (`ai`) and API-key header spec to
`scripts/ingest-app-definitions.mjs`.
- Added a You.com case to `specialMethodsFor` that emits three methods:
browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer
header), and keyless free profile (`mcp-free`, no auth).
- Regenerated `packages/shared/src/app-definitions/youcom.json` and the
generated registry via the ingestion script (`--definitions-only` mode;
no unrelated provider churn).
- Added the official You.com wordmark artwork (light and dark theme
variants, taken from the provider's docs site) under
`ui/public/brands/apps/`, with a manifest entry.
- Updated `packages/shared/src/app-definitions.test.ts`: ledger counts
(47 providers, 44 candidates), store count (48), verification date, and
assertions for the three You.com methods and their endpoints.

## Verification

- `node scripts/ingest-app-definitions.mjs --definitions-only` — passed.
Generated the new definition and registry import only; no other provider
JSON changed.
- `node scripts/check-app-brand-assets.mjs` — passed (71 identities).
- `node --test scripts/app-brand-validation.test.mjs` — passed.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests).
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-connection-removal.test.ts
ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server
suites that require embedded Postgres skipped on this machine by their
own environment gate; the gate is unrelated to this change.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed;
builds `@paperclipai/shared` with the new definition.
- `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592
skipped. Every failure is environmental on this container: the
embedded-Postgres suites refuse to start because the machine runs as
root, the native runtime suites need the Rust runner binary that this
container cannot build, and one media suite needs a native HEIC binary.
No failure touches the app-catalog, connection, branding, or
shared-package surface; those suites pass locally. CI is the
authoritative gate for the full suite.
- `pnpm --filter @paperclipai/server typecheck` — not completed: the
script's `prepare:runner-vendor` prelude builds the Rust runner, which
cannot build on this container. A direct `tsc --noEmit` reports only
pre-existing errors from the missing vendored runner types; no error
touches this change. No server code is changed.
- Live You.com proof on 2026-09-16 (keyless free profile, real network
calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR
endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned
results ✓, `you-discover` call returned results ✓.
- Live proof NOT run: an authenticated OAuth connect and an API-key call
against the full server. This environment has no You.com account or API
key. Per the runbook, this proof stays outstanding and must not be
assumed from the keyless probe. Both paths match the reviewed
`mcp_remote` patterns (DCR and bearer header) used by existing
providers.
- Browser e2e suites not run: opt-in per `AGENTS.md`, and this change
adds catalog data only, with no UI code.

## Risks

- Low risk. The change is catalog data plus generated output. It adds no
runtime code and touches no existing provider.
- The free-profile method is a fixed keyless endpoint. If You.com
changes or removes `?profile=free`, that method breaks and the entry
needs a ledger update. The OAuth and API-key methods do not depend on
it.
- The authenticated tool catalog is discovered live at connect time, so
provider-side tool changes appear through the normal catalog refresh and
quarantine flow, not through this definition.
- Rollback is a single revert; no migration and no state are involved.

## Model Used

- Provider: Zhipu AI, via OpenRouter
- Model: GLM-5.3 (`z-ai/glm-5.3`)
- Context window: 200K tokens
- Capabilities used: tool use (shell, file edits, live HTTP probes),
long-context repository reading
- The change was produced with AI assistance and reviewed by a human
before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-17 10:01:49 -07:00
DottaandPaperclip e26d787928 Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must continue tasks using user answers without losing earlier
requirements or approval gates.
> - The wake prompt mixed human decisions with prior tool evidence and
repeated detailed question instructions.
> - Those instructions belong with the question tool, with a short
routing hint in the wake.
> - The Runner evals need to prove that answers, approvals, and
completed work survive later turns.
> - This PR shortens the prompts, separates authenticated answers, and
adds continuation tests with useful screenshots.

## Linked Issues or Issue Description

Refs #13517. This is a follow-up to the merged onboarding skill and
Runner E2E work. Related #13539 covers responses received while a run is
active; this PR preserves its cases and adds continuation coverage.
Existing continuation/recovery and question PRs were searched; none
covers this prompt/documentation and eval change.

**What existing behavior does this improve?**

The instructions sent when an agent continues a task, the native
human-input tool documentation, and the evidence captured by Runner
full-stack E2E.

**Current behavior**

The wake repeats a long question-tool guide. Human answers appear
alongside untrusted prior results. Screenshot capture can finish at DOM
load while the task still shows a spinner, even when backend behavior
checks pass.

**Proposed behavior**

Keep earlier requirements unless the user changes them. Treat
clarification as distinct from approval. Give authenticated human
responses a scoped field. Keep tool and agent results as evidence. Put
detailed question behavior in the tool descriptor and retain one routing
sentence in the native wake. Wait for the correct task and loaded
conversation before taking screenshots.

**Reason and benefit**

Reduce repeated prompt text and make authority boundaries clear. Test
that real question cards, later answers, approval gates, and completed
child tasks still work. Make screenshots useful for human review.

## What Changed

- Shorten shared continuation instructions for legacy and native
runners. Separate authenticated user responses from tool results and
agent summaries.
- Remove the detailed question guide from native wake prompts. Keep its
behavior in the canonical `request_human_input` descriptor and existing
payload schema. Regenerate semantic contracts and fixture hashes.
- Add five continuation cases across four local profiles. Add a
dedicated choice-then-text case for native Codex and native Claude. All
22 cells join the shared full E2E campaign.
- Cover revised scope, clarification without approval, hostile
instructions in a handoff file, and reuse of a completed child after
restart. Keep production instructions and fixed user facts.
- Capture continuation screenshots only when the intended task and
conversation have rendered. Add provider-free browser regressions for
loaders and wrong-task capture.
- Preserve current master’s extra tool and onboarding cases. The default
campaign now contains 166 cells; 35 manual everyday cells remain
separate.

## Verification

- `pnpm -r typecheck`: passed after replay on current master.
- `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed.
- `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
test:e2e:runner:browser-support`: 4 passed. These tests failed against
immediate screenshot capture and passed after the fix.
- Focused continuation and native-input tests: 36 passed locally. The
tool-authority suite could not initialize embedded PostgreSQL locally,
including one isolated retry; its 17 assertions did not run locally. The
full remote server shards passed on this PR commit.
- `pnpm build`: passed after replay on current master. `pnpm test:run`
was attempted locally but hit the same embedded PostgreSQL
initialization failure; the remaining local run was stopped after
complete remote CI passed. This is not claimed as a full local test
pass.
- [Full PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685):
passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All
server/chat/workspace/serialized shards, browser shards, Runner checks,
typecheck, build, canary and policy checks passed. The isolated native
Runner build and security checks also passed: 57 successful checks, with
two expected Storybook skips.
- Greptile reviewed the exact PR head at 5/5, with no findings or
unresolved review threads. The PR has no merge conflicts.
- [Live question-docs
report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/):
3/3 passed at source `83dd132f2` before replay on master. Native Codex
and Claude each asked a choice, waited, asked a text question, and saved
both answers. Claude also passed a completed-child restart case. All
three native turns are checked for absence of the old question block.
- [Earlier continuation
report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/):
all five continuation cases passed on native Claude. The report retains
campaign and revision provenance and separately shows two unresolved
onboarding behavior failures.
- [Before/after prompt
report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/):
full text, current recorded Claude inputs, and reproducible
reference-token counts. The controlled wake comparison removes 401
reference tokens; the net counted input reduction is 339 after charging
the larger tool description. These are text-size estimates, not measured
billing savings.

## Risks

- Prompt wording affects model behavior. Live results cover the stated
cases, not every provider or conversation. Legacy profiles are
registered but were not rerun for this change.
- The optional continuation field changes prompt data only; there is no
database migration or new production API.
- Authenticated answer projection excludes generated summaries and
agent-resolved interactions. It preserves the answer’s question or
approval scope.
- The screenshot guard can expose UI loading failures that earlier runs
hid. Backend grading alone no longer makes those captures valid.
- The two prior onboarding failures remain separate product issues: work
before acceptance and a missing saved plan. This PR does not claim the
entire onboarding suite passes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, code execution,
and browser verification. The exact deployed model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted tests above; the
full local database-startup limit is documented
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-17 09:32:54 -05:00
5d9b20ccf0 fix(ui): keep task composer available while pause state loads (#13562)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task messages enter through the shared task composer.
> - The task page waits for a separate tree-control query to find active
pause holds.
> - The old loading guard disables the composer until that request
settles. A slow or stalled request blocks messages on both desktop and
mobile.
> - This PR allows sends while that query is pending. The server still
rejects paused board messages before saving a comment or waking an
agent.
> - Query failures and known pause holds still block the composer.
> - Regression tests cover pending sends, successful responses, query
failures, and late root or inherited pauses.

## Linked Issues or Issue Description

Fixes: #13561

Related: #13569 adds pending/error coverage for the same fix. This PR
now covers those cases and the late-pause transitions.

The report overstates two details: `isPending` clears after a successful
response, and the shared loading guard also affects mobile. The
reproduction holds the request pending. It does not prove why the
reporter's request remained unresolved.

## What Changed

- Preserve the original two-line fix that removes the pending-state
composer block.
- Add eight page regression cases using a real QueryClient and a
controlled API promise. Cover desktop and mobile, submission before the
query completes, successful resolution, rejected resolution, and late
root/inherited pause responses.
- Repair an existing server CI failure in a separate commit. Reuse the
shared unique-violation helper so Drizzle-wrapped duplicate inserts
become retryable document conflicts. Add a deterministic regression and
preserve unrelated database errors.
- Repair an existing Inbox test race in a separate commit. Wait for
workspace metadata, which resolves independently of the task list.

## Verification

- Red: restore the pre-fix `IssueDetail.tsx` and run the new `composer
tree control` cases. Both desktop and mobile pending-send cases fail
with `Checking task status…`. The other six cases pass.
- Green: restore the original PR fix. All 343 tests in IssueDetail,
TaskChatThread, and TaskChatComposer pass.
- The server comment/reopen and artifact-review suites pass (179 tests).
They include POST and PATCH pause checks that return 409 before any
comment, task mutation, or wakeup.
- The document error regression fails before the shared-helper fix and
passes after it. The document, artifact-review, and database-error
suites pass (31 tests). The handler matches only the issue-document key
constraint; revision and other constraint errors retain their original
identity.
- The Inbox suite passes (27 tests).
- `pnpm check:token-gates` passes.
- `pnpm -r typecheck` and `pnpm build` pass locally. The full CI matrix
passes on `5af1ed0b43269247aaba406cf4fd4d7fe1a22e75`: 54 successful
checks and two opt-in Storybook skips, including all general/serialized
tests, runner checks, release checks, and eight browser shards.
- One server shard initially hit an unrelated `EADDRINUSE` on test port
52000. Its single rerun passed without code changes.
- Greptile completed successfully on that exact commit with 5/5 and no
outstanding findings or review threads.
- The serial local `pnpm test:run` was started, then stopped after the
complete parallel CI matrix passed. It is not claimed as a completed
local full-suite run. The focused local suites above did finish
successfully.

For a manual reproduction, delay the task's `/tree-control-state`
response, open the task, and enter a message. Send should remain
available during the delay. Resolve the response with an active pause
hold and confirm that the pause takeover replaces the composer. Reject
the request and confirm that the error blocks sends.

## Risks

A user can attempt a send before the pause response arrives. The server
remains authoritative and returns 409 for a paused task before saving or
waking work. The known pause takeover and query-error block remain.
There are no schema, API, or styling changes.

The document change restores the existing conflict/retry behavior for
wrapped database errors. It does not retry unrelated database failures.
The Inbox change affects test synchronization only.

## Model Used

- Original fix: Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k
context, tool use and code editing, as reported by the author.
- Review, regression tests, and CI repairs: OpenAI GPT-6 (`gpt-6-astra`)
through Codex, with reasoning, tool use, and code execution. The session
does not expose its context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: austinpilz <austinpilz@users.noreply.github.com>
Co-authored-by: Dotta <bippadotta@protonmail.com>
2026-09-17 08:40:40 -05:00
Devin FoleyandPaperclip 165b10bd98 fix: enable GitHub Actions MCP toolset (#13553)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The GitHub connector lets agents use repository tools through MCP.
> - GitHub excludes Actions from its default MCP toolsets.
> - Approval of Actions permissions therefore does not make workflow
tools appear in Paperclip.
> - This pull request adds Actions to the requested toolsets for
discovery and execution.
> - Users can refresh existing connections and use workflow tools under
the existing access rules.

## Linked Issues or Issue Description

**What happened?**

GitHub Actions tools remain absent after the GitHub App receives Actions
read/write access and the user refreshes actions in Paperclip. Paperclip
does not request the Actions MCP toolset.

**Expected behavior**

Authorized GitHub connections expose workflow tools, including
`actions_run_trigger` with `method: "run_workflow"`, so agents can
dispatch an existing release workflow.

**Steps to reproduce**

1. Connect GitHub to Paperclip with access to a repository that has a
dispatchable workflow.
2. Grant the GitHub App Actions read/write permission and approve the
installation update.
3. Refresh the connection's actions in Paperclip.
4. Observe that the workflow tools are absent.

**Paperclip version or commit**

Base commit: `fae698031`.

**Deployment mode**

Hosted instance with a managed GitHub connection. The same missing
header affects PAT connections.

No matching public issue or pull request was found in the duplicate
search.

## What Changed

- Send `X-MCP-Toolsets: default,actions` through the shared GitHub MCP
header helper. This covers discovery, refresh, and execution for
existing and new managed or PAT connections, including legacy rows
identified through `transportConfig`.
- Test catalog refresh, tool risk classification, and workflow dispatch
through a mock MCP server.
- Document tool names, workflow arguments, required GitHub permissions,
and the refresh step.

## Verification

- Passed both affected test suites: `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/tool-gateway.test.ts` (391 tests).
- Passed `pnpm check:token-gates` and `git diff --check`.
- Live provider check: the default catalog returned 45 tools.
`default,actions` returned 49 tools, with no tools removed. The four
added tools were `actions_get`, `actions_list`, `actions_run_trigger`,
and `get_job_logs`.
- Live `actions_get` / `get_workflow` call succeeded. No workflow was
dispatched during live verification.
- Passed `pnpm -r typecheck` and `pnpm build` with the existing Rust
toolchain added to PATH.
- Rechecked server typecheck and build after the legacy-connection fix;
both passed.
- The full local test run has reported three skills-cache failures in
`company-skills-service.test.ts`. All three reproduce on the untouched
base commit (`fae698031`) on this macOS host: runtime-cache directory
renames fail with `EACCES`. The full run remains in progress.
- Greptile: 5/5 on `7d391e3c7`, with no unresolved review threads.
- After deployment, use **Refresh actions** on an existing GitHub
connection and verify the workflow tools appear.

## Risks

- Refreshed GitHub catalogs expose more tools. Existing access,
approval, and quarantine rules still apply. `actions_run_trigger` keeps
GitHub's destructive classification because it also supports
cancellation and log deletion.
- GitHub still enforces token and installation permissions. Dispatch
requires Actions write permission and a workflow with
`workflow_dispatch`.
- No database migration or saved connection edit is required.

## Model Used

OpenAI GPT-6 through Codex, with code execution and tool use. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 17:10:34 -07:00
Devin FoleyandPaperclip fae6980310 revert(apps): restore Google connector visibility (#13552)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - PR #13551 temporarily hid Google connectors.
> - We now want to restore their catalog visibility.
> - This PR reverts that change and restores the previous catalog
behavior.

## Linked Issues or Issue Description

Refs: #13551

Revert the temporary removal of Google connectors from the UI.

## What Changed

- Restore Gmail and eight Google Workspace entries to the catalog.
- Restore the matching branding flags and original catalog and service
tests.
- Remove the temporary-hiding documentation note.

This is an exact revert of commit
`cf1e873ab24277d55ffd3ab06074f77014dc4015`.

## Verification

- Passed: 507 catalog, UI, and connection service tests.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Full local Vitest was not repeated because the unchanged base has
confirmed macOS skill-cache permission failures. The full CI suites
passed.
- Passed: all GitHub CI gates; Greptile 5/5 on commit
`4e3dddef0ebfef1f99001e7735822ed4cba852ab`, with no review threads.
- Reviewer check: open Connectors and confirm that Gmail and Google
Workspace entries appear again.

## Risks

Low risk. This restores the previous catalog visibility and setup entry
points. Connector implementations and saved connection data are
retained.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:35:16 -07:00
Devin FoleyandPaperclip cf1e873ab2 fix(apps): temporarily hide Google connectors (#13551)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - We need to temporarily remove Google connectors from the UI.
> - The catalog already separates visibility from retained definitions.
> - This PR uses that setting so Google can return with a small change.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The Connectors catalog and its setup entry points.

**Current behavior**
The catalog shows Gmail and eight Google Workspace connectors.

**Proposed behavior**
Temporarily hide those nine entries. Keep their definitions and existing
connections.

**Reason and benefit**
Make the temporary UI removal easy to reverse.

**Breaking changes**
Fresh catalog setup no longer offers Google. Saved connections keep the
existing management and reconnect paths.

## What Changed

- Add the nine Google connector slugs to the existing hidden list.
- Match the branding manifest visibility flags.
- Update existing catalog and service tests. Keep backend Google
connection coverage and document how to restore visibility.

## Verification

- Passed: 507 targeted tests covering catalog definitions, URL matching,
setup routing, connector UI, branding, and the connection service.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Stopped the full local Vitest run after skill-cache permission
failures. Three failures in `company-skills-service.test.ts` also
reproduce on the unchanged base branch. The final connector service
suite passes all 319 tests.
- Greptile: 5/5 on the current commit, with no open review threads. CI
is retrying one unrelated preview-server readiness timeout. That test
file passes all seven tests locally.
- Reviewer check: open Connectors in a company with no Google
connections. Gmail and Google Workspace entries should be absent.
Existing saved connections remain manageable.

## Risks

Low risk. This uses the existing catalog visibility mechanism. No
connector implementation, credential, or database schema is removed.
Restoring visibility requires updating both the hidden list and branding
manifest.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 15:04:46 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:44:36 -05:00
DottaandPaperclip 11921075a4 Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first task helps a new user define and approve useful work.
> - That workflow needs reusable instructions and tests against the
production experience.
> - Native Codex and Claude must load the assigned skill, including
after resume.
> - Maintainers need recorded conversations and precise failed checks to
judge regressions.
> - This pull request adds the first-task skill and a suite in the
shared Runner E2E harness.
> - It keeps behavior results separate from informational quality scores
and incomplete recordings.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The first onboarding task and the Runner E2E report used to review it.

**Current behavior**

Onboarding embeds its policy in a hidden brief. Native Codex drops the
skill-instructions setting at the Rust boundary. The shared E2E harness
has no onboarding suite or full conversation view.

**Proposed behavior**

Assign and invoke `/first-task` for the onboarding task. Send selected
Codex skills as structured protocol inputs. Run twelve scenarios across
legacy Codex, legacy Claude, native Codex, and native ACPX Claude.
Include all 48 cells in full campaigns. Show recorded chat, question and
approval cards, exact checks, instructions, and billing in the shared
dashboard.

**Reason and benefit**

Measure the real onboarding experience before changing prompts.
Distinguish infrastructure failures, behavior failures, and unexercised
journey steps.

**Breaking changes**

No database migration or production API change. First-task instructions
now live in an assigned skill. The user-edited persona is preserved; the
skill includes the maintainer-approved proposal-mode mapping and
saved-plan requirement.

Related: #11043 is earlier onboarding work. #13422 already fixes native
Claude model pinning, context delivery, and read permissions on master;
this branch includes those fixes through its base. The new Claude
recovery test supplements them.

## What Changed

- Extract and assign the first-task skill while retaining the production
greeting and opening question.
- Carry the Codex skill-instructions flag through thread start and
resume. Resolve explicit task skill references only against assigned
skills and send native skill inputs.
- Invoke an unambiguously selected assigned skill through Claude ACPX’s
native slash-command parser on initial and resumed turns, retaining the
entire task/wake envelope as its argument. Do not carry that invocation
into ordinary tasks.
- Restore the saved single-task proposal modes: confirmation card, or
saved plan with revision-targeted checkbox approval. Explicit plan
requests also require a saved plan.
- Add first-response and complete-journey cases with fixed user facts,
acceptance checkpoints, durable outcome checks, and accounting for child
runs.
- Fail the eval when choice questions have fewer than two real options.
Recognize planning documents without treating them as completed work.
- Add optional, bounded quality judging as explicit post-processing.
- Render full conversations and static interaction cards in the shared
report. Conversations start folded. Show original and regraded results
and incomplete journeys distinctly.
- Keep credential-persistence scanning outside the first-task behavioral
suite; retain public evidence redaction.
- Refresh generated capability references after the API-reference edits.
- Correct shared native question guidance and tool schemas: choices need
at least two meaningful options; open-ended questions use canonical text
fields with the required compatibility payload. Verify both formats
through real tool-authority persistence.
- Disable announcements automatically for every isolated Runner E2E
process and label the gallery environment/provider/target explicitly.
- Remove CI races in the GitHub connection browser test and native
session recovery test by waiting for the actual async work before
asserting its results.

## Verification

- `pnpm exec vitest run
server/src/services/onboarding-first-task-assets.test.ts
server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19
passed.
- `pnpm --dir packages/paperclip-runner exec vitest run
src/drivers/acpx/runtime-host.test.ts
src/drivers/acpx/native-skill-prompt.test.ts
src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command
forwarding and the 1 MiB input boundary both failed before their fixes
and passed afterward. Coverage includes changed skills on reopen,
approval context, and an ordinary subsequent task.
- Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64
first-task fixture and grader tests also pass.
- Full repository typecheck and build passed locally. Server typecheck
and Runner build passed again after the native-command change.
- Full GitHub Actions CI passed on `23e56447b`: all
server/workspace/browser shards, Runner verification, typecheck/release
registry, build, canary, policy, and Docker checks. Greptile reviewed
this exact head at 5/5 with no unresolved threads. The earlier broad
local run had database startup/timing failures that passed isolated
retries; the complete remote suite is green.
- Merge verification against current master: 312 harness tests and 13
native recovery tests passed. Regenerated semantic contracts and fixture
hashes pass their consistency check. Full local typecheck and build also
passed on the stacked queue branch. After merging the latest master and
preserving the GitHub setup timing regression in the split browser
suite, both focused GitHub browser tests passed. Three CI timing/startup
flakes passed local verification and one remote retry; all latest-head
checks are green.
- Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local
mock API confirmed that `/skill-name` expands the assigned skill body
before the model request and retains the task arguments. A prose mention
does not. The probes made no paid model calls. The ACP probe used the
current first-task skill body and retained the wake arguments.
- [Full 48-case campaign and
report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/):
44 passed after three interrupted Codex cases completed in targeted
reruns. Original results, regrades, and all 51 executions remain in the
report provenance.
- [Claude campaign after the shared-question
fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/):
10/12 passed with zero single-option failures. All 12 recorded the
current assigned skill and corrected guidance. The failures exposed
skipped skill invocation and a missing saved plan. This PR adds native
command invocation and explicit saved-plan instructions; the subsequent
report below still shows behavior failures.
- [Fresh 12-case Claude
report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/)
at `78452129e`: 10/12 pass after correcting two false proposal-matcher
failures. The recordings said “Here is the task I will create and
run/complete” in approval cards; the old matcher missed that word order.
Regression tests failed before the fix and pass after it. Original
results and offline regrade provenance remain linked. No agent rerun was
needed. Zero single-option-question failures; two behavior failures
remain: direct work before acceptance on a plain first message, and an
explicit plan request without a saved plan. Neither check was relaxed.
The follow-up `82087ac7e` fixes command-prefix size accounting;
`94aefb1f3` fixes only that proposal matcher.
- Report browser checks confirm folded conversations, rendered cards,
explicit Local/Daytona labels, and no page errors. The published-object
audit scanned 1,306 text files across 2,154 objects with no
credential-format findings or prohibited files. Image pixels and unknown
token formats are outside that scan.

## Risks

- Model behavior is nondeterministic. One campaign is evidence, not a
guarantee. The two remaining Claude behavior failures are visible in the
report and require further product work; this PR does not claim all
onboarding scenarios pass.
- The suite checks persisted Paperclip effects. It cannot prove the
absence of arbitrary external effects.
- Historical recordings can miss later journey steps. These remain
incomplete, never passes.
- Native profiles switch runtime after the production onboarding wizard
because it does not yet expose a native option.
- Quality scores are informational and cannot override behavioral
failures.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, and code
execution. The exact deployed model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:24:32 -05:00
Devin Foley dcb04a8062 fix(claude-local): read a macOS isolated login from its suffixed Keychain item (#13519)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Connecting a Claude subscription during onboarding uses an isolated
login: the wizard points `claude` at a per-connection
`CLAUDE_CONFIG_DIR` and then verifies the credential before saving the
connection
> - The verifier reads `.credentials.json` from that directory — but on
macOS, Claude Code does not write a credentials file at all: it stores
the OAuth credential for a custom config dir in a per-directory Keychain
item named `Claude Code-credentials-<first 8 hex chars of sha256(dir)>`
> - So on macOS the connect step can never verify a successful sign-in,
and onboarding dead-ends at "Could not verify the local subscription"
(Linux works because the CLI falls back to writing the file there, which
is why the Docker-based smokes pass)
> - This pull request teaches the credential readers to consult the
login home's own suffixed Keychain item when the file is missing
> - The benefit is that macOS self-hosted users can connect a Claude
subscription during onboarding, while the standing isolation invariant —
an isolated login must never fall through to the machine-level operator
login — is preserved, because only the per-directory suffixed item is
ever read

## Linked Issues or Issue Description

No existing issue. Description follows the bug-report template:

**What happened?**
On macOS, connecting a Claude subscription during onboarding (or from
Connections) always fails with "Could not verify the local subscription.
Run the sign-in command shown for this connection, finish signing in,
then try Connect again" — even after `claude auth login` completes
successfully in the isolated `CLAUDE_CONFIG_DIR`.

**Expected behavior**
After finishing the browser sign-in for the printed command, clicking
Connect verifies the subscription and saves the connection.

**Steps to reproduce**
1. On macOS, run onboarding on a fresh instance and reach "Connect a
model" → Claude → Subscription.
2. Run the printed `export CLAUDE_CONFIG_DIR=… && claude auth login`
command in a terminal on the same machine and complete the browser
sign-in.
3. Return and click Connect. Verification fails every time. Inspecting
the isolated directory shows `.claude.json` with a fully populated
`oauthAccount` but no `.credentials.json`; `security
find-generic-password -s "Claude Code-credentials-<suffix>"` shows the
credential landed in the Keychain, where the verifier never looks.

**Paperclip version or commit**
Reproduced on `2026.915.0-canary.11` (`dffc2b3ca`) with Claude Code
2.1.231.

**Deployment mode**
Self-hosted, authenticated instance on macOS.

**Installation method**
`npx paperclipai onboard` (also affects any macOS install; Linux is
unaffected).

## What Changed

- `packages/adapters/claude-local/src/server/quota.ts`:
- New exported helper `readIsolatedClaudeKeychainToken(loginHome)` —
computes the suffixed service name (`Claude Code-credentials-` + first 8
hex chars of `sha256(loginHome)`) and reads only that item via
`/usr/bin/security`; returns null off macOS
- `readClaudeToken` with a custom `CLAUDE_CONFIG_DIR` now consults that
directory's suffixed item after the file reads miss (previously it
refused the Keychain entirely for custom homes). The unsuffixed operator
item is still gated behind the explicit `allowKeychain` opt-in with no
custom home, unchanged
- `server/src/services/local-ai-credentials.ts`: for anthropic isolated
logins, fall back to the suffixed Keychain item after the hardened
credentials-file reads miss. The file path is untouched and still
preferred; the hardened file reader (`readLocalAiCredentialFile` with
its uid/mode/symlink checks) is not bypassed
- Tests: adapter keychain suite extended (suffixed lookup for custom
homes, no unsuffixed fallback when the suffixed item is absent,
off-macOS null); server verifier suite extended (keychain fallback when
the file is missing, file preferred over keychain, absent-login failure
still never touches the ambient reader)

Security note: the suffix binds each Keychain item to exactly one auth
home, so reading it can only surface the login performed inside that
home. The account-isolation invariant the old code enforced by refusing
the Keychain outright ("never substitute the server operator's login for
a user's isolated login") is preserved — the unsuffixed item is never
consulted for an isolated login, and a new test pins that.

The suffix derivation was confirmed against a live login on macOS: a
real `claude auth login` into an isolated home left no credentials file,
wrote the full `oauthAccount` to `.claude.json`, and created a Keychain
item whose suffix equals the first 8 sha256 hex chars of the exact
`CLAUDE_CONFIG_DIR` string; reading it back with the same `security`
invocation returned the live token, which the new code path then
verifies via the existing quota probe.

## Verification

- `pnpm exec vitest run src/server/quota-keychain.test.ts`
(claude-local): 10 tests pass; full claude-local suite: 287 passed, 1
skipped
- `pnpm exec vitest run src/__tests__/local-ai-credentials.test.ts`
(server): 11 tests pass
- Reverting only the verifier change makes the two new server tests fail
— the suite reproduces the live bug
- End-to-end on macOS: a dev server built from this branch, fresh data
dir, full onboarding walk with a real `claude auth login` into the
printed isolated dir — the connect step verifies and saves the
connection

## Risks

- Low. The change is additive and fail-closed: when the suffixed item is
absent (Linux, older Claude Code versions, no login performed), behavior
is byte-identical to today — the file reads run first and the failure
message is unchanged
- The `security` call runs with the existing 10s timeout and swallowed
errors, matching the established unsuffixed-item code path
- No migrations, no API surface changes

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:28 -07:00
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00
Nicky Leach ae06329971 test(server): hoist the route module graph in the issue ownership authz suite (#13524) 2026-09-15 22:53:44 -07:00
Nicky LeachandPaperclip e1f245a660 fix(server): recover sandbox leases stranded active after a restart (#13515)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server starts and stops provider sandboxes through environment
leases
> - A restart can leave a terminal run with an `active` lease
> - The normal orphan recovery path cannot find that lease after the run
ends
> - This pull request adds a bounded sweep that changes the stranded
lease to `pending_cleanup`
> - The existing cleanup sweep then stops the provider sandbox on the
same heartbeat tick
> - The benefit is that stranded sandboxes stop and do not continue to
create provider cost

## Linked Issues or Issue Description

**What happened?**

A restart can occur after the server writes a terminal run status but
before it releases the related environment lease. The lease then stays
`active`, and later recovery does not select it. A second path skips the
lease when its environment row does not exist.

**Expected behavior**

The heartbeat recovery path must find an `active` lease that no live run
can release. It must move that lease to `pending_cleanup`, and the
cleanup sweep must stop the provider sandbox.

**Steps to reproduce**

1. Start a run that owns a provider sandbox lease.
2. End the run and stop the server between the run-status write and the
lease-release write.
3. Restart the server and allow the heartbeat recovery sweep to run.
4. Confirm that the lease reaches `pending_cleanup` and the provider
sandbox receives a stop request.

**Paperclip version or commit**

This pull request targets the current `master` branch at the base commit
used for review.

**Deployment mode**

The change applies to local development and server deployments.

**Installation method**

Built from source with the repository test commands.

**Agent adapter(s) involved**

Not adapter-specific. The change applies to core heartbeat recovery.

**Database mode**

The change uses the existing database tables. It adds no migration.

## What Changed

- Add `sweepOrphanedActiveLeases()` to heartbeat recovery.
- Select only stale `active` leases that have no live run owner.
- Skip leases with a different live lease for the same provider
resource.
- Preserve retained leases and write a failure reason for recovered
leases.
- Limit each sweep to 20 rows.
- Run the recovery sweep before the pending-cleanup sweep.
- Add focused tests for the recovery guards and same-tick cleanup.

## Verification

- `pnpm vitest run
server/src/__tests__/heartbeat-orphaned-active-lease-sweep.test.ts` — 10
tests pass.
- `pnpm vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts` — 22 tests
pass.
- `pnpm --filter @paperclipai/server typecheck` — exits 0.
- The complete CI suite remains the final check for the repository.

## Risks

The sweep changes only stale `active` leases that no live run can
release. The stale threshold, live-resource guard, retained-lease guard,
and page limit reduce false recovery. The change adds no endpoint,
schema change, or migration.

## Model Used

OpenAI Codex, GPT-5, tool-enabled coding agent with repository
inspection, GitHub CLI, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 21:52:02 -07:00
a8d32e5e61 feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent work runs in sandboxes that provider plugins supply
> - Operators can choose a provider to run agent work
> - CreateOS adds another provider with workspace-preserving pause and
resume
> - This pull request adds a CreateOS provider plugin
> - The benefit is that operators can preserve a workspace between runs
without keeping its compute active

## Linked Issues or Issue Description

Refs #13203 and the earlier closed #13096.

This continues the CreateOS contribution from @bhautikchudasama and
@ashwaq06. The branch preserves the original implementation commit.
Thank you to both contributors.

When squash-merging, preserve the original author's credit in the squash
commit body:

```text
Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com>
```

The original fork rejects maintainer pushes. This branch includes the
merge-conflict resolution and review fixes. The request is described
below using `adapter_request.yml`.

**Agent or provider**

CreateOS sandbox API (https://api.sb.createos.sh).

**Why this adapter is useful**

CreateOS can pause a sandbox and resume it by ID. The workspace survives
the pause. This adds a reusable-lease option to the existing sandbox
provider system.

**How the agent is invoked**

Build and install the local plugin as described in its README. Open
Instance Settings, then Environments. Select the `createos` driver.
Supply an API key and shape. The driver then supplies sandbox leases for
agent runs.

**Are you willing to implement it?**

Yes. This pull request is the implementation.

## What Changed

- Adds the `createos` sandbox provider under
`packages/plugins/sandbox-providers/createos`.
- Calls the CreateOS HTTP API directly. The package adds no vendor SDK.
- Implements the environment lifecycle hooks, incremental process
output, and binary workspace sync.
- Registers the optional bundled provider and its trusted host
credential fallback. The fallback is limited to the official API origin;
custom endpoints require an explicit key.
- Lists the package in the release manifest with `publishFromCi: false`
until its first npm publish is bootstrapped.
- Waits through delayed pause/resume state updates without duplicate
action requests.
- Cancels queued API requests promptly while preserving request spacing.
- Uses direct CLI invocation in the setup guide so paths and IDs are
passed without an extra shell expansion.
- Includes current master and retains its existing Git-subfolder
containment fix.

## Demo

Fresh setup and a run against a CreateOS sandbox.


https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110


https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084

## Verification

All 25 jobs in [CI run
34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542)
passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including
typecheck, build, native runner verification, server and workspace
tests, browser tests, and the canary release dry run. Greptile reviewed
the same commit at 5/5 with no unresolved review threads.

GitHub reports no merge conflicts. The remaining merge gate is
code-owner approval for the new `package.json`, as required by
`.github/CODEOWNERS` and the `master` ruleset. Reviewers have been
requested automatically.

Local checks passed:

- Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke
skipped), and `pnpm build`.
- Host: focused credential and bundled-plugin tests (17 passed), plus
CLI invocation safety (39 passed).
- Release: package manifest check and release policy tests (18 passed).

The full local `pnpm test:run` attempt caught the README command issue;
its focused rerun now passes. The full local run stopped after its
general-server group: 7,804 tests passed, with unrelated embedded
PostgreSQL startup failures and 10 failures in unchanged
runtime-skill-cache tests (`EACCES` on directory rename on macOS). It
did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm
build` reach the runner package and stop because this machine has no
Rust/Cargo installation. The corresponding CI checks passed on
provisioned runners, as linked above.

The live CreateOS smoke requires explicit provider credentials and was
not run during this review. It is available with `CREATEOS_LIVE_TEST=1
pnpm test` in the provider directory. The author supplied the demo links
above.

## Risks

The provider is opt-in and is not installed by default. It is available
through a local-path install or explicit image inclusion. npm
publication remains disabled until a maintainer bootstraps the package
and enables publishing.

Sandbox creation has no idempotency key. An ambiguous create response
can leave a resource that requires provider-account inspection. Process
tracking is in memory; durable lease recovery belongs to the host. The
provider does not advertise guaranteed expiry, interactive login,
snapshots, duplex channels, or ingress. Live native-runner qualification
remains outside this PR's tested claims.

## Model Used

Original provider implementation: human-authored by @bhautikchudasama,
as reported in #13203. The original description reports Claude Opus 5
assistance.

Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review,
editing, and tool execution. The precise runtime model variant and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full-suite
environment limits documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:09:43 -07:00
DottaandPaperclip 9adeabb590 fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let people give agents access to external tools.
> - A personal connection needs the current user's authorization.
> - A new MCP URL must be probed before Paperclip can discover its
sign-in method.
> - Requiring a personal grant before that probe prevents sign-in from
starting.
> - This pull request permits the initial probe for a creator-owned
draft with no credentials.
> - People can complete personal setup while later requests retain
authorization checks.

## Linked Issues or Issue Description

**What happened?**

Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and
"This connection needs the current user's authorization". Paperclip
checked for a personal grant before contacting the provider. The health
wrapper also changed the expected authorization error into a server
error.

**Expected behavior**

Discover OAuth and start browser sign-in. Create a personal grant after
consent. For a public endpoint, discover its tools and create the empty
personal grant after a successful probe. Keep missing authorization on
later health checks as HTTP 422.

**Steps to reproduce**

1. Add an unknown remote MCP URL with no saved credentials.
2. Select "Just me".
3. Check the link. Before this fix, the request fails before sign-in or
tool discovery.

**Paperclip version or commit**

The three original regressions fail against `6cfe4acff` with the service
fix removed and pass with it restored.

**Deployment mode**

The defect was reported in production and reproduced in local server
tests with isolated PostgreSQL.

Related work: Refs #11831. Refs #11144. Searches found no duplicate fix.

## What Changed

- Allow an initial credential-free probe only for the creating user's
personal draft with unknown authentication and no supplied credentials.
- Leave OAuth grant creation to the callback. Create an empty personal
grant only after a public probe succeeds.
- Preserve the personal grant for URLs that already contain a
credential.
- Make empty personal grant creation conflict-safe without overwriting a
concurrent grant or duplicating its creation audit.
- Commit the empty grant and audit atomically. Retain the established
public draft identity after a later catalog failure, so a failed retry
cannot remove a successful retry's grant.
- Run the following catalog/default-profile step in its own transaction,
so failures discard partial catalog, profile, binding, and audit changes
without deleting the established identity.
- Verify archived personal connections retain their owner: another user
is rejected before probing, while the original owner can resume setup.
- Preserve `user_authorization_required` and HTTP 422 in health
failures.
- Cover the connect and OAuth callback routes, real loopback HTTP,
credential-bearing URLs, and later health checks.
- Document personal setup and the test fixtures.

## Verification

- Red/green: the original three tests failed with the exact reported
error before the fix and passed after it.
- The credential-bearing personal URL regression also failed before its
guard was added.
- Focused suite: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/generic-mcp-connection.test.ts
src/__tests__/tool-access-service.test.ts` passed all 385 tests across
the two suites. An earlier run had a socket hang-up in an existing
agent-permissions test; the unchanged suite passed on rerun.
- The concurrent rollback regression failed before its fix because the
successful retry's grant was deleted. It now verifies the grant and
draft survive and a later normal health check succeeds.
- Database fault injection during profile-entry insertion reproduced
partial catalog writes before the transaction fix. The regression now
verifies unchanged catalog rows, no partial profile/bindings, a retained
grant, and successful retry.
- CI's first serialized-server shard 3 attempt failed an existing
peer-agent mutation test (the real run-context guard ran despite the
test's mock). The test passed in isolation and all 108 tests in that
suite passed unchanged locally. The single failed-shard rerun passed
without code changes.
- Final-commit CI: all 32 applicable checks passed on
`639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected
tests, the full test matrix, browser suite, build, typecheck, release
checks, and security checks. The two Storybook-only checks were not
applicable and skipped. Greptile is 5/5 with all review threads
resolved. [Successful CI
run](https://github.com/paperclipai/paperclip/actions/runs/35027478353).
- `pnpm -r typecheck` passed.
- `pnpm smoke:mcp-fixtures -- --require-paperclip` passed.
- `pnpm build` passed.
- Full local `pnpm test:run` was attempted: its general-server group
finished with 12,372 passed, 4 failed, and 70 skipped tests. The run
started before review edits; its two MCP failures used the old cached
service (including an insert without the new conflict clause). All 385
focused tests pass on the final code. The other failures were existing
workspace-cleanup and runtime-port tests; their unchanged suites passed
on rerun (66 passed, and 25 passed/3 skipped). The local command stopped
before later groups. The final-commit CI matrix is the full-suite merge
gate; this local run is not claimed as green.

## Risks

- The initial probe must not become a general authorization bypass. It
is restricted to the creating user's draft. Normal health checks retain
authorization enforcement.
- Public endpoints get a personal grant with no secrets only after they
answer successfully. Credential-bearing URLs keep their existing grant.
- The concurrent-probe regression seeds the catalog and default profile
to isolate grant creation. Existing first-time catalog/profile creation
races are outside this change; this does not claim to make the entire
setup flow concurrency-safe.
- No database migration or UI change is required. OAuth tests use a
simulated provider; the public endpoint test uses real loopback HTTP.

## Model Used

OpenAI Codex, GPT-6-based assistant for regression tests and PR
preparation; a GPT-5-based Codex assistant assisted with the initial
implementation. Exact runtime model IDs and context-window sizes are not
exposed in this session. Both used reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:11:59 -05:00
Devin FoleyandPaperclip 544c3476a8 feat(server): wrap a bare Cloud UI snippet body in a <script> element (#13496)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A Cloud-managed instance injects an operator-owned HTML snippet
before `</body>` through `injectCloudUiSnippet`.
> - The Cloud control plane delivers that snippet to each instance as an
environment variable through a provider API.
> - The provider edge firewall now base64-decodes the request payload
and blocks any value whose decoded form contains a `<script` marker.
> - A working snippet needs a script tag, so every delivery is now
blocked and the operator cannot ship the snippet at all.
> - This pull request treats a resolved value that does not start with
`<` as a bare script body and wraps it in a `<script>` element at
injection time.
> - The benefit is that the operator can deliver a tag-free body that
the firewall passes, and the instance restores the script element on the
page.

## Linked Issues or Issue Description

No public issue exists. The problem is described below.

Related PRs (searched the PR list; none duplicate this change):
- Refs #13168 — added `injectCloudUiSnippet`, the mechanism this
extends.
- Refs #13245 — added the base64 `_B64` path on the assumption that
base64 clears provider WAFs. That assumption no longer holds; this PR is
the successor.
- Refs #13441 — the in-product feedback approach that the Cloud-owned
snippet replaced (closed).

**What happened?**

`injectCloudUiSnippet` injects `PAPERCLIP_CLOUD_UI_SNIPPET` (or the
base64 `_B64` form) verbatim before `</body>`. A working value must
therefore contain a `<script>` tag. The Cloud control plane delivers
this value as an environment variable through a provider API that sits
behind an edge firewall. The firewall now base64-decodes the payload and
rejects any value whose decoded form contains `<script`. The delivery
request fails, so the snippet cannot reach the instance.

**Expected behavior**

The operator can deliver the snippet through the provider API, and the
instance runs it.

**Steps to reproduce**

1. Build a snippet that contains a `<script>` element.
2. Deliver it to a Cloud instance through the provider variable API, in
plain or base64 form.
3. The firewall rejects the request. The instance never receives the
snippet.

## What Changed

- `injectCloudUiSnippet` now wraps a resolved value that does not start
with `<` in a `<script>...</script>` element. A value that already looks
like markup is injected byte-for-byte, so existing full-`<script>`
snippets are unchanged.
- Added two unit tests: one for the wrap path (plain and base64), one
that confirms `$&`-style replacement tokens in the body survive the
wrap.

## Verification

- `pnpm exec vitest run src/__tests__/cloud-ui-snippet.test.ts` — 17
pass.
- `pnpm exec vitest run src/__tests__/static-index-html.test.ts` — 2
pass.
- Confirmed the changed file has no type errors. The full `pnpm run
typecheck` needs the Rust runner toolchain (`cargo`), which is absent on
this machine; CI runs it in full.
- Verified out of band that a tag-free body clears the provider firewall
and wraps into valid, executable standalone JavaScript with no premature
`</script>` close.

## Risks

Low risk. The change adds a branch that only affects values that do not
start with `<` — previously injected as inert text, never as a running
script. Values that start with `<` keep their exact bytes. The content
is trusted operator HTML, consistent with the existing contract. Roll
back by reverting this commit.

## Model Used

Claude — `claude-fable-5` (Fable 5), extended thinking, with tool use
and code execution in Claude Code. A human author reviewed and verified
the change before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 15:09:19 -07:00
DottaandPaperclip d49f168381 fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must publish generated files so users can inspect their
results after a sandbox stops.
> - Legacy sandbox bridges blocked attachment listing and could not
carry multipart binary uploads through the queue transport.
> - The native runner has a separate verified file registration path
that needs the same durable result.
> - This pull request repairs legacy binary transport, makes native
download receipts explicit, and reveals new outputs in the task
Artifacts tab.
> - Users can open generated files from either runner without a
transport flag change.

## Linked Issues or Issue Description

Refs #13355 for the existing native file publication path. Related
filename fixes: #2615 and #4788. Related sandbox persistence work:
#13376. This change repairs attachment delivery through the existing
API; it does not add workspace persistence.

**What happened?**

The upload helper first lists task attachments to avoid duplicates. Both
legacy bridge allowlists rejected that GET request with 403. A direct
multipart upload also failed: the queue bridge accepted only JSON,
excluded attachment uploads, and converted bytes to UTF-8 text. Enabling
HTTP/2 alone did not fix the missing listing route. These failures
occurred before attachment storage.

**Expected behavior**

Both runners can publish a workspace file, register its work product,
bind it to a response, and return a working download. The file stays
accessible after sandbox deletion. A new output opens the task Artifacts
tab. The agent receives accurate errors and decides how to retry or
report a failure.

**Steps to reproduce**

1. Run a legacy agent in Daytona with the duplex bridge disabled.
2. Invoke the bundled upload helper with Bash on a PNG or PDF.
3. Repeat with the duplex bridge enabled.
4. Register the same file through the native runner with generic API
tools disabled.
5. Retry registration, delete the sandbox, and compare the downloaded
bytes with the original file.

**Paperclip version or commit**

The failing baseline was `f2c5e54dc`. This branch is rebased onto
`6cfe4acff`.

**Deployment mode**

Source checkout with a local API and real isolated Daytona sandboxes.

## What Changed

- Allow authenticated attachment listing, upload, and content download
through both legacy bridge transports.
- Add optional base64 body encoding to queue envelopes. Preserve the
existing UTF-8 contract when the encoding field is absent. Decode binary
bodies before forwarding them.
- Preserve multipart headers. Bound raw bytes, encoded envelopes, and
in-flight reservations. Retain timeout and uncertain-write behavior.
- Preserve helper deduplication and return structured uncertain-write
failures. Document explicit Bash invocation in live skills.
- Add attachment IDs and content/download paths to native registration
receipts. Reuse verified local and remote file reads, attachment
storage, work-product registration, and response binding.
- Preserve Unicode upload filenames and provide a valid
Content-Disposition header.
- Open the task Artifacts tab when new stored outputs arrive, including
a closed desktop panel or mobile drawer. Deduplicate upload and
registration events by object ID. Preserve manual selection on
refetches, edits, and panel remounts.
- Remove task artifact filters, the company Artifacts footer link, and
the unassigned group heading and timestamp.

### Screenshot

![Generated images and a document in the task Artifacts
tab](https://pages.paperclip.ing/sandbox-file-delivery-2026-09-15/artifacts-tab.png)

This is the local display fixture. The image was generated separately
and published through the attachment and work-product APIs.

## Verification

- Post-rebase `pnpm -r typecheck` and `pnpm build` pass.
- The post-rebase local `pnpm test:run` passed 12,369 tests before one
existing conversation reset test timed out; all 33 tests in that suite
pass when rerun with isolated test configuration. The aggregate command
stopped before its remaining groups. GitHub runs the complete suite in
separate shards.
- All [GitHub verification
checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350)
pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks
and two configured skips. The native-session recovery assertion
initially raced its fire-and-forget Sentry report; all 13 tests pass
locally, and the same-commit CI rerun passes all 170 suites (3,079
tests).
- Live post-rebase Daytona: all three file-delivery tests pass. They
cover the real Bash helper with the queue bridge, the helper with
HTTP/2, and native `register_deliverable` with generic API tools
disabled.
- Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate
registration, response binding, authorization controls, and
byte-for-byte downloads after sandbox deletion.
- Local focused coverage includes transfer bounds, malformed encoding,
interrupted transfers, remote path containment, and native file
verification. The attachment route suite passes all 32 tests, including
an eight-case filename-header matrix for Unicode and special characters,
inline and forced downloads, and full and partial responses.
- Browser verification confirms image previews, persisted downloads,
automatic Artifacts selection, and preserved manual selection after
edits and reloads. Desktop/mobile component coverage passes. The latest
UI cleanup passes its 10 affected tests and token gates.
- Coverage limit: the Daytona tests call the real helper and native
registration path directly. They do not replay a complete model-led
image-generation task through the browser.

Live command (requires a configured Daytona credential):

```sh
PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts
```

## Risks

- Binary queue bodies use more memory because base64 adds encoding
overhead. Transfer and process limits must remain aligned.
- An interrupted write can have an unknown result. The bridge reports
this state and preserves stable retry identities.
- New artifacts intentionally change the active task tab. Existing
history and repeated updates must not take focus again.
- Transport flag defaults, server authorization, frozen skill snapshots,
and completion policies remain unchanged. No schema migration is
required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser testing. The runtime does not expose a more
specific model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 15:46:57 -05:00
DottaandPaperclip f4cdc7b231 fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The control plane prepares task workspaces before it starts an
agent.
> - Workspace preparation reads Git state so it can preserve edits and
exclude private files.
> - A failed scan was treated as a non-Git folder and lost its actual
failure code.
> - The resulting generic setup failure could not recover, even when the
cause was temporary.
> - This pull request keeps the cause and uses the existing bounded
retry schedule before provider startup.
> - Tasks can recover without human intervention, while permanent
failures and exhausted retries stop with useful guidance.

## Linked Issues or Issue Description

**What happened?**

A Git scan error during managed repository preparation became
`Configured repository folder is not a Git checkout`, followed by
generic `setup_failed`. The agent never started. Generic recovery could
not distinguish a temporary timeout from a bad workspace configuration.

**Expected behavior**

Keep the closed scan error code. Retry temporary timeouts and queue
saturation under the existing shared budget. Preserve edits, exclusions,
ownership, and pause gates. Stop permanent failures and exhausted
retries with a specific explanation. Do not replay historical generic
setup failures.

**Steps to reproduce**

1. Configure a task project with a local Git source that must be copied
into its managed repositories.
2. Make the ignored-file scan exceed its timeout before the agent
starts.
3. Before this fix, the snapshot returns null and the run ends as
non-retryable `setup_failed`.
4. Use the disposable browser fixture in
`tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and
test the full recovery path.

**Paperclip version or commit**

Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`.

**Deployment mode**

Built from source. The defect is in core workspace setup, not a specific
model provider.

Related work: Refs #13442 (managed repository preparation), Refs #11572
(bounded Git scheduler), Refs #12997 (separate adapter startup retry
work), Refs #13469 (separate terminal-workspace scan performance work).

## What Changed

- Return the non-Git fallback only for repository discovery. Propagate
failed scans of a confirmed repository.
- Replace full ignored status output with an ignored-only directory
listing. Preserve NUL-delimited paths and exclusions.
- Preserve typed, sanitized scan errors through workspace preparation
and persist pre-provider failure details.
- Retry only timeouts and queue saturation, using the existing durable
two-retry budget and issue gates. Prevent generic recovery from adding
another budget.
- Show workspace-specific failure copy and actionable exhausted-recovery
notices.
- Add red-green unit tests, real-database restart and retry-boundary
tests, and opt-in browser acceptance fixtures with real Git subprocess
timeouts.
- Document the recovery contract and browser verification procedure.

## Verification

- Red: injected scan failures returned null instead of rejecting; setup
lost the timeout code; task-thread and recovery notices had generic
copy.
- Green: 119 focused adapter/backend tests, 20 recovery-boundary tests,
and 136 task-thread tests.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed on the final production code.
- `pnpm check:token-gates` — passed.
- The initial local `pnpm test:run` overlapped source edits and was
interrupted after two late-added assertions saw pre-fix behavior; it is
not counted as a green full run. A fresh final-head run passed all 229
tests across the six affected adapter/backend/UI suites. The clean
latest-head CI full test matrix passed: all five general-server shards,
all five serialized-server shards, and all three general-workspace
shards.
- Latest-head CI also passed all three browser shards and their
aggregate gate, typecheck and release registry, build, runner
verification, canary dry run, policy, Docker context integrity, and
security gates. Greptile: 5/5, with the review thread resolved.
- Browser: created a task in a disposable instance. A real Git timeout
scheduled recovery, the next run completed through the run-scoped API
without manual Retry, and Done survived reload. The deterministic
process worker checked preserved source edits and excluded private
files; no model calls were made.
- `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec
playwright test --config
tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6
minutes). The persistent case made exactly three failed attempts, never
started the worker, showed the cause-specific notice, stayed stopped for
another scheduler tick, and retained Blocked after reload.
- Extra red-green coverage: 50 recovery tests passed after fixing an
exhausted-bootstrap classification that incorrectly implied unknown
provider actions. Missing or uncertain evidence still retains the safety
hold.
- Verified the documented Git executable override during repository
seeding.

## Risks

- A confirmed repository scan failure now fails closed instead of
falling back to directory sync. This prevents unfiltered copying but
makes previously hidden errors visible.
- Temporary host problems can create up to two additional setup
attempts, 30 seconds apart. Permanent scan errors do not auto-retry.
Generic recovery cannot reset this budget.
- The durable retry path still enforces ownership, pause, and work
eligibility. Integration tests cover restart, duplicate promotion,
pause, exhaustion, and non-retryable categories.
- No schema migration, new runtime setting, new retry budget, production
deployment, or historical task replay.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, repository
tools, shell execution, and browser testing. The exact deployment model
ID and context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 12:37:21 -05:00
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00
Nicky LeachandPaperclip 9ed55f6931 fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip starts local command-line sessions and stores provider
credentials through managed connections
> - One OpenAI Codex or xAI Grok subscription connection held a
credential lease for the full agent run
> - A second run then waited for the first run, and concurrent runs
could overwrite a newer credential
> - The write-back must compare fresh credentials while it holds the row
lock
> - This pull request removes the run lease, keeps the revocation guard,
and bounds Codex timestamps against the host clock
> - The benefit is safe concurrent use of one subscription connection
with newest-credential selection

## Linked Issues or Issue Description

**What happened?**

A managed OpenAI Codex or xAI Grok subscription connection held a
credential lease for the full agent run. A second run waited for the
first run to finish. The write-back gate also rejected any row change
before it compared credential freshness.

**Expected behavior**

Concurrent runs should start on one subscription connection. The server
should keep the newest valid credential and reject a credential write
after a person revokes the connection.

**Steps to reproduce**

1. Start two runs that use one OpenAI or xAI subscription connection.
2. Let both provider tools refresh the credential.
3. Finish the runs in either order.
4. Confirm that the newest valid credential remains in the connection.

**Paperclip version or commit**

`ec25bf1e4a81d1729a6d7276e6a486587cff4a0b`

**Deployment mode**

Local dev (`pnpm dev`), built from source.

**Agent adapter(s) involved**

Codex. The server path also covers xAI Grok subscription connections.

**Database mode**

Embedded PGlite for local development, and external Postgres for
deployments.

**Access context**

Both board and agent runs can use managed connections.

**Additional context**

The provider command-line tool refreshes credentials inside the sandbox.
The server copies the result back after the run. Two long runs can still
refresh one token hours apart, so the provider can reject the second
refresh. The server cannot observe that provider call.

## What Changed

- Remove the full-run credential lease for OpenAI Codex and xAI Grok
subscription connections.
- Lock and re-read the connection row before credential write-back.
- Accept only a strictly newer credential, while keeping the connection
revocation guard.
- Reject Codex freshness timestamps more than five minutes ahead of the
host clock.
- Add tests for both completion orders, xAI cleanup, revocation, the
Codex time bound, and agent hiring.
- Update the connection and run-log documentation.

## Verification

- `server/src/__tests__/ai-connections.test.ts` passes with 42 tests.
-
`packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts`
covers the five-minute boundary and the one-millisecond overflow.
- `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts`
passes.
- `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI
and Anthropic.
- The project type check reports no new error in changed files.
- GitHub Actions must pass on this pull request.

## Risks

The write-back now permits concurrent runs, so the provider may reject a
later refresh when both long runs use one token. The server keeps the
revocation guard and rejects future-dated Codex timestamps. No schema
change occurs.

## Model Used

OpenAI Codex, GPT-5. Context window and exact deployment build are not
exposed in this run. The model used tool calls, code inspection, and
test verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 01:10:54 -07:00
Nicky LeachandPaperclip c276d3fdc3 feat(observability): report terminal run failures to Sentry (#13446)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records the state of each agent run.
> - Terminal run failures need clear error tracking for operators.
> - The server did not report terminal `failed` or `timed_out`
transitions to Sentry.
> - This pull request reports each genuine terminal failure transition
with safe diagnostic data.
> - The benefit is faster diagnosis without changing run control flow or
exposing credentials.

## Linked Issues or Issue Description

**What happened?**

The server wrote terminal run failures but did not report them to
Sentry. Operators could not see these failures in error tracking.

**Expected behavior**

The server should report each genuine transition to `failed` or
`timed_out` to Sentry.

**Steps to reproduce**

1. Run an agent task that reaches a terminal failure state.
2. Inspect the Sentry events for the server.
3. Observe that the terminal run failure has no matching Sentry event.

**Paperclip version or commit**

The change targets the current `master` branch.

No public GitHub issue or pull request covers this change.

## What Changed

- Add `captureRunFailure()` as a fail-open Sentry entry point.
- Add `reportRunFailure()` to filter status, resolve the adapter, redact
text, and report the failure.
- Call `reportRunFailure()` beside each of the eight terminal status
writers.
- Report six diagnostic values: the instance host, task identifier, run
identifier, error message, error code, and agent adapter.
- Group events by error code and agent adapter while keeping the
redacted message in the event.
- Report only genuine transitions and avoid duplicate finalization
events.
- Keep Sentry failures outside run control flow.

## Verification

- `pnpm vitest run
server/src/services/__tests__/run-failure-report.test.ts
server/src/__tests__/run-failure-sentry.test.ts
server/src/__tests__/native-session-resumption.test.ts`
- `pnpm vitest run
server/src/services/execution-control-reconciliation.test.ts`
- `pnpm --filter @paperclipai/server typecheck`
- The full continuous-integration suite must run on this pull request.

## Risks

- The report path can add diagnostic events when Sentry is configured.
- The report path returns without action when Sentry is not configured.
- Redaction runs before length limits and before the event leaves the
process.
- The change has no migration and no schema change.

## Model Used

OpenAI Codex, GPT-5, tool use and code review support. The exact context
window and reasoning configuration are not exposed by the runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 23:32:15 -07:00
Devin FoleyandPaperclip 667c79ded2 fix: prevent retry-exhaustion events from exhausting attention-feed memory (#13451)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Its attention feed shows failed runs whose retry budget is
exhausted.
> - Startup retention builds this feed before startup completes.
> - The feed joins every exhaustion event to the full run context, then
removes duplicate runs in JavaScript.
> - Recovery can revisit an exhausted run and append the same event
again. This multiplies the data loaded into memory.
> - This pull request selects one small row per run in PostgreSQL and
makes exhaustion writes idempotent.
> - Existing duplicate events can stay in the database without
multiplying run contexts in server memory.

## Linked Issues or Issue Description

Refs #13367. This fixes the repeated-event allocation path in the
retention feed. Other full-feed sources and sweep cadence remain
separate concerns.

**What happened?**

The attention query loaded one full run context for every matching
exhaustion event. Deduplication ran only after the driver had loaded
those rows. A run with thousands of exhaustion events therefore produced
thousands of context copies. The startup retention sweep can exhaust the
server heap while reading this result.

**Expected behavior**

The query should return one row per exhausted run and only the context
fields that the feed needs. Repeated checks of the same exhausted retry
budget should reuse the original event.

**Steps to reproduce**

1. Create a failed run with a 32 KB context and 2,500 matching
exhaustion events.
2. Build the attention feed, including dismissed items, as startup
retention does.
3. Inspect the database result before JavaScript feed processing. The
old join returns 2,500 copies of the run context.
4. Call bounded retry scheduling repeatedly for a run at its retry
limit. The old writer appends another exhaustion event on every call.

**Paperclip version or commit**

The attention query was introduced by #9380 and is present in stable
`v2026.831.1`. The retention caller is also present in that stable
release. The later recovery callback added by #13075 provides a repeated
path into the exhausted-budget writer. This change is based on
`c0fda8fac` after rebasing onto current master.

Related PR:
[#12162](https://github.com/paperclipai/paperclip/pull/12162) changes
retention cadence and newer-run suppression queries. This change
addresses the exhaustion-event join and duplicate event writes.

## What Changed

- Select the newest company-scoped exhaustion event per run with a
PostgreSQL `DISTINCT ON` subquery before joining run data.
- Project only `issueId` and `taskId` from run context. Preserve JSON
types, fallback behavior, run ordering, company filters, and run/agent
status filters.
- Reuse an exhaustion event for the same run, reason, attempt, and retry
limit under the existing run-row lock. Recognize historical events
without a migration.
- Skip sequence allocation and live publication when a receipt already
exists.
- Document the run-log behavior and add PostgreSQL regression tests.

## Verification

- Focused attention, retry-scheduling, and event-sequencing suites: 72
tests pass on the rebased head `39ce77158`.
- The large-history fixture verifies two database result rows under 4
KB, newest-message selection, task-ID fallback, and company/status
filtering. This assertion runs before feed deduplication.
- Concurrent event tests verify one receipt across two database clients,
reuse of historical receipts, distinct reason/attempt/budget keys, and
stable event sequences.
- Repeated scheduling through a new service instance produces no extra
run-log or live events.
- `pnpm --filter @paperclipai/server exec tsc --noEmit`: passes.
- `pnpm -r typecheck` and `pnpm build`: pass on the rebased head,
including native runner checks.
- Greptile: 5/5 on `39ce77158`, with no review threads.
- [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/34929002304):
all 25 workflow jobs pass, including all server/workspace test shards,
serialized server suites, browser tests, typecheck, build, native-runner
verification, and the release dry run. Security checks also pass. The
branch has no merge conflicts.
- `pnpm test:run`: stopped with unrelated failures. Two chat integration
cases passed when rerun separately (2 passed, 993 unselected). Three
skill-cache cases failed on both this branch and the unmodified parent
commit, with `EACCES` during directory rename. The full suite is not
reported as green.

## Risks

- No schema migration or data cleanup is required. The query still scans
matching event history in PostgreSQL; its result size now scales with
exhausted runs.
- This does not bound every source in the attention feed or change
retention scheduling.
- An exhaustion receipt is emitted once per retry decision. Consumers
that observed repeated copies will now receive one event.
- Native source-event replay handling stays on its existing path.
- A deployed application boot has not been verified.

## Model Used

OpenAI GPT-6, used through Codex with reasoning, repository inspection,
code editing, and test execution. The runtime does not expose a more
specific model snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the focused tests locally and they pass (full-suite
limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 21:58:34 -07:00
Devin FoleyandPaperclip c0fda8fac5 fix(apps): restore action test picker scrolling and agent eligibility (#13414)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps action tests let operators use an agent's permissions.
> - The agent picker must scroll inside the test dialog.
> - Its body portal sits outside the dialog's scroll boundary and blocks
wheel input.
> - Admin permission bypasses also skip agent lifecycle checks.
> - This PR fixes scrolling and rejects agents that cannot receive
assignments.

## Linked Issues or Issue Description

**What happened?**

The Act as picker does not scroll with the mouse wheel inside an action
test dialog. Admins can also see terminated agents.

**Expected behavior**

The list scrolls normally. Terminated and pending-approval agents are
absent. Direct requests cannot test an action as one of those agents.

**Steps to reproduce**

1. Create enough agents to overflow the list. Terminate one agent.
2. Open a connected app's Permissions tab. Click Test on an action.
3. Open Act as and use the mouse wheel over the list.
4. Check whether the terminated agent appears as an admin.

**Paperclip version or commit**

Reproduced on master at f2c5e54dc. Rebased onto d351e08de.

**Deployment mode**

Local dev, built from source. This is a shared Apps bug and does not
require Railway credentials.

Related search result: #9918 added search to a separate secrets picker.
It does not cover this action test dialog. No duplicate action-test
picker PR was found.

## What Changed

- Keep the action tester's agent popover inside its dialog's scroll
boundary.
- Check company membership and the shared agent lifecycle policy before
assignment permission bypasses.
- Reject terminated and pending-approval agents in lists, previews, and
test calls.
- Reuse company-scoped rows during listing to avoid extra per-agent
queries.
- Add route tests for both admin modes and a browser wheel-scroll
regression.

## Verification

- 335 focused tool-access and TestPanel tests passed after rebase. After
the review cleanup, both admin regressions and writable-agent selection
passed again (3 tests).
- The browser regression failed before the fix because wheel input left
scrollTop at zero. It passed after the fix, including search and
selection. It executes no provider tools.
- Full typecheck, build, and token gates passed during implementation.
Token gates passed again after rebase.
- The full test run reported a failure in the GitHub installation
recovery chat test. That test passed in isolation. The full run was
stopped after the failure, so later groups were not completed.

Manual check: open an action's Test dialog, open Act as, scroll, search,
and select an agent. Terminated agents must be absent.

## Risks

The portal change affects only the picker inside the action test dialog.
The browser test covers scrolling and selection. Paused agents remain
eligible under existing assignment rules. There is no migration or
provider policy change.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
The exact serving model ID and context window were not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 21:18:12 -07:00
Nicky LeachandClaude Opus 5 b64469e403 feat(workspaces): add an operator default for isolated execution workspaces (#13444)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The execution workspace subsystem decides if a task run uses the
shared project checkout or an isolated per-task git worktree
> - The mode comes from the project policy, then the task settings. A
project that stores no policy always falls back to the shared checkout
> - An operator who wants every project to use isolated workspaces must
therefore edit each project one at a time, and must repeat this for each
new project
> - There is no instance-level control, so a fleet operator cannot set
this default at all
> - This pull request adds a managed experimental flag that moves the
default for projects that store no policy of their own
> - The benefit is that an operator sets the workspace default one time,
and every current and future project follows it

## Linked Issues or Issue Description

No public issue exists. The description follows the feature request
template.

**Subsystem affected**

Execution workspaces. The files are
`server/src/services/execution-workspace-policy.ts` and the run dispatch
path in `server/src/services/heartbeat.ts`.

**Problem or motivation**

`resolveExecutionWorkspaceMode` reads only the project policy, the task
settings, and a legacy field. Its last statement returns
`shared_workspace`. A project that stores no policy always gets the
shared project checkout.

An operator has no way to change this default for many projects at the
same time. The operator must edit each project, and must edit each new
project again later. Tasks in one project therefore share one checkout,
and they run one at a time when the environment driver makes the
scheduler serialize them.

**Proposed solution**

Add the managed experimental flag `enableIsolatedWorkspacesByDefault`.
When the flag is on, a project that stores no policy of its own resolves
as if it selected isolated workspaces. A project that stores a policy
keeps that policy.

The new helper substitutes a project policy. It does not move the last
statement of `resolveExecutionWorkspaceMode`. Two behaviors make this
necessary:

- A task that has no project must keep its current behavior. An isolated
workspace needs a repository to cut a worktree from.
`isUnrunnableWorktreeCombo` blocks an isolated task that has no
`projectId` and no `projectWorkspaceId`. A moved fallback would resolve
isolated for project-less tasks, such as agent chat, and stop them
before dispatch.
- The mode and the strategy must agree.
`buildExecutionWorkspaceAdapterConfig` supplies the default
`git_worktree` strategy only when one layer asserts workspace control. A
moved fallback would leave isolated mode with a `project_primary`
strategy.

**Alternatives considered**

- Change the last statement of `resolveExecutionWorkspaceMode` to
`isolated_workspace`. This is one line, but it changes the default for
every deployment. It is also not gated, so it would apply where isolated
workspaces are off.
- Write the policy to each project row with a script. This does not
cover new projects, and it does not cover new instances.
- Add an instance defaults section to the managed-config document. This
needs a new document key, new validation, and new delivery code. A
boolean flag reuses the delivery machinery that exists today.

**Roadmap alignment**

`ROADMAP.md` does not list execution workspace defaults. This change
adds an operator control to an existing capability. It does not add a
new capability.

**Additional context**

The flag is `tier: "managed"`. A cloud operator can therefore deliver it
with the managed-config machinery that exists today. No new delivery
code is needed.

## What Changed

- Add `enableIsolatedWorkspacesByDefault` to the feature catalog with
`tier: "managed"`. Both defaults are off.
- Add the flag to the experimental settings schema, the type, and both
branches of `normalizeExperimentalSettings`.
- Add `applyDefaultIsolatedExecutionWorkspacePolicy` to
`execution-workspace-policy.ts`. It substitutes `{ enabled: true,
defaultMode: "isolated_workspace" }` only when the flag is on, the task
has a project, and the project stores no policy.
- Apply the helper in the run dispatch path in `heartbeat.ts`, after the
existing `gateProjectExecutionWorkspacePolicy` call. The `hasProject`
argument reads the resolved project row, not the raw `projectId` of the
task.
- Gate the new flag behind `enableIsolatedWorkspaces` at the call site.
The new flag does nothing on its own.
- Add a toggle card to the instance experimental settings page. The card
shows only when isolated workspaces are on.
- Add eight tests for the new helper.

## Verification

Commands:

```
pnpm --filter @paperclipai/shared typecheck
pnpm --filter @paperclipai/ui typecheck
cd server && ../node_modules/.bin/tsc --noEmit
```

The server typecheck script also builds a Rust binary. I ran `tsc`
directly because this machine has no `cargo`. The server package reports
no type errors.

Tests:

```
./node_modules/.bin/vitest run \
  server/src/__tests__/execution-workspace-policy.test.ts \
  server/src/__tests__/instance-settings-service.test.ts \
  server/src/__tests__/instance-settings-cloud-defaults.test.ts \
  server/src/__tests__/instance-settings-managed-overlay.test.ts \
  server/src/__tests__/instance-settings-routes.test.ts \
  server/src/__tests__/managed-config.test.ts \
  server/src/__tests__/heartbeat-workspace-busy.test.ts \
  server/src/__tests__/heartbeat-workspace-session.test.ts \
  server/src/__tests__/heartbeat-workspace-ready-comment.test.ts \
  server/src/__tests__/execution-workspaces-service.test.ts \
  server/src/__tests__/issue-runtime-workspace-binding.test.ts \
  server/src/__tests__/run-trust-preset.test.ts \
  packages/shared/src/feature-catalog.test.ts \
  packages/shared/src/settings-visibility.test.ts \
  packages/shared/src/validators/instance.test.ts \
  ui/src/pages/InstanceExperimentalSettings.test.tsx \
  ui/src/components/Sidebar.test.tsx
```

All of these files pass. The new tests cover each of these cases:

- The helper substitutes an isolated policy for a project that stores
none.
- The helper changes nothing while the flag is off.
- The helper changes nothing for a task that has no project.
- The helper keeps a stored policy, including a policy with `enabled:
false`.
- The resolver returns `isolated_workspace` for an unpolicied project.
- An explicit task setting still wins over the operator default.
- The substituted policy produces the `git_worktree` strategy.
- A project-less task does not become an unrunnable worktree.

To confirm the behavior by hand:

1. Turn on Isolated Workspaces, then turn on Use Isolated Workspaces By
Default.
2. Open a project that has no execution workspace policy.
3. Start a task in that project.
4. The run gets its own worktree. Tasks in that project no longer wait
for each other.

## Risks

Low to medium. The details:

- The flag defaults to off, and it is inert unless
`enableIsolatedWorkspaces` is also on. An instance that does not turn on
both flags sees no change.
- A project that stores a policy keeps it. This includes a policy with
`enabled: false`, which the helper reads as a decision to stay on the
shared checkout.
- When an operator turns the flag on, the workspace configuration
fingerprint changes for projects that store no policy. Their next run
creates a new workspace. This is correct, because the mode did change,
but the first run after the change does more setup work.
- A task that is in flight when the flag changes resumes with a
different workspace path than the path its session remembers. An
operator should let current runs finish before turning the flag on.
- Isolated workspaces use more disk, because each task gets its own
worktree.

## Model Used

Claude Opus 5 (`claude-opus-5`) in Claude Code, with extended thinking
and tool use. The model read the repository, made the change, and ran
the typechecks and tests above.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 20:59:56 -07:00
Nicky LeachandPaperclip ee81cee76d fix: allow concurrent agent runs on one Anthropic subscription connection (#13445)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Agent runs use provider connections and subscription credentials.
> - File-backed subscriptions need a lease while a run writes new
credentials.
> - Anthropic subscriptions pass a token and do not write a credential
file.
> - The current lease blocks concurrent Anthropic runs without
protecting shared state.
> - This pull request limits the lease to file-backed subscriptions and
tests both paths.
> - The result lets Anthropic runs share one connection safely while
file-backed credentials remain serialized.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This change improves concurrent use of one subscription connection.
Anthropic runs no longer wait for a lease that protects no shared file.

**Subsystem affected**

server/ — REST API and orchestration services, with related server tests
and connection documentation.

**Current behavior**

The server takes a connection lease for every subscription invocation.
The lease blocks a second Anthropic run while the first runtime remains
open.

**Proposed behavior**

The server takes the lease only when the subscription writes a provider
authentication file back to the grant. Anthropic runs can proceed at the
same time.

**Reason and benefit**

Anthropic passes its token through an environment variable and writes no
file. Removing this unnecessary wait improves concurrency and preserves
credential safety for OpenAI and xAI.

**Breaking changes**

None. OpenAI and xAI keep the existing lease and retry behavior.

## What Changed

- Limit the AI credential lease to subscriptions that write a provider
authentication file.
- Add coverage for concurrent Anthropic runs and hired-agent runs.
- Keep the existing OpenAI contention assertions as a sensitivity
control.
- Update the AI connection documentation to describe the lease boundary.

## Verification

- Run `server/src/__tests__/ai-connections.test.ts`.
- Run `server/src/__tests__/agent-hire-ai-connections.test.ts`.
- Run
`server/src/__tests__/heartbeat-ai-subscription-contention.test.ts`.
- Run the server type check.
- Review the pull request checks after GitHub starts continuous
integration.

## Risks

The main risk is an incorrect provider classification. The write-back
condition remains unchanged, and OpenAI contention tests keep the
file-backed lease behavior under test. No schema or API contract changes
occur.

## Model Used

OpenAI GPT-5. Exact runtime model ID and context window are not exposed
in this execution. The model used tool calls, shell commands, and GitHub
operations.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 20:53:56 -07:00
Devin FoleyandPaperclip 52d120f68d fix(release): wait 30 minutes for npm to expose a published version (#13436)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every release publishes a batch of npm packages and then waits for
each one to become visible before continuing
> - npm accepts a publish immediately but exposes it later, so that wait
exists to keep a release from continuing past a package nobody can
install yet
> - Today that wait was too short twice in a row, and each timeout
aborted the release with the batch half-published
> - Version numbers are derived from what is already on npm, so every
retry moves to a new number and meets the same lag
> - Two canary attempts burned two versions this way and shipped nothing
> - This pull request raises the per-package budget from ten minutes to
thirty
> - The benefit is that ordinary registry lag costs waiting instead of a
failed, half-published release

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
`publish_canary` failed twice in a row with the batch half-published:

```
Warning: npm accepted @paperclipai/server@2026.914.0-canary.2, but the version did not become registry-visible.
Error: stopping release: npm did not publish and expose @paperclipai/server@2026.914.0-canary.2
```

The version was accepted at 17:49:09 and became visible at 18:04:29 —
about five minutes after the poll gave up. `shared`, `db` and
`adapter-utils` published at that version; `server`, `paperclip-runner`
and the root package did not.

**Expected behavior**
Ordinary registry propagation delay costs the release some waiting, not
a failure. A package that becomes visible after 15 minutes must not
abort the batch because of the old 10-minute wait. Longer registry
outages can still leave a partial batch.

**Steps to reproduce**
1. Publish any channel while npm is propagating slowly.
2. A package takes longer than `NPM_PUBLISH_VERIFY_ATTEMPTS *
NPM_PUBLISH_VERIFY_DELAY_SECONDS` to become visible.
3. The release aborts, that version is half-published, and the retry
picks a new version number and meets the same lag.

**Paperclip version or commit**
Present on master. Observed on 2026-09-14 across canary runs in workflow
run 34869494325.

## What Changed

- Increase npm visibility checks from 60 to 180, retaining the 10-second
delay: about 30 minutes per package.
- Increase canary, nightly, beta, and stable publish job timeouts from
90 to 150 minutes.
- Add offline regression tests using the workflow's actual settings.
They cover the observed 15-minute 20-second delay, immediate visibility,
exhausted retries, and job timeout sizing.
- Load the access router in test setup so its cold transform does not
consume the first permission test's 10-second timeout. The permission
assertions are unchanged.

## Verification

- `node --test scripts/release-lib.test.mjs`: 14 passed.
- `pnpm run test:release-registry`: 129 passed.
- `pnpm exec vitest run
server/src/__tests__/access-routes-permissions-upgrade.test.ts`: 3
passed.
- Regression proof in temporary fixtures: restoring 60 attempts fails
the observed-delay test; restoring 90-minute jobs fails the
timeout-budget test.
- [CI run
34916632804](https://github.com/paperclipai/paperclip/actions/runs/34916632804):
all jobs passed, including typecheck/release registry, build, all
general and serialized server shards, browser tests, runner
verification, and canary dry run. The PR has 31 successful checks and
two expected Storybook skips at `a27f5e896`.
- [Previously failing serialized
shard](https://github.com/paperclipai/paperclip/actions/runs/34916632804/job/104215809616):
all three access-route permission tests passed in CI after preloading
the router.
- Greptile's final review is 5/5 with no outstanding findings. Both
review threads are resolved.
- Local limits: `pnpm -r typecheck` and `pnpm build` stop at the runner
package because this machine has no Rust `cargo` executable. The
duplicate full local `pnpm test:run` was interrupted while the complete
CI matrix ran. Targeted local results are listed above; full validation
is from CI.

The previous CI failures were unrelated to npm propagation:

- [PR
review](https://github.com/paperclipai/paperclip/actions/runs/34891770396)
required a test file for this fix.
- [Serialized server shard
3](https://github.com/paperclipai/paperclip/actions/runs/34891773736/job/104136392196)
timed out in the first access-route permission test at 10 seconds. The
other two tests in that file passed.
- The canary dry run passed in that same CI run.

## Risks

Low risk. Production behavior changes only in release waiting budgets.

- Polling exits as soon as npm exposes the version, so healthy publishes
do not wait longer.
- An unavailable version now takes about 30 minutes to report. Polls
remain bounded and still fail the release on exhaustion.
- The 150-minute jobs leave roughly 30 minutes for setup/build plus four
full polling windows. npm command runtime and later release steps also
consume that budget; a broader outage can still interrupt a batch.
- The permission-test change moves module loading into a bounded setup
hook; it does not relax authorization assertions.
- No schema changes or operational migrations.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

- OpenAI GPT-6 (Codex), with reasoning, tool use, and code execution,
for the CI follow-up. Context-window size is not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted checks listed
above; full validation passed in CI
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes —
workflow comments and verification details
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 20:16:29 -07:00
DottaandPaperclip a2e7ffdc34 fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents can run in remote sandboxes.
> - Connection checks must use the selected execution target.
> - Recovery must respect the controller that owns an active run.
> - A host path or PID does not describe a remote sandbox.
> - This pull request checks sandbox paths on the target and preserves
live controller leases.

## Linked Issues or Issue Description

**What happened?**

Selecting an AI account for a sandbox agent could fail because the
Claude ACP environment check tried to create the sandbox directory on
the Paperclip host. The recovery sweep could also interrupt a sandbox
run while its controller lease was still valid. It treated a PID absent
from the local host as proof that the run had stopped.

**Expected behavior**

ACP checks directories on the selected execution target. Recovery leaves
a run with a live controller lease alone. Its final database write
rejects a stale snapshot after renewal, a claim, a controller change, or
a runtime change.

**Steps to reproduce**

1. Test a Claude ACP sandbox agent with a directory that cannot be
created on the host. The check fails before this fix.
2. Give a running sandbox task a valid controller lease and a PID absent
from the host. Run the stale-lock sweep without an in-memory handle. The
sweep interrupts the run before this fix.
3. Renew or replace the controller between the sweep's read and write.
The old snapshot must not end that controller's run.

**Paperclip version or commit**

Rebased onto `origin/master` at
`0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases
still fail against this base.

Refs #13438, which supplies the managed hiring and task-connection
behavior, and #13433, which preserves non-assignee subscription comment
wakes. This PR preserves both upstream changes and addresses the two
remaining sandbox failures.

## What Changed

- Resolve and create Claude ACP test directories through the
execution-target helpers.
- Preserve active legacy controller leases during stale-lock recovery,
including finalization after a task becomes terminal.
- Recheck the controller, lease, runtime mode, and native ownership in
the terminal database write.
- Add two sandbox-directory cases and five database-backed
controller-lease cases.
- Document the target used for ACP directory checks.

## Verification

- Red: all seven new cases fail against `0e9b24c82` without these two
implementation changes.
- Before the final upstream sync, 192 focused tests passed. Full `pnpm
test:run` coverage completed using the repository's group/shard runner:
all general server and workspace groups passed, and all 147 serialized
server suites passed across the initial run and isolated continuations.
Five cold-import timeout suites passed with
`--experimental.fsModuleCache`; their assertions and deadlines were
unchanged.
- The hiring routes, default-selection service, and upstream hiring
tests match `origin/master` exactly. The two remaining fixes are
unchanged by the final rebase.
- On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372
focused tests pass across 16 suites covering both upstream changes and
these fixes. Two timeouts in the combined run (database setup and an
existing ACP case) pass in isolated reruns with fresh test homes and
temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on
this head.
- All 32 active checks pass on final head `bd28d5cef`, including the
full test matrix and browser shards, in [CI run
34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849).
Two Storybook checks are skipped by path filters.
- Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or
unresolved review threads.

## Risks

- A failed remote directory check still blocks connection adoption.
- A live controller retains finalization authority after its task
becomes terminal. Cleanup waits for ownership to expire and must pass
the final ownership check.
- No schema or credential-storage changes.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, repository inspection, code
execution, and API tools. The exact serving model ID and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 22:13:57 -05:00
DottaandPaperclip 8f1905d34d fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Projects can now attach several source repositories.
> - Task preparation still treated these sources as alternative
workspaces.
> - Sandbox sync preserved Git history only for the selected repository.
> - A task needs every attached repository to complete work across the
project.
> - This pull request prepares all distinct project repositories and
preserves their separate Git histories through sandbox restore.

## Linked Issues or Issue Description

**What happened?**

A user reported that a project with two repositories received only the
first repository in Daytona. Repository-only project rows also reached
the agent with null local paths. Managed checkouts with matching
repository names could resolve to the same directory.

**Expected behavior**

Local and sandbox tasks receive every distinct repository attached to
their project. Repository-only sources work without preconfigured local
folders. Each repository keeps its own Git history and working files.

**Steps to reproduce**

1. Create a project with two repository sources and no local folder
paths.
2. Assign a task to the project and run it in Daytona.
3. Inspect the task workspace and the repository paths exposed to the
agent.
4. Observe that the original implementation supplies only the selected
checkout.

Related change: #13010 added multiple repository selection. The open
repository-catalog proposals #11234 and #11228 cover a different data
model. This fix uses the existing project workspaces.

## What Changed

- Materialize each additional distinct repository as an editable
checkout inside the task root. Seed configured local sources with their
current working files and retain task edits across runs.
- Pass materialized repository paths to local agents and native sandbox
task prompts. Apply existing run-scoped Git credentials to each remote
clone.
- Preserve each repository's Git history, dirty files, and restore
baseline during sandbox staging and durable recovery. Apply each
repository's ignore rules and the operator's workspace exclusions.
- Keep same-name managed repositories in separate directories. Report
additional clone failures before the task starts.
- Add task-level, checkout, sandbox round-trip, environment-hint, and
recovery-descriptor regression coverage. Document checkout and restore
behavior.

## Verification

- Red: the original implementation fails the sandbox test because the
second repository has no Git directory. It also fails the same-name
checkout test and both real-database task tests because repository hints
have no local path.
- Green: focused tests pass for one and two repository-only sources,
local source edits, clone failures, per-repository credentials, separate
Git histories, ignored files, and recovery from remote or durable seed
state.
- Live Daytona smoke passed with two disposable repositories through the
production provider sync functions. Both repositories arrived with Git
history. Commits from both restored locally. Ignored files stayed
excluded. The disposable sandbox was deleted.
- Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm
build`.
- Final focused coverage: 254 assertions across the six changed test
areas passed across the serial run and an isolated rerun of the existing
process-kill timing test. The live Daytona smoke also passed.
- The local `pnpm test:run` overlapped source edits and retained stale
transformed code. Its first phase reported 12,240 passed assertions,
nine failed assertions, three hook failures, and one worker error; later
phases did not run locally. This run is not claimed as green. Fresh
focused tests verify the changes, and every general/workspace and
serialized-server CI shard passes on the final commit.
- Final CI is green on `93ab76763`: all test shards, all three browser
shards, typecheck, build, runner verification, canary dry run, and
security checks. The initial unrelated chat-delivery browser timing
failure passed in the final CI run. Optional Storybook visual checks
were skipped.
- Greptile is 5/5 on the final commit with no unresolved review threads.
Its checkout-race finding was reproduced with a failing test, fixed, and
rechecked.

## Risks

- Additional repositories need disk space and clone time. Access failure
for an attached repository stops preparation.
- Additional checkouts live under `.paperclip-repositories/` and keep
independent histories. Changes stay in those task copies; they do not
overwrite configured source folders.
- Detached or reconfigured repository copies are retained under
`.paperclip-runtime/detached-repositories/`. Sandbox recovery retains
per-repository merge baselines.
- No database migration, UI contract change, or new credential
delegation is required. Referenced projects retain their separate
read-only behavior.

## Model Used

OpenAI Codex, based on GPT-6, with repository inspection, code
execution, and tool use. The runtime does not expose a more specific
model deployment ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 18:37:51 -05:00
DottaandPaperclip 0e9b24c821 fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Concurrent task runs can share a managed AI subscription with one
credential lease.
> - #13438 added durable retries for busy subscriptions and task-lock
checks.
> - A comment wake can run for an agent who is not the task assignee.
Such a run never owns the task lock, so the new checks suppress its
retry.
> - This pull request preserves those comment wakes while keeping the
lock checks for assignee runs.
> - It also shows the subscription wait in task status and preserves the
failure count through the database projection.

## Linked Issues or Issue Description

Refs #13438.

**What happened?**

When a subscription is busy, a non-assignee comment wake is cancelled
without a successor. Repeated subscription waits also display an
inflated attempt count because the execution query omits the preserved
failure count.

**Expected behavior**

An eligible comment wake waits and resumes without claiming the
assignee's task lock. The task shows “Waiting for AI subscription”.
Waiting does not consume provider-failure retries. Assignee retries
still stop when the task lock is cleared or transferred.

**Steps to reproduce**

1. Configure an agent with a managed subscription and concurrent runs.
2. Hold its credential lease in one run.
3. Mention the agent on a task assigned to another actor.
4. Release the lease and inspect whether the comment wake has a
scheduled retry.

The test uses a real embedded Postgres database and a real credential
lease. Provider execution uses a fixture. The reassignment test applies
a database mutation after the real checkout. No live provider account is
needed.

**Paperclip version or commit**

Based on `f912ecaac` from #13438. Before the reconciliation fix, the
rebased regression at `2063cfd13` failed because the comment wake had no
scheduled retry. The original configuration failure was reproduced on
`5282cabde` before #13438 merged.

**Deployment mode**

Built from source with embedded Postgres for local verification.

## What Changed

- Record non-assignee comment-wake authority at admission while holding
the task and run locks. A later reassignment cannot grant this
exception. Preserve it through repeated subscription waits.
- Keep master’s lock checks for assignee retries, including the check
inside the scheduling transaction.
- Show the subscription wait and preserve its failure count in the
execution projection query.
- Limit pre-provider wait receipts to fresh executions. A persisted
native execution input must retain its recovery path.
- Add lease, retry, projection, cancellation, reassignment, pause,
revocation, service-recreation, and ownership regression coverage.
- Use master’s 60–120 second retry interval and document the resulting
behavior.

## Verification

- Red: after rebase, the comment-wake test failed with no scheduled
retry; the other 50 contention and retry-scheduling tests passed.
- Red/green: reassigning the task immediately after real checkout
reproduced an unwanted successor for both assignment and comment wakes.
Both tests pass after recording authority at admission.
- Green: all 285 targeted tests pass across 12 suites, including
#13438’s four cancellation-race phases, its hiring/connection suite,
run-dispatch integration tests, and both reassignment regressions.
- `pnpm -r typecheck` passes after rebase.
- `pnpm build` passes on the final revision.
- The full general and serialized test suites pass in [CI run
34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681)
on `99f1e4c1d`. Full-suite verification ran in CI; local verification
used the 285 targeted tests.
- All 32 active PR checks pass, including build, typecheck, native
runner verification, canary dry run, and all browser shards. Two
Storybook checks are skipped by path filters.
- Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All
review threads are resolved.
- `git diff --check` and `node scripts/check-module-boundaries.mjs`
pass.

## Risks

- The non-assignee exception requires authority recorded under admission
locks or its server-created subscription retry. Tests cover
caller-supplied flags, assignment and comment reassignment races,
repeated comment waits, and assignee lock protections.
- A wait uses master’s 60–120 second interval. A long-held lease can
produce multiple wait records.
- Existing native sessions can have provider effects. They must not
receive a fresh-execution receipt that permits replay.
- No schema migration or credential permission change is required.

## Model Used

OpenAI Codex, model `gpt-6-astra`, with repository search, code editing,
shell execution, and test tools. The session does not expose its
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 16:58:37 -05:00
DottaandPaperclip f912ecaacf fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents hire other agents and assign tasks to them.
> - Managed AI connections must follow those hires across legacy and
native runners.
> - Missing accounts should pause task execution and let the user
connect from the task.
> - Subscription contention must wait without asking for new
credentials.
> - This pull request fixes these paths and the native tool and Daytona
staging failures found during live tests.
> - The result is a working hire, subtask, and connection setup flow on
local and remote runners.

## Linked Issues or Issue Description

**What happened?**

A managed Claude or Codex agent could hire a teammate without a usable
AI binding. Cross-provider hiring could fail before the user had a
chance to connect the new provider. First-time task setup did not show
the existing AI credential form inline. A busy subscription could
request a new connection. Native API replies could stop the parent after
a hire had already committed. Fresh Daytona sandboxes could fail to
extract read-only skill directories created on macOS.

**Expected behavior**

Compatible hires inherit the managed connection choice. A hire for
another provider uses the responsible user's default. If that account is
missing, the hire succeeds and the task asks for a connection.
Completing setup in the task resumes work automatically. Explicit child
auth settings and existing unmanaged login paths keep precedence.
Shared-account access checks remain in force.

**Steps to reproduce**

1. Connect a Claude or Codex parent with a managed AI account.
2. Ask it to hire one agent of each provider and create a self-assigned
subtask.
3. Assign work to both hires without connecting the second provider
first.
4. Connect the missing provider from its task card.
5. Check that all tasks finish and same-provider work uses the original
account.
6. Repeat with native runners and fresh Daytona sandboxes. The opt-in
browser suite in `tests/hiring-ai-connections/README.md` performs these
steps.

**Paperclip version or commit**

The live failures were reproduced from `f2c5e54dc`. The branch is
rebased onto `5282cabde`.

**Deployment mode**

Isolated local development instance. Legacy CLI and native runners.
Local execution and ephemeral Daytona sandboxes.

Related work: Refs #13247 for managed AI connections. Refs #13268 for
legacy credential-reference inheritance, which this branch preserves.
Refs #13432 for a concurrent managed-inheritance fix. This PR also
covers cross-provider task setup, subscription waits, native API
replies, and Daytona extraction. It permits missing responsible-user
defaults at hire time; restricted shared selections still fail.

## What Changed

- Apply managed connection defaults to both agent creation routes.
Preserve explicit auth choices and legacy credential-reference
inheritance.
- Allow hires before their responsible user connects the provider. Keep
approval gates, company boundaries, and shared-account access checks.
- Reuse the production AI credential form inside the pending task card.
Resume the task after setup.
- Retry subscription lease contention without consuming the
provider-failure allowance or creating a connection request.
- Require task execution-lock ownership when scheduling, promoting, and
dispatching subscription retries. Recheck ownership under the issue row
lock.
- Rename the HTTP operation identity at the native tool boundary so it
cannot override the runner's operation identity.
- Delay directory permission restoration during Daytona extraction.
Preserve the final read-only modes.
- Add database-backed regressions, real browser acceptance tests, and
Storybook states. Document setup and run-log behavior.

## Verification

- Six real browser scenarios passed: both parent providers on legacy
local and legacy Daytona; native Codex locally; native Claude on
Daytona. Each scenario hires both providers, completes a self-subtask
and assigned work, and connects the missing provider inline with
automatic continuation.
- Successful runs verify the account, responsible user, runner mode, and
Daytona lease. All 18 test sandboxes were deleted.
- Live authentication used API keys. Subscription inheritance, lease
contention, and retry have integration coverage. Fresh subscription
OAuth sign-in was not automated.
- Red/green tests reproduced missing bindings, missing inline forms,
subscription contention, native API reply failure, and GNU tar
permission failure.
- Seven Storybook browser checks passed. They cover both providers,
method selection, narrow layout, completion, cancellation, and invalid
credentials.
- Full local suite coverage completed before rebase. Initial timing and
fixture startup failures passed unchanged on isolated reruns. The first
full command did not exit cleanly; the remaining workspace and
serialized groups were completed separately.
- After rebase, 107 hiring/auth/retry tests and 59 native API,
task-card, and Daytona tests passed. The full workspace typecheck,
production build, and token gates passed again. Storybook build passed
before rebase.

- Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and
four explicit cancellation-race cases passed. Eight cross-provider cases
cover stale auth keys on both creation routes and both runner types.
Server typecheck and build passed.

- Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks
passed. Two optional Storybook jobs were skipped. The full server,
workspace, browser, native runner, build, typecheck, and release checks
passed.
- Three unchanged tests initially failed on a busy port, a chat row-lock
race, and preview-server readiness. Each affected job passed after one
CI rerun. Isolated local checks also passed: 41 credential tests, the
chat-concurrency case, and 25 preview-runtime tests.
- Greptile reviewed the final commit at 5/5. Both review threads are
resolved. GitHub reports no merge conflicts.

## Risks

- A missing personal account now defers authentication to the first
task. Explicit incompatible bindings and restricted shared accounts
still fail at hire time.
- An inherited personal default uses the responsible user's existing
authorization to install access for the new agent. It never copies
credentials or another user's identity.
- Subscription contention retries after a delay and rechecks task
eligibility. It does not consume the provider-failure budget.
- Native hiring uses the existing managed API-tools opt-in. Remote
native runners require a matching Linux binary and provider pack, as
documented in the acceptance README.
- No schema changes. Live tests make paid provider calls and remain
opt-in.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository
inspection, code execution, browser automation, and API tools. The
context-window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 16:15:48 -05:00
02c7175e72 feat(agents): hired agents inherit provider credential references (#13268)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent adapters need provider credentials to start work
> - A hired agent can lack the credential reference that its hiring
agent already uses
> - The child agent then cannot authenticate, even when the company has
a valid credential
> - This pull request copies matching credential references from the
hiring agent to the hired agent
> - The benefit is that hired agents can start with the provider access
that the hiring agent already uses

## Linked Issues or Issue Description

This change relates to [PR
#9920](https://github.com/paperclipai/paperclip/pull/9920), which covers
credential inheritance for other agent creation paths. This pull request
covers hiring-specific inheritance and fixed Claude OAuth binding
checks.

### What existing behavior does this improve?

The agent hire route builds the child adapter configuration from the
hire request only.

### Current behavior

A hired agent does not receive matching provider credential references
from its hiring agent. The child agent cannot run when the request omits
the credential.

### Proposed behavior

The hire route inherits matching credential references from the hiring
agent. The request keeps priority. A Claude hire that supplies any
Claude credential inherits none.

### Reason and benefit

The child agent can use the provider access that the hiring agent
already uses. The change copies references only and never copies raw
token values.

### Breaking changes

None. The change affects only hires that need an inherited reference.

## What Changed

- Copy matching credential references from the hiring agent into the
hired agent adapter configuration.
- Preserve pinned versions, `required`, and `allowMissingOverride`
fields on each copied reference.
- Keep hire-request credentials ahead of inherited credentials.
- Reject inherited fixed Claude OAuth bindings unless the parent agent
passes company, adapter, and exact-binding checks inside the same
transaction.
- Add route and service tests for inheritance, precedence, and binding
validation.

## Verification

- `pnpm exec vitest run --project @paperclipai/server
src/__tests__/agent-hire-auth-inheritance-routes.test.ts
src/__tests__/agents-claude-oauth-binding.test.ts` — 60 passed.
- `pnpm exec vitest run --project @paperclipai/server
src/__tests__/agent-hire-idempotency-routes.test.ts
src/__tests__/agents-service-secret-bindings.test.ts
src/__tests__/secrets-service-user-secret-owner-scoped.test.ts` — 28
passed.
- `tsc --noEmit` in `server/` — the error count matches the merge base,
with no error in either changed source file.
- `git diff --check` — clean.

## Risks

Low risk. The route copies references, not raw tokens. The request keeps
precedence. Company, adapter, and exact-binding checks protect the
inherited Claude OAuth path.

## Model Used

Codex, OpenAI GPT-5, with code execution and review support. The
implementation commit predates this pull request handoff.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 09:26:27 -07:00
DottaandPaperclip 78ce96a48b fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner supplies each agent with instructions, assigned
skills, and tools.
> - Native Claude lost its skill snapshot before provider launch. It
also changed exact model IDs to aliases.
> - The read permission mode denied ordinary Paperclip reads because
this runner has no interactive approval handler.
> - This pull request restores the missing context and permits only
assigned tools that Paperclip defines as reads.
> - Other requests that need approval stop with a clear action for the
operator. They do not retry automatically.
> - Agents can complete read tasks. Users can see why a restricted task
stopped.

## Linked Issues or Issue Description

**What happened?**

Native Claude could not find assigned skills. The ACP layer could
replace an exact model ID with an alias and then fail identity
verification. The default `approve-reads` mode denied Paperclip read
tools. A denied write showed a generic transport failure.

The tool descriptions also called live operations “mock” operations. The
completion prompt said to call a completion tool once, although the
protocol can reject a claim and require a corrected call.

**Expected behavior**

Load assigned skills before launch. Keep the selected model ID. Allow
assigned Paperclip reads. Stop an operation that needs approval with
clear instructions when no approval handler exists.

**Steps to reproduce**

1. Configure a native Paperclip Runner agent with ACPX Claude and an
exact model ID.
2. Assign a skill and ask the agent to use it.
3. Select the read permission mode and ask the agent to read task
context and list documents.
4. Ask the agent to write a document. Check the task state and recovery
message.

**Paperclip version or commit**

Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`.

**Deployment mode**

Local native runner with real Claude, Rust runnerd, the Paperclip
server, and embedded PostgreSQL.

Related: #13196 fixed remote skill staging in the legacy `claude_local`
adapter. The native runner uses a separate path, which this pull request
fixes.

## What Changed

- Carry the runtime context through Rust and the ACPX sidecar. Load
assigned Claude skills after the provider lifetime lease is held.
Refresh the files on each open.
- Write the exact requested model ID into the isolated Claude settings.
- Grant exact MCP permissions for the intersection of assigned tools and
Paperclip's read catalog. Provider hints cannot grant access. Existing
task-control permissions stay in place.
- Stop requests that need an unavailable approval handler. Preserve the
typed error through the server. Show “Approval required” on the task and
require operator action without automatic retry.
- Label the setting “Allow Paperclip reads.” Remove “mock” from live
tool descriptions and regenerate the contracts.
- Change one completion-prompt sentence to require one accepted result.
Add regression tests and update the runner documentation.

## Verification

- Red/green regression tests cover read admission, unavailable approval
handling, server recovery, and the task error message.
- A real Claude read trial failed before the fix and completed after it:
[red
trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1),
[green
trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1).
- Full browser tests used real Claude and the production tool authority.
Reads completed. A write stopped with “Approval required.” No document
was created and no automatic retry was scheduled. All nine checks
passed: [events and
screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc).
- A separate full browser test assigned a skill, invoked it, and
completed with a marker absent from the task prompt: [skill
evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52).
- Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests,
27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm
build`. The server/UI red-green checks and `pnpm check:token-gates` also
passed. The full suite passed in CI, including all general and
serialized test shards, all three browser shards, and runner
`check:all`. The duplicate unsharded local full-suite run was stopped
after CI passed.

Braintrust links require project access. The traces contain
provider/runner events and app outcomes. They do not contain raw model
HTTP requests.

## Risks

- `approve-reads` now allows assigned Paperclip reads. Other operations
that need approval stop the turn. Users must review the operation and
change permissions before retrying.
- The Claude settings depend on the pinned ACP and SDK behavior.
Automated tests and real Claude trials cover this boundary.
- The native Codex path is unchanged. The ACPX Codex fallback shares the
clearer approval failure handling.
- Local execution was tested end to end. Remote execution was not run.
This change has no database migration.

## Model Used

OpenAI GPT-6 through Codex. The agent used reasoning, code execution,
and browser tools. The exact deployment ID and context-window size were
not exposed in the session. Live acceptance tests used
`claude-sonnet-5`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:56:24 -05:00
DottaandPaperclip 728f7185f6 feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Self-hosted boards need a way to show occasional product
announcements.
> - An app release should not be required to publish or withdraw a card.
> - Native card controls keep publishing consistent; the hero can use a
static image or isolated HTML/CSS animation.
> - This pull request renders a validated JSON feed with native
components.
> - It stores dismissals per account on each instance, so a closed card
stays closed across companies and browsers.
> - Named staging feeds let authors test content before production
publication.

## Linked Issues or Issue Description

**Subsystem affected**

Board application shell, announcement delivery, and user preferences.

**Problem or motivation**

Operators need a small, optional announcement card. Users need reliable
dismissal state. Authors need to test remote content without changing
the production feed.

**Proposed solution**

Add one non-modal AnnouncementWell. Fetch validated JSON and
content-addressed media through the instance server. Keep card controls
native, with optional sandboxed HTML/CSS animation in the hero. Use
stable announcement IDs for dismissal, an explicit empty manifest and
quiet 404 handling. Provide a staged publishing helper and isolated
test-drive guide.

**Alternatives considered**

Hosting the entire card as a page would move navigation and dismissal
into remote content. This change limits HTML to a scriptless, isolated
visual hero and keeps controls native. Browser-only storage would lose
dismissals across browsers, so the instance stores account preferences.

**Roadmap alignment**

ROADMAP.md has no overlapping announcement feature. A GitHub title
search found no related announcement pull requests. This work implements
a maintainer-requested feature.

## What Changed

- Add shared feed types, strict validation of every object, supported
routes, expiration and version checks.
- Add a board-only current-feed API, constrained media proxy, and
idempotent dismissal API. Store the first dismissal and its company
audit entry in one transaction.
- Cache upstream data for one hour. Use conditional requests, request
deduplication, response limits, public destination checks, and a
three-second deadline. Treat a remote 404 as an empty feed with a
fifteen-minute retry cooldown.
- Keep announcement visibility stable when focus moves to browser chrome
or another app pane; only tab visibility starts a return check.
- Add a responsive native announcement card. Respect onboarding, dialogs
and toast placement. Sync pending dismissals across tabs and retry after
reconnect or return.
- Add idempotent migrations for dismissals and validated publication
IDs, design-guide examples, static and animated Storybook examples, and
focused tests. The publication registry supports offline retries without
accepting caller-invented IDs.
- Add HTML/CSS animated heroes with static posters, automatic playback,
reduced-motion handling, strict DOMPurify validation, an empty iframe
sandbox and CSP that blocks scripts/network resources.
- Add validated staging publication, content-addressed assets, an empty
production manifest, preview fixtures, and authoring/operator
documentation.

## Verification

- The preceding implementation passed 98 targeted
shared/server/publisher/route/OpenAPI/UI tests and 127 tests including
the master rebase. The playback-control removal passes all 21
announcement UI tests, covering the rendered sandbox, fallback, reduced
motion, dismissal and slow/stale state lookups. The preceding
shared/server tests cover HTML validation and response sandbox headers.
- The playback-control removal passes UI typecheck, production UI build,
Storybook build and token gates locally. Browser verification confirms
the animated card has only its dismiss button and two links, with no
page errors. The full canonical CI matrix passed on current head
`00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two
optional Storybook deployment checks skipped. This run needed no
retries. Greptile reviewed this same head at 5/5 with no outstanding
findings.
- The local canonical general-server run passed 12,063 tests before
reporting embedded-PostgreSQL startup failures in an unrelated fixture.
All 31 tests in that fixture passed across isolated retries. The UI
group passed 6,219 tests and other workspace groups passed 3,201; two
CLI database-startup failures also passed individually. Serialized
server suites were verified by the full CI matrix rather than repeating
them locally. No source changes were needed for these environment
failures.
- The real S3/CloudFront staging manifest and both media asset headers
were verified. Production remains empty/unpublished. The guide
distinguishes the preview host's disabled edge cache from production
cache requirements.
- In the isolated test-drive, the animation visibly moves without
playback controls. A 390×844 browser viewport keeps the card above
navigation. Reduced motion makes no animation request. Both themes
render correctly and browser page errors are empty. Browser fault
injection verified that scripts cannot execute and CSS cannot make
network requests; a missing animation leaves its poster and controls.
- Refresh leaves the animated card visible. Closing it persists after
reload and the API returns null. Earlier live checks verified dismissal
across browsers, company-relative CTA navigation, modal
deferral/restoration, and new-ID eligibility after restarting the same
database.
- The deployed empty feed and a real remote 404 return HTTP 200 with
null from the board API, with a usable dashboard and no announcement
popup or browser warnings.
- Authoring documentation covers staging, animated HTML constraints,
test-drive, withdrawal, ID reuse and cache-refresh steps.

## Risks

- Animation supports self-contained visual HTML/CSS and inline SVG,
without JavaScript or external resources. A static image is required.
Older builds that do not recognize the optional animation field quietly
hide that unsupported feed.
- The default feed makes an outbound request from an instance when a
board is used. Operators can disable it. Requests contain no account
IDs, company data, cookies or interaction events.
- Feed publication and withdrawal can take about 65 minutes to reach
returning users because of CDN and instance caches. Expiration also
removes visible cards locally.
- Dismissals follow an account within one instance. No-login instances
share the existing local-board identity. Separate installations do not
share state.
- Both tables are additive. A unique key prevents duplicate dismissals;
the transaction prevents duplicate first-dismissal audit entries. The
publication registry retains only validated IDs. AGENTS.md and the
implementation spec document the required exception to company scope for
these instance-level records.
- Publication was limited to separate public staging prefixes on the
existing preview host. Production remains empty/unpublished. No AWS
policies or infrastructure were changed.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and
context-window size are not exposed in this session. Capabilities used:
reasoning, code editing, shell execution, tests, browser interaction,
and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:19:54 -05:00
DottaandPaperclip f1d57863d2 fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels.

Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-14 10:06:06 -05:00
DottaandPaperclip 4cd7b40255 fix: recover new messages after historical native runs stop (#13405)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native recovery must distinguish a fresh user request from replay of
failed work.
> - Older runs can lose their process fields before a local stop receipt
exists.
> - A suspended durable session can still prove that the exact runner
and provider session are idle.
> - The message admission path ignored that evidence and kept new user
messages blocked.
> - This pull request uses the existing exact-state verifier for those
historical runs.
> - The user can start one fresh turn while the old history and unknown
outcomes remain intact.

## Linked Issues or Issue Description

**What happened?**
A user sent a new message after a native run exhausted recovery.
Paperclip saved the message but said the previous run had no verified
stop record. The old runner was suspended, with no active provider turn
or pending output. Its process fields had been cleared before stop
receipts were added.

**Expected behavior**
A new user message starts a fresh turn when the exact retained session
proves it is suspended and the other execution gates pass.

**Steps to reproduce**
1. Retain a failed native run with a terminal controller, cleared
process fields, and no process receipt events.
2. Retain its exact suspended runner state and idle provider state. Keep
its recovery hold.
3. Send a new user comment. Before this fix, admission returns no
successor.

**Paperclip version or commit**
Reproduced on master at d351e08de.

**Deployment mode**
Self-hosted server. The regression uses an embedded PostgreSQL test
database and real durable state files.

Related work: Refs #13270, Refs #13338. This adds compatibility for
older stopped runs. Refs #13332 concerns separate recovery-hold scope
rules.

## What Changed

- Add exact suspended-state evidence to explicit native message
admission for runs that predate process receipts.
- Reuse the existing failed-retry verifier for run, runner, workspace,
provider identity, and pending-work checks.
- Reject this fallback if any server-authored process receipt or launch
event exists.
- Add a red/green regression with real message admission, duplicate
delivery, dry-run behavior, and blocked-state cases.
- Document the new-message recovery rule.

## Verification

- Red: the historical suspended-state regression failed on master
because admission returned null. Eight rejection cases passed.
- Green: 469 tests passed across explicit native continuation and native
session execution.
- Full repository `pnpm -r typecheck` and `pnpm build` passed locally.
- The complete Vitest CI matrix and all browser test shards passed. The
duplicate local `pnpm test:run` was stopped after these CI results; it
did not complete locally.
- The first runner verification worker received an infrastructure
shutdown signal during compilation. The single retry passed.
- Greptile reviewed commit `6836d7310` at 5/5 with no findings.
- Inspected the affected server's database and durable state read-only.
It has the historical missing-PID shape and an exact suspended runner
with no active provider turn, pending tools, or undelivered output.

## Risks

- Incorrect idle evidence could allow overlapping work. The fallback
requires an exact suspended root and rejects active or pending work,
mismatched identities, missing files, and newer process evidence.
- Normal task, controller, lease cleanup, decision, and active-run gates
remain in force.
- This does not resume old provider actions or reset recovery attempts.
Unknown outcomes remain unknown.
- No schema change or deployment is included.

## Model Used

OpenAI Codex, GPT-6. The session does not expose the exact model
snapshot or context-window size. Used reasoning, repository inspection,
code editing, shell execution, and test tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 09:45:34 -05:00
Devin Foley 13368c5183 fix: unblock clean-machine onboarding for api_key AI connections (nightly smoke) (#13372)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The release pipeline gates each nightly on a Docker onboarding
smoke. The smoke proves a clean machine can finish onboarding and hire
the first agent.
> - #13247, #13248, #13344, and #13351 changed the Connect step. Connect
now creates an AI connection that the server verifies live with the
provider.
> - The managed adoption check also demanded a CLI hello probe. A clean
machine has no provider CLI and cannot complete a subscription login.
Onboarding dead-ends and the nightly gate fails.
> - This pull request lets a live-verified API key adopt on the engine's
own verdict, and re-verifies the key with the provider at adoption time.
> - It also drives the release smoke through the API-key path against a
provider mock that lives inside the test harness.
> - The benefit is a green, deterministic release gate with no paid
credential in CI, and a working first run for API-key users on clean
installs.

## Linked Issues or Issue Description

No public issue exists. The failure surfaced in the nightly release
gate. Related PRs (no duplicates found): #13247, #13248, #13344, #13351
(the Connect changes), and #12423, #12135, #12151 (earlier release-smoke
updates).

**What happened?**

The nightly Release cut failed its gate: [run
34749840498](https://github.com/paperclipai/paperclip/actions/runs/34749840498),
job `smoke_nightly / smoke`, on published canary `2026.913.0-canary.2`.
The wizard never left the "Connect a model" step. The subscription path
waits for a human to run `claude auth login` on the server. The API-key
path saves and live-validates the key, but the environment test then
fails with `Command not found in PATH: "claude"` and
`ai_connection_validation_incomplete`, and the wizard blocks the hire.

**Expected behavior**

A clean machine with a provider-accepted API key completes onboarding
and hires the lead agent. The release smoke passes without a real paid
credential in CI.

**Steps to reproduce**

1. Run `scripts/docker-onboard-smoke.sh` with
`PAPERCLIPAI_VERSION=2026.913.0-canary.2`.
2. Sign in, complete onboarding to "Connect a model", select "Use API
key instead", pick Claude, enter a valid API key, and press Connect.
3. The environment test fails on the missing `claude` CLI and blocks the
hire.

**Paperclip version or commit**

`2026.913.0-canary.2` (nightly candidate `c9e3bb7ca`).

## What Changed

- `server/src/routes/agents.ts`: `testManagedEnvironment` no longer
forces the CLI-lane hello probe for a resolved `api_key` binding. It
re-verifies the key against the provider's live endpoint instead (the
same `validateAiApiKey` check the save performed, which needs no CLI). A
key the provider rejects fails adoption with
`ai_connection_api_key_rejected`. Subscription adoption keeps the strict
hello-probe requirement.
- `scripts/docker-onboard-smoke.sh`: the harness now serves
`api.anthropic.com` itself. A sibling container (the already-built smoke
image) runs a small HTTPS mock. The app container gets `--add-host` for
that one hostname and trusts the mock's certificate through
`NODE_EXTRA_CA_CERTS`. The private key stays mode 600 in the mock
container; the app container mounts only the certificate. The mock
serves only `GET /v1/models` and returns 404 for every other path.
`SMOKE_PROVIDER_MOCK=false` disables it.
- `tests/release-smoke/docker-auth-onboarding.spec.ts`: the spec drives
the API-key path — switch the credential mode before the source tile
(the link hides when the row collapses), enter the key, and Connect.
Loopback targets use a placeholder key that the mock accepts. Any other
target must set `PAPERCLIP_RELEASE_SMOKE_ANTHROPIC_API_KEY`, and the
test fails on arrival without it.
- `server/src/__tests__/agent-test-environment-routes.test.ts`: three
new route tests cover accepted keys (no CLI probe consulted),
provider-rejected keys, and subscriptions that cannot complete a hello
probe.

## Verification

- `npx vitest run src/__tests__/agent-test-environment-routes.test.ts` —
26/26 pass.
- `npx vitest run src/__tests__/ai-connections.test.ts
src/__tests__/ai-legacy-compatibility.test.ts` — 42/42 pass.
- `tsc --noEmit` reports no errors in the touched files.
- Full local harness + suite run against the exact failing canary: the
app container reaches the mock (request visible in the mock log), the
placeholder key validates, and the connection saves as the default. The
flow then stops at the forced CLI hello probe — the exact server check
this PR removes, still present in the published canary. The next canary
that includes this fix is the end-to-end proof.
- Hardening check: from inside the app container, the mock answers with
status 200 and `key.pem` is not visible.

## Risks

- Behavior shift: `api_key` adoption no longer requires a CLI hello
probe. It re-verifies the key with the provider at adoption instead.
Subscription adoption is unchanged.
- The mock returns 404 for unexpected provider calls, so a future
onboarding change that calls a new endpoint fails the smoke loudly
instead of passing silently.
- Release-smoke runs against non-loopback targets now require an
explicit key and fail fast without one.
- No database migration. No dependency change. No provider routing
change.

## Model Used

Claude Fable 5 (`claude-fable-5`) through the Claude Code CLI, with
extended thinking and tool use (shell, file edits, Playwright runs,
GitHub CLI). No other models were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-13 17:02:02 -07:00
DottaandPaperclip f2c5e54dca fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path

> - Paperclip lets people manage AI agents and their tasks.
> - A task keeps its instructions, progress, and files when its assigned
agent changes.
> - The replacement runner lost the interrupted run's context and could
overwrite an existing draft.
> - A saved message also stayed attached to the former agent and could
reopen the task after the replacement finished.
> - File tasks could report Done with only a local path that the user
could not open.
> - This pull request transfers handoff context and saved messages, and
makes requested files accessible through the existing attachment
contract.
> - Users can change agents and collect completed work without repeating
instructions or confirming bookkeeping.

## Linked Issues or Issue Description

Refs #13338. Builds on merged #13354 for queue admission and #13353 for
remote workspace retry. #10123 concerns restricted recovery-model
escalation; this change instead covers ordinary native handoff and file
completion.

**What happened?**

Codex wrote a draft before a user assigned the task to Claude. The
replacement lacked continuation context and replaced the draft. A queued
user message could later restart the former agent and reopen the
completed task. Separately, a runner could finish a requested file but
return only a machine-local path. Remote native runs had no bound file
publication tool.

**Expected behavior**

The replacement reads and preserves existing work, receives saved
messages once, and keeps each message's author. The former agent stays
stopped. A requested file has a working attachment or accessible work
product before Done. Text-only tasks do not require attachments.

**Steps to reproduce**

1. Ask Codex to save three newsletter names and then wait.
2. Queue an instruction to keep those names and expand the draft.
3. Use Interrupt and assign to select Claude.
4. Verify the original names survive, the result has a working download,
and only the source and replacement runs exist.
5. Ask either provider for a Markdown checklist and open the file from
its completed response.

## What Changed

- Carry the exact same-task interrupted run's summary, semantic
receipts, and history into handoff context. Tell the replacement to
inspect existing files before editing.
- Adopt saved ordinary task comments into the successor's receipt under
the task lock. Preserve authors and separate mention, chat, and
interaction contracts.
- Prevent a former-assignee comment wake from reopening a completed task
or starting a stale execution.
- Reject workspace-only, fabricated, and cross-task file completion
references with actionable runner feedback. New file output also needs a
matching current-run publication receipt and asset filename/size/hash,
or an accessible work product registered by the current run. Prior
output can remain context alongside a current file, or be verified and
re-registered internally. Authorized chat attachment reuse retains its
verified current-run clone receipt; older receipt shapes require an
intact matching source.
- Bind remote file reads to the active environment runner and reuse the
existing attachment and work-product publication path.
- Enforce workspace confinement, regular single-link files, stable
identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the
native session fingerprint for the updated tool contract.
- Contain rejected remote signals and protocol-failure cleanup,
including logging failures. Preserve the original cleanup rejection for
its owner; a rejected operation never supplies stop acknowledgement or
cleanup proof.
- Allow exactly one maximum-size base64 file through the native SSH
command adapter, preserving a finite output cap.
- Document handoff and accessible file completion rules.

## Verification

- Each observed bug has a failing regression before its fix. Final
post-rebase integration passed 732 tests across 13 files before the
final receipt and signal guards; final affected results are below.
- Publication provenance and compatibility: 8 provenance regressions and
2 compatibility regressions failed before their fixes; the final four
affected suites pass 53 tests, including mixed old/new references and
real authorized chat reuse. Controls cover old attachments and work
products, filename/size/hash/origin mismatch, missing/wrong receipts,
current-run publication, same-run durable proof, internally
re-registering preserved bytes, and no-new-file follow-ups.
- Remote signal rejection: the real Node subprocess previously exited 1
when the production launcher signalled a deleted sandbox. It now stays
alive for both a failed signal and failed logging; all 349 executor
tests pass. The failed signal still provides no termination proof.
- Remote file reader and SSH command boundary: 33 tests passed,
including real Linux descriptor reads and the actual SSH adapter
subprocess output cap (network executable replaced by a deterministic
fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap
fails.
- Live local Claude and Codex Stop journeys preserve the saved file,
deliver queued instructions once, and reach Done with two total runs.
The handoff journey preserves the original names and download with
exactly two runs. Both local providers deliver exact checklist files
without a completion confirmation.
- Live combined Daytona verification passed: the original failed task's
Retry reused its sandbox; a selected Git subfolder produced an exact
downloadable file; warm and deliberately resumed Claude runs took about
33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134
seconds of contention/backoff. Both cloud downloads retained exact bytes
after the two owned sandboxes were deleted.
- The sandbox-deletion retest identified a separate ignored promise in
protocol-failure cleanup. Two real Node subprocess regressions failed
under fatal unhandled-rejection policy before the fix; all 36 protocol,
lifecycle, and integrity tests now pass. The original close promise
still rejects to its owning runtime. The final live retest passed: a
normal Claude Daytona task completed in 132.352 seconds, then its
sandbox was deleted. Thirteen samples over 361 seconds confirmed the
same controller stayed healthy, the task stayed Done with unchanged run
IDs, and its attachment retained exact bytes. The post-deletion browser
download passed with zero page errors; all five owned sandboxes are
confirmed absent.
- Full repository typecheck and build passed on final commit
`de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads.
[Final-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758)
passed: 32 successful checks and two conditional skips. The earlier
mixed-source full local test invocation was deliberately stopped before
rebase, so no pristine green full local aggregate is claimed. Its known
failures passed in later affected suites.

## Risks

- Handoff may adopt only ordinary comments from its validated former
owner. Other delivery contracts must remain independent.
- File verification fails closed if a remote file changes during
reading. The runner must retry publication or explain a blocker.
- The updated session fingerprint starts a fresh provider process where
needed to install the new tool contract.
- No schema migration or historical status reconciliation is included.

## Model Used

OpenAI `gpt-6-astra` through Codex, with reasoning, code execution,
browser testing, and tool use. The context-window size is not exposed in
this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-13 08:41:20 -05:00
DottaandPaperclip 827ba8a434 fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox runs prepare credentials and files before an agent starts.
> - Stop must work during that preparation.
> - ACPX registered cancellation, but did not handle it while a setup
command was waiting.
> - Daytona cleanup waited for that same command before stopping the
sandbox.
> - This change stops the run's sandbox first and requires proof before
abandoning setup.

## Linked Issues or Issue Description

**What happened?**

Stop left a sandbox run active when remote credential setup stalled. The
run kept its connection lease until the sandbox was stopped separately.

**Expected behavior**

Stop terminates the selected run's sandbox, prevents later setup from
launching the agent, and lets run cleanup finish. It must not report
success without proof from the provider.

**Steps to reproduce**

1. Start an ACPX agent in Daytona.
2. Hold a command during remote credential or file setup.
3. Select Stop before the agent starts.
4. Before this fix, cleanup waits for the held command and never reaches
sandbox stop.

Related: #13351 exposed this during connection acceptance testing.
#12150 addresses scheduler load and session initialization limits, a
separate startup problem.

## What Changed

- Handle cancellation during ACPX sandbox preparation with a host-owned
stop callback.
- Pass an explicit active-work cancellation flag through environment
cleanup.
- Stop Daytona before draining setup commands. Keep normal graceful
cleanup.
- Require an exact run and lease termination receipt. Keep ownership of
outstanding requests when stop cannot be verified.
- Reject late setup work and defer sandbox resume until old requests
settle.
- Add regression tests and document the cancellation boundary.

## Verification

- Red: both the stalled ACPX setup test and the Daytona cancellation
test failed before the fix because Stop never reached the provider.
- Green: adapter and Daytona suites passed, along with cancellation
boundary and database-backed receipt/isolation tests.
- Two real Daytona probes ran a five-minute setup command. Cancellation
returned matching stopped receipts in 6.01 and 6.533 seconds. No later
setup command ran. Both test sandboxes were deleted.
- Repository typecheck and build passed locally. The full required CI
suite passed, including all server/workspace tests, browser shards,
Paperclip Runner verification, and canary packaging. The duplicate local
full-suite run was interrupted after CI passed; it is not claimed as a
completed local pass.
- No UI changes. The live probe uses the actual adapter cancellation
boundary and Daytona plugin; it is not a browser acceptance test.

## Risks

- Provider stop failures remain unacknowledged. The adapter keeps
ownership while its original requests remain active.
- A retry can receive a settling-work error until old provider requests
finish.
- Cancelled sandboxes are stopped and retained under their existing
provider expiry policy.
- Local execution and cancellation after the agent turn starts keep
their current behavior.
- Other providers must return an exact termination receipt to permit
early setup cancellation. No database migration.

## Model Used

OpenAI GPT-6 through Codex, with code execution and tool use. The exact
deployment identifier and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-13 08:35:43 -05:00
DottaandPaperclip c9e3bb7ca4 fix: preserve queued work after native Stop and honor steering support (#13354)
## Thinking Path

> - Paperclip lets people manage AI agents and their tasks.
> - The runner owns execution, while the task keeps user instructions
and status.
> - Stop must stop the current response without losing instructions that
the user already sent.
> - The queue stored its original reason inside saved context. Recovery
checked the outer deferred reason and left the message waiting.
> - Claude also exposed Steer through a shared method even though its
driver did not support it. A rejected request could remove its own error
row.
> - This pull request keeps queued work until execution has stopped,
uses the driver's real capability, and preserves completion event order.
> - Users can continue work without repairing task state or repeating
messages.

## Linked Issues or Issue Description

Refs #13338. Related recovery work: #13353.

**What happened?**

A message sent during a native run stayed queued after Stop. Claude
exposed an unsupported Steer action. A steering failure could hide the
queued row and its error. A terminal event could also precede the final
provider result, and subtree Stop omitted the board actor.

**Expected behavior**

Stop ends the current execution. Once Paperclip proves that execution
has stopped, it delivers the saved instruction once through normal
admission. Pause and recovery holds still prevent dispatch. Unsupported
controls stay disabled, and a rejected action leaves an actionable error
visible. Final results precede terminal events.

**Steps to reproduce**

1. Start a Claude or Codex task that writes a file and then waits.
2. Send a follow-up instruction while it runs.
3. Press Stop. Check that the queued instruction runs once and preserves
the file.
4. Check Claude's Steer control and simulate a server rejection on the
only queued message.

## What Changed

- Recover saved native comments after acknowledged Stop using their
original wake reason.
- Require durable remote termination receipts or verified local process
termination before dispatch.
- Preserve actor identity, queued-message deduplication, Pause, and
recovery gates.
- Derive steering support from the driver descriptor and reject
unsupported calls.
- Keep the queue mounted until a steering request succeeds so its error
remains visible.
- Emit provider results before terminal events and pass the board actor
into subtree Stop.
- Document Stop and steering behavior.

## Verification

- Final-head continuation suite: 104 passed, after failing regressions
for saved wake reasons, cleanup proof, and deduplication of every queued
message. Steering UI: 121 passed. Driver capability: 26 passed. Runner
backend/transport: 205 passed; Rust library: 285 passed.
- Real Claude and Codex browser journeys both preserved the saved file,
delivered the queued instruction once after Stop, and reached Done with
exactly two total runs. The process Stop browser fixture also passed.
- Local full repository typecheck and build passed on `afaa35139`;
server typecheck and the affected 104-test suite passed after the final
queue changes. Token gates passed. Final-head CI verifies the complete
integrated source.
- Local aggregate evidence has explicit limits: the general-server
invocation overlapped the queue fixes and finished with 11,989 passed, 2
failed, and 80 skipped; both failures are covered by the final 104-test
pass. The UI and CLI then passed all 6,184 and 485 tests; the complete
145-file serialized rerun passed all 2470 tests. The shared-package lock
fixture passed unchanged on rerun, but the package phase subsequently
stopped at an embedded-Postgres bootstrap resource failure. No single
pristine green local full aggregate is claimed.
- Greptile reviewed `fa66e2bd5` at [5/5 with no unresolved
findings](https://github.com/paperclipai/paperclip/pull/13354#issuecomment-5650334242).
[Final-head CI completed
successfully](https://github.com/paperclipai/paperclip/actions/runs/34733781888/attempts/2):
33 successful checks, 2 conditional skips, including all server,
package, UI, browser, runner, typecheck, and build gates. The first
attempt hit a preview-readiness/port-collision fixture; its unchanged
local control passed 25 tests with 3 skips, and one supported unchanged
CI retry passed the affected shard and aggregate gates.

## Risks

- Queue recovery must never overlap an old execution. Unknown cleanup
state remains blocked.
- Driver descriptors are now authoritative for steering; a wrong
descriptor disables the action instead of attempting it.
- No schema migration or historical status reconciliation is included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, browser
testing, and tool use. The exact hosted model ID and context-window size
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 22:29:54 -05:00
DottaandPaperclip 6809314a3f fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path

> - Paperclip manages AI agents and their tasks.
> - Sandbox tasks need a workspace and a provider session before work
can start.
> - A selected Git subfolder is a valid workspace, but it is not a
repository fetch source.
> - A resumed sandbox must keep one resource identity when the provider
fills in its region.
> - Workspace reuse does not prove that a provider session has started.
> - A recorded continuation must not leave an old failure blocking Retry
for a newer attempt.
> - This pull request fixes those setup and recovery boundaries while
preserving ownership checks.

## Linked Issues or Issue Description

**What happened?**

A Daytona task failed before provider startup when its project folder
was inside a parent Git repository. After the folder was repaired, Retry
resumed the same sandbox and verified its workspace sentinel, but
workspace preparation reported that the lease was no longer active. A
later attempt could fail because the resumed workspace had no provider
checkpoint. The old run also kept a recovery projection after the server
recorded its explicit successor, hiding Retry for the new failure.

**Expected behavior**

Sync the selected folder without importing parent files or history. Keep
the existing sandbox identity stable. Allow a new provider session only
with proof that its exact session has never started. Let the current
failed attempt retain Retry once the old recovery has a recorded
successor.

**Steps to reproduce**

1. Select a subfolder of a Git repository as a project workspace and
start a Daytona task.
2. Leave the target region unset. Release its reusable sandbox after
setup fails, then resume and realize the workspace.
3. Retry a provider setup that failed before any session directory or
checkpoint was created.
4. Record an explicit successor for a native failure, fail that
successor during setup, and inspect Retry in the task thread.

**Paperclip version or commit**

Reproduced against `8d1f0c20a` with new failing regressions before each
production fix.

**Deployment mode**

Local controller with a Daytona environment. Automated tests use
deterministic provider fixtures and real filesystem operations.

Related: #13338, #13163, #13264, #13349. The open work-folder stack in
#13264 includes a broader fresh-session authority change. This patch
addresses the independently reproduced startup failure with existing
durable bootstrap proof and atomic directory creation. It does not
include the work-folder migration or credential changes from that stack.

## What Changed

- Classify only the selected Git repository root as a fetch source. Sync
subfolders as directories and preserve their enclosing Git ignore rules
on upload and restore. Recognize the shared scheduler’s completed
non-repository result so ordinary folders still sync; timeout,
cancellation, and output-limit failures remain closed.
- Remove the placement target from Daytona account cache identity. Keep
API endpoint, credential digest, company, environment, driver, and
sandbox ID boundaries.
- Preserve closed-lease admission until a sentinel-verified resume
reopens that same resource.
- Show the already-recorded explicit successor of a resolved recovery.
Keep the old failure evidence and unresolved holds. No historical status
writes occur.
- Permit a resumed workspace to create a new session directory only with
matching durable identity, zero connections and events, untouched
bootstrap commands, no backup, and an absent remote session. Claim the
directory atomically. Existing, partial, or ambiguous state still blocks
startup.

## Verification

- Final head `c43c7403f`: [CI completed
successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2),
with 32 successful checks and 2 conditional skips. This includes every
server, UI, package, serialized-route, browser, native-runner,
typecheck, and build gate. Greptile reviewed the same head at [5/5 with
no unresolved
findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168).
- New regressions failed before each of the four production fixes.
Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git
regression also failed before its fix; 135 affected Git/sync/Codex tests
then passed.
- Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native
executor, projection, and TaskChatThread: 494 passed, including
existing/partial state, wrong identity, prior connections or turns,
backups, unavailable proof, and unresolved recovery controls.
- Local full-repository typecheck, build, and token gates passed on the
final head. Local CLI: 485 passed. Complete single-worker package rerun:
3,224 passed, 19 skipped.
- Local verification is an aggregate with recorded retries, not one
pristine green invocation: the general-server run began on `e721a920a`
and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex
scheduler failure is fixed above; the socket and workspace-runtime
timeout failures passed unchanged in focused reruns. Both Inbox failures
passed unchanged in the full 27-test Inbox file; database/shared-package
failures passed in the single-worker package rerun. The supplemental
local serialized-route rerun remains in progress; all five corresponding
final-head CI lanes passed.
- The first final-head CI attempt hit a Daytona fixture-readiness race
and a signoff-browser heartbeat receipt timeout. One supported unchanged
failed-job rerun passed both and the aggregate gates. Live combined
user-journey verification is tracked in the related follow-up; this PR's
provider tests use deterministic fixtures and real filesystem checks.

## Risks

- A selected subfolder uses directory sync and does not carry parent Git
history. Its ignored files stay local.
- The target region remains a creation setting and part of workspace
reuse policy; it does not split the account identity of an existing
sandbox.
- Incomplete or conflicting provider state still fails closed. This
change does not erase a session, infer completed work, bypass a user
decision, or replay uncertain actions.
- A later remote setup failure can leave a claimed partial session
directory. It remains blocked rather than being overwritten.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact hosted model ID and context-window size are not
exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 22:19:42 -05:00
DottaandPaperclip 8d1f0c20af fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - AI Connections select the account used for each run.
> - A responsible-user binding must follow the person whose work the
agent performs.
> - The saved sign-in method currently blocks users with another method
for the same provider.
> - This pull request resolves a personal default by company, user, and
provider.
> - Each user can use a subscription or API key with the same bot and
model.

## Linked Issues or Issue Description

Refs #13247, #13248, #13346, #13347.

**What happened?**
A bot configured with a Claude subscription rejects another responsible
user’s Claude API key. Inline repair also limits that person to the
original sign-in method.

**Expected behavior**
The same bot uses each responsible user’s default Claude account,
whether it is a subscription or API key. Explicit shared account
selections remain fixed.

**Steps to reproduce**
1. User A connects a Claude subscription and creates a bot using the
responsible user’s connection.
2. User C connects a personal Claude API key.
3. User C runs the same bot. Before this fix, credential resolution
fails.

## What Changed

- Add a personal provider-default table. Preserve legacy per-method
preferences and backfill the most recently updated preference, including
unavailable defaults. Repeated migration does not replace a selection. A
database trigger propagates old-server default updates without treating
new accounts as replacement defaults.
- Resolve responsible-user bindings by provider. Retain the method as a
wire compatibility hint for old servers. Explicit selections still
require the exact method and grant.
- Use the selected account’s method for credential isolation, refresh
locking, and run attribution.
- Update onboarding, agent setup, the picker, and inline task repair.
Keep existing authentication components and harness/model settings.
- Add mixed-method runtime, migration, repair, and Storybook coverage.
Include upstream’s duplicate Anthropic option fix through the base
branch.

## Verification

- Focused resolver, migration, connection-intent, onboarding, agent
setup, model, and connector UI suites: 331 tests passed.
- Onboarding and new-agent regression suites passed during the initial
focused run.
- UI typecheck, token gates, and Storybook build passed.
- Live browser checks passed for Claude and Codex API-default execution,
switching both back to subscriptions, and both existing shared-account
bots. Bot configuration remained unchanged.
- One Daytona startup command stalled before Claude launched. The test
run was cancelled, its sandbox stopped, and the same account/task passed
on retry. Startup cancellation remains a separate environment finding;
this PR does not change that command transport.
- Browser review: all eight assertions passed in the new mixed-method
story, including shared selection, return to responsible-user selection,
and unchanged harness/model.
- Repository build and typecheck passed after refreshing upstream
dependencies. Final resolver and historical rollback verification: 38
tests passed. All latest-head CI gates passed, including
server/workspace/serialized suites, browser E2E, build, typecheck, and
runner verification. The extra serial local full-suite run was stopped
after equivalent CI passed; focused local checks completed.

## Risks

- Users with both historical method defaults get their most recently
updated preference as the initial provider default. They can change it
explicitly in Connections.
- A revoked or unavailable default blocks. Connecting an additional
account does not silently replace it.
- Existing legacy authentication is unchanged. Managed responsible-user
bindings intentionally stop pinning a method.
- Live staging: the same Claude and Codex bots completed real API-key
runs after changing only the personal default, then completed
subscription runs after restoring the original defaults. Read-only
database verification confirms unchanged bot configuration and actual
method attribution. Distinct-user concurrency is covered by automated
real-database tests with synthetic credentials, not two live human
logins.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution,
and browser interaction. The runtime does not expose a more specific
model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 20:27:26 -05:00
DottaandPaperclip 422287eecd fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task messages, provider execution, and
task outcomes.
> - First-time user tests exposed gaps in recovery, completion
permissions, message delivery, and Stop behavior.
> - These gaps left usable output hidden, completed work waiting for
bookkeeping, or safe work unable to continue.
> - This pull request fixes the shared lifecycle and receipt paths while
preserving process ownership and action checks.
> - Users can continue work with accurate task state and durable
messages.

## Linked Issues or Issue Description

**What happened?**

A stopped local Codex execution could remain blocked even after its
processes had stopped and its complete transcript proved that no
external action needed replay. Claude under Conservative permissions
could fail to call task completion tools. Recovery could reuse an
assistant item ID and overwrite prior output. A delivered comment could
remain marked uncertain after navigation. Stop could look like Pause or
a new recovery incident. Workspace contention could look like
cancellation. A direct reply reopening Done could enter a clarification
loop.

**Expected behavior**

Recover automatically only with verified termination and complete action
receipts. Preserve answers and messages. Keep task completion available
under Conservative permissions without broad tool access. Show crashes
as Blocked, actual human decisions as In Review, and ordinary workspace
contention as waiting. Stop the current response and allow a new
direction.

**Steps to reproduce**

1. Create ordinary response tasks with local Codex and Claude Code, then
send follow-up messages through the task composer.
2. Interrupt a disposable local Codex runner during text-only work.
Verify automatic continuation and retained output.
3. Stop a response, send a new request, answer a clarification, and
reopen completed work with another message.
4. Navigate or reload while a comment submission is pending. Confirm the
exact persisted request receipt settles it without removing newer draft
text.
5. Run two tasks in a shared Daytona workspace. Confirm waiting does not
appear as failure.

**Paperclip version or commit**

Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`.
Current integration base: `6cef9743c`. Both operator-interruption and
workspace-waiting guards are preserved; native restart and legacy
permission rules remain documented.

**Deployment mode**

An isolated source-built test-drive instance, with real local Codex and
Claude Code providers and disposable Daytona environments.

Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163.
This PR addresses additional failures from ordinary task journeys,
including controller restart handoff and repeated warm sandbox setup.
Historical task status reconciliation is excluded.

## What Changed

- Persist runner ownership immediately at spawn and resume an explicitly
adopted runner even when the controller crashed before the first driver
checkpoint. Detach the controller safely across graceful restarts,
including session startup. Prevent an old finalizer from suspending or
signaling an adopted runner. Checkpoint idle warm sessions before
shutdown. Preserve the same run and queued follow-up messages.
- Scope saved legacy queue successor checks to the queue owner while
preserving ordinary task locks, operator identity, assignment gates, and
exactly-once delivery.
- Preserve managed Codex credential files when an old session is
detached for restart; normal owned cleanup still copies refreshed auth
back and removes the scoped copy.
- Reuse the bound warm shared sandbox and fully verify an existing
staged provider pack before using it. This avoids repeated uploads when
the pack is already valid.
- Add a narrow local Codex replacement path with stopped-process proof,
a closed transcript inventory, exact completion receipts, and
fresh-session lineage. Preserve no-replay holds when evidence is
incomplete. Recovery may clear only the same run's recorded Blocked
status version; manual re-blocking and dependency changes invalidate
that receipt, while queued comments do not. Later blocks stop scheduled,
queued, and final dispatch; queued/final checks re-read dependencies
even when the task status stays In Progress.
- Permit only task delivery and human-input tools through the isolated
Claude runner's exact task bridge.
- Scope assistant item identity to the provider turn and ignore only
authority-free Codex skill-change notifications during startup.
- Reconcile composer submissions by client request ID across response
loss, navigation, and reload. Retain text typed during delivery.
- Keep acknowledged run-only Stop neutral and show workspace contention
as waiting. Project exhausted native failures as Blocked.
- Restore the guarded task-page retry action for failed legacy runs,
including the server-supported explicit new-attempt path for stopped
conversation adapters. Preserve native/process recovery holds and avoid
promising Retry while a decision or execution gate hides it.
- Refresh delivered artifacts and handle direct user replies that reopen
completed work without a clarification loop.
- Check the embedded PostgreSQL PID, data directory, and actual port
before connecting or migrating.
- Document accepted behavior and add focused regressions at lifecycle,
route, transcript, and UI boundaries.

## Verification

- Final head `fece606ac2` passes the complete GitHub CI matrix: **34
green checks, two expected Storybook skips, no failures or pending
checks**, including `ci / verify`, `ci / e2e`, full runner verification,
typecheck, build, every server/workspace shard, and all browser shards.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/34727183287).
Greptile is **5/5 with no open findings**. The final two commits only
refine test fixtures; both affected suites pass 24/24 locally and in CI,
with server typecheck green.
- Complete local Vitest coverage uses the canonical groups/shards: all
635 general server suites, all 145 serialized suites, and all workspace
packages. The aggregate began on `0a8001c18` while the final queue fix
arrived: 23,903 passed, five failed, 87 skipped. The five
port/socket/timing failures passed unchanged in follow-ups (60 tests in
the exposure/file suites and 412 tests covering the serialized failures
and unrun tails). The final queue/operator-identity suites separately
passed 52/52. This is aggregate coverage plus explicit reruns, not a
pristine single-command final-head run.
- After integration with current master,
queue/operator-identity/continuation suites passed 162/162 and affected
UI suites passed 140/140. ACP Stop/continuation and legacy
task/Inbox/message browser suites passed 9/9, including both task
recovery Retry and thread Try again, automatic saved-message delivery,
exactly one new run, Done, and retained output after reload. The default
process Stop/Pause/Resume browser case passed (the native-provider case
is opt-in and skipped by default). The complete Board attachment/receipt
browser suite passed 11/11 on a disposable instance, covering both
composers, exact receipts after lost responses, no replay, bound
attachments, and newer drafts after reload.
- Blocking-intent regressions cover pre-existing Blocked, a mismatched
run/cause, an explicit manual re-block, changed dependencies, a queued
comment after failure, and a block arriving between scheduling and
provider dispatch. The negative cases reproduced before the fix. All 478
affected executor/recovery/dispatch tests passed; both database suites
ran separately after availability-probe skips in the first combined
command. The final late-dependency check passed all 143 affected
recovery/dispatch tests (zero skips) after two new negative cases
reproduced the bug.
- Focused runtime regressions cover awaited runner ownership
publication, authenticated adoption before the first checkpoint,
old-finalizer detachment, idle and busy warm-session shutdown, rejected
checkpoint propagation, provider-pack verification, and managed-Codex
credential preservation. Four managed credential detachment cases
reproduced the bug before the fix; normal owned cleanup still succeeds
exactly once.
- Live local Claude: SIGKILL 2.6 seconds into startup recovered the same
run automatically in 53 seconds, then a normal follow-up completed in 24
seconds. SIGTERM 2.5 seconds into startup preserved the same run (54
seconds) and its queued follow-up (21 seconds). Answers remained visible
and the task reached Done.
- Live Claude Daytona: a warm follow-up retained its sandbox and fell
from 121 seconds to 44 seconds. A separate cold turn took 127 seconds;
after controller shutdown and checkpointing, its follow-up completed in
33 seconds with the same sandbox, workspace, native session, and runner.
Both answers remained visible and the task was Done.
- Other live journeys covered task completion and follow-up with local
and Daytona Codex, local Codex crash recovery, Stop then new direction,
clarification response, live artifact refresh, and shared-workspace
waiting.
- Validation limits: the opt-in native composer Stop/Pause→subtree
Resume fixture exposes terminal/result ordering and subtree-cancellation
attribution bugs that can leave a child task blocked; that new finding
is assigned to a separate follow-up and is not claimed fixed here.
Default CI skips this optional native-provider fixture. Managed-Codex
credential handoff and the queue-agent integration use automated
regression evidence. Cold custom provider-pack uploads still add startup
latency.

## Risks

- Automatic replacement remains deliberately narrow: local Codex,
verified stopped identities, unchanged retained state, and a complete
text/completion-only turn. Unknown actions, partial history, or changed
ownership remain blocked.
- Claude completion permission handling changes an upstream package
patch. The exact isolated task bridge must remain pinned; unrelated
tools keep their existing permissions.
- New task failure projection changes user-visible status. No historical
status backfill or database migration is included.
- This is a broad lifecycle fix across server and UI. Live proof covers
graceful local Claude restart during startup and idle Claude Daytona
session recovery across controller shutdown. Live abrupt SIGKILL during
local Claude startup also recovered the same run. Unknown ownership or
missing action evidence still blocks reuse. Cold custom provider-pack
uploads still add startup latency; this change avoids unnecessary repeat
uploads.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, browser
automation, and tool use. The exact hosted model ID and context window
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:41:15 -05:00