Commit Graph
1532 Commits
Author SHA1 Message Date
dependabot[bot] cf8ad63c80 build(deps-dev): bump tsx from 4.23.12 to 4.23.15 (#12965)
Bumps [tsx](https://github.com/privatenumber/tsx) from 4.23.12 to
4.23.15.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/privatenumber/tsx/releases">tsx's
releases</a>.</em></p>
<blockquote>
<h2>v4.23.15</h2>
<h2><a
href="https://github.com/privatenumber/tsx/compare/v4.23.14...v4.23.15">4.23.15</a>
(2026-09-20)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>exclude bare builtins from namespace inheritance (<a
href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a">38e1588</a>)</li>
<li>expose require.cache and require.extensions to tsImport CommonJS
modules (<a
href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a">2da3407</a>)</li>
<li>make namespaced register() overloads portable for declaration emit
(<a
href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15">562c434</a>)</li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.23.15"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
<h2>v4.23.14</h2>
<h2><a
href="https://github.com/privatenumber/tsx/compare/v4.23.13...v4.23.14">4.23.14</a>
(2026-09-20)</h2>
<h3>Bug Fixes</h3>
<ul>
<li>restore the CJS bridge namespace for Node 24 require(esm) under
tsImport() (<a
href="https://redirect.github.com/privatenumber/tsx/issues/802">#802</a>)
(<a
href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f">6e5236b</a>)</li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.23.14"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
<h2>v4.23.13</h2>
<h2><a
href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.13">4.23.13</a>
(2026-08-30)</h2>
<h3>Bug Fixes</h3>
<ul>
<li><strong>cache:</strong> bound shared transform cache memory (<a
href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>)
(<a
href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e">28e1f12</a>)</li>
</ul>
<hr />
<p>This release is also available on:</p>
<ul>
<li><a href="https://www.npmjs.com/package/tsx/v/4.23.13"><code>npm
package (@​latest dist-tag)</code></a></li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/privatenumber/tsx/commit/ca66105a17a2a4c6503fe3a12b5b9ec408286011"><code>ca66105</code></a>
test: fix drive-less file URLs in ESM resolver fixtures</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/2da34075afaed43e2b7fd0aca5fbebaaf337ff3a"><code>2da3407</code></a>
fix: expose require.cache and require.extensions to tsImport CommonJS
modules</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/38e158857e50bca311be7c232a5057e5a2e5347a"><code>38e1588</code></a>
fix: exclude bare builtins from namespace inheritance</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/562c434a5c8695e74327bbeb51cfeb9b86fc7e15"><code>562c434</code></a>
fix: make namespaced register() overloads portable for declaration
emit</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/edfb1f05a3f40b879a41a03a0801c2abd3a3ecf9"><code>edfb1f0</code></a>
build: upgrade pkgroll and externalize CJS loader reference</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/70e78284837c859f09b96cd10cd71d007aa4b795"><code>70e7828</code></a>
test: upgrade tinyspy for disposable API</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/9ed2022dfa9ea1be9511fe6abcde8110c25055a7"><code>9ed2022</code></a>
ci: avoid duplicate release notifications</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/872e77ffc5e96ca5c4727e74c0694debcb26219b"><code>872e77f</code></a>
refactor: use disposables for cleanup</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/6e5236b065738d3687a06396d064774cfede390f"><code>6e5236b</code></a>
fix: restore the CJS bridge namespace for Node 24 require(esm) under
tsImport...</li>
<li><a
href="https://github.com/privatenumber/tsx/commit/28e1f12d04cd2afe1db17f8555b14fe5fb567c6e"><code>28e1f12</code></a>
fix(cache): bound shared transform cache memory (<a
href="https://redirect.github.com/privatenumber/tsx/issues/835">#835</a>)</li>
<li>See full diff in <a
href="https://github.com/privatenumber/tsx/compare/v4.23.12...v4.23.15">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-02 08:39:37 -07:00
Nicky LeachandPaperclip b2c565038b test(shared): make the worktree port registry lock suite deterministic (#12798)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Shared worktree services use lock leases and worker-thread
heartbeats
> - The lock test suite measured wall-clock timing across two threads
> - Processor contention allowed a heartbeat tick to change the value
during an assertion
> - This pull request removes that timing race and restores a regression
guard
> - The benefit is a stable test suite that still detects slow
heartbeats

## Linked Issues or Issue Description

**What happened?** The worktree port registry lock suite failed at
random under continuous-integration processor contention. The failure
reported a fresh timestamp where the test expected an old timestamp.

**Expected behavior** The suite must pass when the heartbeat runs at its
supported interval. It must also fail when the heartbeat interval
regresses.

**Steps to reproduce**
1. Run `npx vitest run src/worktree-port-registry.test.ts` in
`packages/shared`.
2. Repeat the run under bounded processor contention.
3. Set the heartbeat interval to 3000 ms and run the asynchronous
critical-section test.

**Paperclip version or commit**
`a661caf74e704f7700a8b8a1e79b76ebd04e3483`

**Deployment mode** Built from source.

**Installation method** Built from source.

**Agent adapter(s) involved** Not adapter-specific (core test).

**Database mode** Not database-related.

**Additional context** Related open pull requests are #11994, #11985,
and #11922. This pull request keeps all five tests active and does not
use `skip`, `skipIf`, or `todo`.

## What Changed

- Build the fallback-probe lock state by hand so no live heartbeat
changes the timestamp during the assertion.
- Count distinct heartbeat refreshes in the asynchronous
critical-section test.
- Close the fake probe and settle the pending lock attempt in a
`finally` block.
- Keep production code unchanged.

## Verification

- `npx vitest run src/worktree-port-registry.test.ts` — 5 of 5 tests
pass.
- `npx vitest run` — 72 files and 704 tests pass at submit time.
- `npx tsc --noEmit` — exit code 0.
- Ten target-file runs pass under bounded processor contention.
- A 3000 ms heartbeat interval fails with `expected 2 to be greater than
or equal to 3`.
- An inverted cleanup assertion exits normally in 379 ms without a
leaked worker.

## Risks

Low risk. This pull request changes one test file. It changes test setup
and assertions only.

## Model Used

OpenAI GPT-5 through Codex. Exact model ID: GPT-5. The model used tool
calls and code execution. The context window is not disclosed by the
runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:24:00 -07:00
Nicky LeachandPaperclip 9d0f7e2ddd fix(adapter-utils): make the directory merge lock crash test deterministic (#14881)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A workspace restore merges a directory, and a cross-process lock
serializes that merge
> - The lock must recover after the process that holds it crashes
> - One test proves that recovery: it kills the holder process and then
acquires the lock
> - That test failed intermittently for two independent reasons, and
this pull request removes both
> - First, it spawned the holder through the tsx command-line entry
point, which re-spawns the evaluated code in a further child process, so
the kill signal reached only the wrapper and the real holder kept the
lock
> - Second, it replaced the global clock to force a timeout, which left
the acquisition with zero real retries, so a single transient busy
result failed the test
> - The benefit is a deterministic crash-recovery test and a reliable
continuous-integration signal

## Linked Issues or Issue Description

**What happened?**

The test `recovers a killed holder even when its recorded PID has been
reused` in `packages/adapter-utils/src/directory-merge-lock.test.ts`
failed intermittently in continuous integration. The failure reported
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` with `waitMs: 3` and
`knownLocalHolder: false`. A rerun of the same job on the same commit
passed.

**Expected behavior**

The test must pass every run. It must acquire the lock after the holder
process dies.

**Steps to reproduce**

1. Check out `master`.
2. Run `npx vitest run
packages/adapter-utils/src/directory-merge-lock.test.ts`.
3. Repeat the run. The named test fails intermittently.

**Paperclip version or commit**

`32e9f3ba0ec000578936731990d23bb0e77493fa`

**Deployment mode**

Built from source. The failure appears in the general test job of
continuous integration.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Relevant logs or output**

```
ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT
workspaceRestoreLock: { ownerState: 'alive', knownLocalHolder: false, waitMs: 3,
                        ownerSameProcess: true, ownerAgeMs: 60030, ownerPredatesProcess: true }
```

## What Changed

The test file had two independent defects. This pull request removes
both.

**1. The kill signal did not reach the real lock holder.**

The test spawned its holder through the tsx command-line entry point.
That entry point re-spawns the evaluated code in a further child
process. `SIGKILL` therefore killed only the wrapper, and the process
that had opened the lock database survived as an orphan that still held
the lock. The test now loads tsx as an `--import` hook, so the spawned
process is the real holder and the kill releases the lock at once. This
also stops the test from leaking an orphan process.

**2. The forced clock left the acquisition with zero retries.**

A test helper replaced `Date.now` to force a timeout. The implementation
reads `Date.now()` one time, to compute its deadline, so that single
read consumed the forced value and every later read returned a time
already past the deadline. The retry loop therefore got one attempt and
no retries. That is correct for a test that asserts a timeout, but the
crash-recovery test asserts a *successful* acquisition, so any transient
busy result on the first attempt failed it.

The fix removes the clock replacement from the whole file and gives each
test a real, short, explicit wait budget:

- `withDirectoryMergeLock` takes a new optional wait-budget parameter.
It threads through to the lock acquisition function. The production
default is the existing 30-second budget, and no production call site
changed.
- The five tests that assert a timeout pass a real 200-millisecond
budget. Each one still times out for the real reason, because the lock
is genuinely held or the legacy lock directory genuinely exists. Each
one now exercises at least four real retries of the 50-millisecond retry
interval.
- The crash-recovery test passes a real 5-second budget. A failure now
reports the structured `ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` diagnostic
well inside the test timeout, instead of a bare test timeout.

**No test timeout increased.** Every `it(..., N)` timeout in the file
equals its value on `master`.

## Verification

- Measured the first cause rather than assumed it: the spawned wrapper
process reported one process id, and the process that opened the lock
database reported a different process id and named the wrapper as its
parent. The real holder kept the lock for about 50 to 60 milliseconds
after the kill.
- Reproduced the failure deterministically before the change, with no
artificial processor load: 15 of 15 runs failed. Confirmed the fix: 15
of 15 runs passed.
- Ran the lock test file 15 times in series: 12 of 12 tests passed every
time.
- Confirmed the clock replacement is gone: a search for a `Date.now`
override in the file returns nothing.
- Confirmed the production default is unchanged at 30 seconds, and that
the diff touches no production call site.
- Proved the diagnostic still surfaces: with a temporary edit that held
the lock with a genuine live holder, the test failed with
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` and the full `workspaceRestoreLock`
block at about 5 seconds, inside the 15-second test timeout. The
temporary edit was reverted.
- `workspace-restore-merge.test.ts` passed 56 of 56. The adapter test
files that cover every production caller passed 116 of 116 and 52 of 52.
`agent-directory-working-copies.test.ts` passed 70 of 70.
- The `adapter-utils` and `server` type-checks passed with no error.
- Confirmed that no spawned process survives the test run.

## Risks

Low risk. The production change is one optional parameter with the
existing default, so every production caller keeps the real 30-second
budget and no production call site changed. The remaining change is
limited to one test file. The `--import` form of the tsx hook is already
used elsewhere in this repository, in the container image command and in
an end-to-end test configuration. Test coverage does not drop: the owner
record is diagnostic only, the SQLite reserved lock remains the
authority that the tests exercise, and the timeout-asserting tests now
exercise the real retry loop instead of a replaced clock. The file costs
about 0.5 to 0.9 seconds more wall clock than `master`, which is the
cost of the short real waits that replace the instant forced timeout.

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`), used with extended thinking and
tool use for the diagnosis, the measurement, and the change.

## Checklist

Check every box that the state of the pull request satisfies. The local
test runs and the type checks are complete. Reconcile the
continuous-integration and review boxes after the checks reach their
terminal state.

## Test plan

- [x] Continuous integration is green on every check, including the
general test job.
- [x] The general test job passes the file
`packages/adapter-utils/src/directory-merge-lock.test.ts`.
- [x] Greptile returns 5 of 5 with no open item.
- [x] `mergeable: MERGEABLE` is terminal.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:17:03 -07:00
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00
DottaandPaperclip ec3bacc9bd fix(chat): hide ignored provider information (#14929)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task and agent chats show agent progress and problems that need
attention.
> - Codex also sends account, skill, and unrelated thread notifications.
> - The runner correctly ignores that information but reports it as a
warning.
> - Chat then shows an internal diagnostic as an actionable provider
notice.
> - This pull request keeps the diagnostic in run logs and removes it
from chat.
> - Real provider warnings, errors, and agent replies remain visible.

## Linked Issues or Issue Description

**What happened?**

Chat showed “Received a provider update” and a warning with the text
“ignored
unrelated provider information”. Its details said “User Actionable: Yes”
even
though no user action was needed. Saved conversations retained the same
noise.

**Expected behavior**

Keep ignored provider information in the run log. Do not show it as chat
activity
or a user warning. Preserve real warnings and errors.

**Steps to reproduce**

1. Start a conversation with the native Codex runner.
2. Have the provider send an account update, skill change, or unrelated
thread
   notification during the turn.
3. Inspect live chat and reload its saved history.

The regression tests also reproduce the old stored notice without a live
account.

**Paperclip version or commit**

Source implementation on master at `e00d10d5d`. The duplicate search
found no
open PR for this fix. Related prior work: #13109 improved
provider-notice
presentation. #12367 added Codex thread normalization. This change
addresses
the internal information that those paths still projected as chat
warnings.

**Deployment mode**

Native Paperclip Runner with the Codex app-server provider. The issue
was seen
in hosted chat and can be reproduced with local provider fixtures.

## What Changed

- Map ignored unrelated Codex information to `harness.diagnostic` in the
Rust
  and TypeScript normalizers.
- Retain a bounded allowlist of redacted provider method and thread/turn
identifiers.
- Use the same Unicode character limit and truncation marker in both
normalizers.
- Share the text redactor through a pure helper. Keep provider
connection code
  out of the standalone demo's source closure.
- Omit that diagnostic and the matching legacy notice from live chat.
- Omit the matching legacy notice from saved chat history.
- Test diagnostic retention, account-notification integration, live and
saved
  chat, and continued visibility of real warnings, errors, and replies.
- Document the local run-log event and historical display behavior.

## Verification

- Passed: 68 tests in the two affected UI transcript suites.
- Passed: 60 TypeScript tests across provider events, transport
behavior, and
  the standalone demo boundary.
- Passed: 13 Rust provider-event tests and the Codex
account-notification
  integration test.
- Passed: `pnpm check:token-gates` and Cargo formatting checks.
- Passed: full `pnpm build` and `pnpm -r typecheck`. After the review
fix,
the provider package build, typecheck, and both provider-event suites
passed again.
- Full local `pnpm test:run` failed: 608 files / 10,904 tests passed, 30
server
suites failed, and 104 files / 4,012 tests were skipped. Most failures
were
  embedded PostgreSQL startup errors. Two tests timed out in
`heartbeat-comment-wake-batching` and
`workspace-git-snapshot-streaming`.
  PostgreSQL startup also failed in `heartbeat-run-event-sequencing` and
`native-finalization-migration`. These server files are unchanged by
this PR.
Isolated heartbeat reruns were skipped locally. The stable test script
stopped
  after this general-server group, so later groups did not run locally.
- The original review thread is resolved. Greptile is 5/5 on current
head
  `683dab7cce57187c57e84c83f5e9da4ad75c9c04`.
- All current-head CI gates passed, including the full
server/chat/workspace
test matrix, Rust and TypeScript runner suites, browser E2E, build,
typecheck,
and release canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37021330663).
- Replay the exact old warning in either transcript adapter. It must
produce
no chat row. A genuine provider warning or error must still produce a
row.

## Risks

- Low risk. The display filter matches one diagnostic code or the
complete
  legacy warning shape. Other provider notices remain visible.
- New ignored-information events use the existing harness-diagnostic
event
type. They retain diagnostic evidence without original account payloads.
- No database migration, API permission, provider execution, or recovery
  behavior changes. This affects the local run log, not Telemetry or
  OpenTelemetry exports.

## Model Used

OpenAI Codex, GPT-6. The exact backend model ID and context-window size
are
not exposed in this session. Used reasoning, repository inspection, code
editing, tool use, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the affected tests locally and they pass (the broad
local run has PostgreSQL startup errors and timeouts documented above;
the full CI matrix passed)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 09:59:34 -05:00
DottaandPaperclip e00d10d5d5 fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path

> - Paperclip manages AI agents and controls the credentials used for
their work.
> - Managed AI connections resolve each responsible user's provider
default.
> - Agent settings created another account but kept the old default
selected.
> - A rejected provider test left the old account marked as connected.
> - Claude ACP reported a typed login failure as a generic
terminal-access error.
> - This pull request repairs the selected account or selects the new
login explicitly.
> - Agents can save and run with the repaired credential, and failed
logins request sign-in.

## Linked Issues or Issue Description

- Fixes #14831.
- Refs #13867. Environment failures remain separate from
credential-health failures.

## What Changed

- Add an agent-settings action to reconnect an unavailable personal
default in place. Keep its connection, grant, default, and agent access.
- State that a new account becomes the user's provider default. Select
its returned grant before changing the agent binding. Keep the actual
sign-in method.
- Show default-update errors and allow retry without another provider
login.
- Show the agent-access choice. Connection managers start with
company-wide access for their own tasks. Other members start with access
for the current agent.
- Use the server's connection-manager permission in the shared list
response. This includes members with a custom management grant.
- Mark credentials as needing attention after an explicit login
rejection in Test or Save. This includes API-key 401 and 403 responses.
Network, quota, and server failures keep the credential health
unchanged.
- Reuse the credential-generation check so an old failure cannot
invalidate a newer reconnect.
- Route Claude's typed provider `access` failure to the existing
login-recovery flow. Replace its generic terminal-access fallback with a
sign-in message.
- Add regression tests and update the AI Connections documentation.

## Verification

- Red: the UI tests failed on the missing reconnect action, unused
returned grant, missing access choice, and lost default-update error.
The server tests failed because rejected credentials stayed connected.
The real ACP fixture returned `acpx_turn_failed` for typed login
failures.
- Green: 156 tests passed across the AI connection, hiring, agent field,
and New Agent suites. All 37 environment-route tests passed. The Claude
ACP authentication fixtures also passed.
- `pnpm check:token-gates` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- The full local `pnpm test:run` passed 707 files and 14,503 tests, then
exited with an agent-conversation timeout and embedded PostgreSQL
startup failures in unchanged suites. The isolated conversation and
migration tests passed on rerun. Later local test groups did not run
after this failure.
- [All CI gates
passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356)
on commit `38513dfe2`. This includes the full test matrix, browser
tests, typecheck, build, Runner checks, and canary dry run.
- Greptile reviewed commit `38513dfe2` and returned 5/5 with no open
findings.
- The regression tests use a real embedded database and a real ACP
fixture process. Live provider sign-in requires a valid account and was
not run.

## Risks

- Connecting a new account from agent settings changes the user's
provider default. The dialog states this before sign-in.
- The displayed access choice can allow all company agents to use the
account for its owner's tasks. Reconnect keeps the existing access.
Server permissions still control installs.
- Claude's typed `access` category maps to the provider's
`auth_required` signal. Tool and workspace request failures retain their
existing classification.
- No database migration or provider credential format changes are
required.

## Model Used

- OpenAI GPT-6 through Codex. The exact served model identifier and
context window are not exposed in this session. Capabilities used:
reasoning, repository tools, code editing, and command execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:57:52 -05:00
DottaandPaperclip 408f70e69f fix(runner): preserve stock Codex base instructions (#14920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner connects Paperclip tasks to Codex app-server.
> - Paperclip passed its runtime context as `baseInstructions`.
> - That field replaces the stock Codex base prompt.
> - This pull request sends the same Paperclip context as additive
developer instructions.
> - Codex keeps its stock prompt and still receives Paperclip task
instructions and tools.

## Linked Issues or Issue Description

**What happened?**

The native Codex driver and Rust provider sent Paperclip context through
`baseInstructions` on thread start and resume. Codex used this text in
place of its stock base instructions. Direct-chat resume also sent an
empty replacement base. The Runner Lab session path used the same
replacement field.

**Expected behavior**

Codex should retain its stock base prompt. Paperclip should add its
runtime context through `developerInstructions`. Other provider facades
should retain their current instruction handling.

**Steps to reproduce**

1. Create a native Codex session through Paperclip Runner.
2. Inspect the `thread/start` request in the native provider trace.
3. Resume the session and inspect `thread/resume`.
4. Before this fix, these paths set `baseInstructions`. After this fix,
the Codex paths set `developerInstructions` and omit `baseInstructions`.

**Paperclip version or commit**

Reproduced against master at `cad26c6bfb736039c8ed5743da650a44792a083c`.

**Deployment mode**

Built from source. Native Codex app-server and runnerd paths. A local
protocol probe used codex-cli 0.153.4 and a localhost Responses stub.

No duplicate fix or matching public issue was found in the GitHub
search.

## What Changed

- Send additive developer instructions on Codex start and resume in the
TypeScript driver, Rust provider, and Runner Lab session path.
- Carry the additive fragment through runnerd, including runtime asset
path mapping.
- Preserve existing instruction fields for other provider facades,
including OpenCode.
- Add start/resume/direct-chat regression coverage and check the actual
Rust provider request.
- Document the historical option and trace field names. Record progress
and follow-ups in the working checklist.

## Verification

- `pnpm -r typecheck` — passed.
- `pnpm build` — passed.
- Targeted Codex driver lifecycle, driver, and live-session Vitest
suites — 139 tests passed.
- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml --locked -p
paperclip-runner-core --test codex_provider` — 91 passed, 2 ignored
subprocess helpers.
- Real app-server probe: a localhost Responses stub captured identical
14,732-character stock base instructions on fresh start and cold resume.
Both requests retained the Paperclip marker in developer input. Both
stub turns completed. No paid inference was used.
- Runnerd transport Vitest suite — 182 tests passed.
- The initial `pnpm test:run` attempt reported local dependency-loading,
embedded PostgreSQL startup, and macOS `/var` versus `/private/var` path
failures. It was stopped after those failures. Loading-suite reruns
passed 1,428 tests after the build; native interaction/finalization
reruns passed 38 tests. A seven-suite diagnostic rerun passed 463 tests
and isolated the remaining path and PostgreSQL setup failures.
- With `TMPDIR=/private/tmp`, workspace, gateway, interaction, and
attachment suites passed all 356 tests. The remaining environment-image
and native-session-resumption suites passed all 44 tests with the same
canonical temp path. All affected suites passed on rerun. The original
full local command was stopped after failures and is not claimed as
passing.
- All 55 PR checks passed at `83281439456181396f3707eecda5d2ebc90bd14d`.
Greptile scored 5/5 with no open review threads.
- No paid live campaign or Product E2E browser suite was run. This
change has protocol and regression coverage; it does not claim improved
task quality.

## Risks

- Stock Codex behavior may differ from behavior under the previous
Paperclip replacement prompt. Restoring that behavior is the intended
change.
- Existing Codex threads retain their saved replacement base prompt.
They need a provider session reset to receive the stock base. This PR
does not reset active sessions or alter recovery rules.
- The legacy `baseInstructions` option and trace field names remain for
compatibility. They now describe the additive Paperclip fragment for
Codex.
- The separate Codex-through-ACP dependency patch remains a follow-up in
the harness coverage checklist. This PR covers native app-server
execution.

## Model Used

OpenAI Codex, GPT-6. The exact runtime model variant and context window
are not exposed in this session. Used reasoning, repository inspection,
code editing, shell execution, and test tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:53:47 -05:00
cad26c6bfb fix(tool-gateway): bound MCP discovery memory and concurrency (#14864)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents discover governed tools through the MCP gateway.
> - A listing repeated policy and full run-row reads for each catalog
tool.
> - Parallel listings multiplied those allocations during run startup.
> - One 900-tool baseline listing used 11,489 queries and about 1.7 GiB
of extra heap in a fixture.
> - This pull request shares reads within a listing and bounds whole
listings across the process.
> - The benefit is lower discovery memory use while execution still
checks current policy.

## Linked Issues or Issue Description

Refs #13115. Its on-demand target change affects the same listing loop.

**What happened?**

MCP discovery repeated roughly 13 reads per tool. Full run snapshots and
repeated connection configurations caused large allocations. Per-listing
bounds alone did not limit concurrent listings across gateways.

**Expected behavior**

Discovery reads shared inputs once per listing. The process bounds
active and queued listings. Catalog payload size and policy evaluation
still grow with the catalog. Tool execution checks current access rules.

**Steps to reproduce**

Create a company with a remote MCP connection, 900 catalog tools, large
schemas, and a large run snapshot. Send concurrent tools/list requests
using a run-bound gateway token. Run the committed benchmark for a
deterministic reproduction.

**Paperclip version or commit**

Baseline: f2e0f19630. Measured on Node
26.4.0 on macOS.

## What Changed

- Keep Michael Nguyen's per-listing policy cache, scalar policy reads,
16-decision bound, and catalog/connection split under repeatable read.
- Admit two whole listings per process and queue at most 32. Return 503
with tool_discovery_busy when full.
- Propagate disconnects to discovery. Stop scheduling new reads and
drain started reads before freeing the slot.
- Project context IDs in gateway authentication and GitHub runtime
discovery. Avoid loading task descriptions and results.
- Keep discovery audit counts and a SHA-256 digest instead of the full
name list. Retain access and activity audit records.
- Sweep expired named gateway tokens at startup and on the existing
scheduler. Delete at most 500 per pass. Add an idempotent expiry-index
migration. Resolve audit token references atomically so cleanup cannot
break admitted requests.
- Return GET 405 with Allow: POST on stateless MCP gateway and
runtime-tools endpoints. Keep runtime-tools authentication.
- Add red-green regressions, HTTP concurrency and policy-revocation
coverage, and browser approval assertions.
- Commit the benchmark harness, raw results, and resource-bound
documentation.

## Verification

- Baseline listing regressions failed with 722 queries for 50 tools and
6,422 for 500. The connection-row duplication regression also failed.
- Follow-up regressions failed before the fixes for full snapshot reads,
abandoned listings, full name-list audits, and expired tokens.
- Listing and scheduler tests: 14 pass. Existing gateway/policy suites
passed after preserving the cleanup return contract.
- Two HTTP journeys pass: initialize, GET/SSE rejection, 16 concurrent
500-tool listings, provider call, policy revocation, and denied retry.
Token cleanup during provider dispatch also completes successfully and
blocks subsequent requests.
- pnpm test:e2e tests/e2e/mcp-user-stories.spec.ts --grep
'@mcp-runnable': 8 pass. The approval journey clicks Allow once in the
browser. Review screenshots wait for loaded content.
- pnpm -r typecheck: passes. pnpm build: passes.
- The general server group completed with 14,755 passes and 18 failures.
The 17 startup mock failures were fixed; all 21 startup tests pass on
rerun. The one Discord timing failure passed on the unchanged baseline
and on rerun (74 tests). UI and CLI groups pass 7,632 tests. Shared and
skill groups pass 853 tests. The remaining database and adapter groups
pass 3,010 tests with one worker after a macOS shared-memory limit
interrupted a parallel run. All 149 serialized route files pass (2,762
tests).
- At 900 tools, one listing falls from 11,489 to 36 queries and from
about 1.7 GiB to 37 MiB of extra heap. Sixteen concurrent listings used
169–187 MiB of extra heap. Four connections used 39 queries per listing.
- Latest-head verification: 54 successful checks and two expected skips
on 5c0793090c. Fresh Greptile review: 5/5
with no unresolved threads.
- Reproduce with server/scripts/benchmark-tool-gateway-listing.ts. See
doc/mcp-discovery-performance.md and
doc/benchmarks/2026-10-01-mcp-discovery.json.

## Risks

- The process-wide FIFO queue can increase discovery latency. Excess
callers must retry 503 responses.
- Cancellation applies to discovery. Started database reads finish
before their slot is released.
- Audit consumers must use visibleToolCount and visibleToolsHash instead
of visibleTools.
- The expiry index can briefly lock the token table during migration.
Sweeps preserve unexpired and non-expiring tokens.
- Measurements use isolated fixtures and deterministic providers. They
do not establish a production heap limit or affected installation count.
Rate-limit reads remain uncached.

## Model Used

- Original listing optimization: Anthropic Claude Opus 5.5
(claude-opus-5-5), Claude Code, extended thinking and tool use, as
recorded by the original author.
- Follow-up fixes and verification: OpenAI Codex, GPT-6, with shell
execution, file edits, database fixtures, and browser tests. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes / Closes /
Refs OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 07:45:27 -05:00
Devin FoleyandPaperclip 4e52463203 fix(daytona): recover output from stalled log streams (#14889)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Daytona driver streams sandbox command output to the host.
> - A log socket can stop delivering bytes without closing or rejecting.
> - Missing permission requests and cancellation results can leave a run
active and block queued work.
> - This pull request switches an idle log stream to saved-output
polling for the same command.
> - The host can receive the missing output without another command
dispatch.

## Linked Issues or Issue Description

**What happened?**

The driver waits for the SDK log-stream promise before it can start
recovery. If that promise never settles, the host can miss new output
that is already in the provider's saved logs. A run can remain active
after a watch completes. Interrupt can then time out with “Execution is
still stopping; termination has not been verified.”

**Expected behavior**

Recover command observation when the live stream stalls. Forward new
permission requests and cancellation results. Require a recorded command
exit before reporting completion. Keep quiet commands running under the
caller's existing lifetime controls.

**Steps to reproduce**

1. Start a session command and leave the callback log promise pending.
2. Put new output in the snapshot API without invoking the stream
callback.
3. Keep the command running until the host receives that output, then
expose its cancellation output and exit code.
4. Verify that the host receives each byte once and dispatches the
command once.

**Paperclip version or commit**

Base commit `479554120d`.

**Agent adapter(s) involved**

Daytona session commands, including sandbox ACP agent sessions.

Related: #14799 handles closed streams; this change handles sockets that
never close. #14485 handles input delivery retries. #13262 adds
permission diagnostics.

## What Changed

- After 15 seconds with no stdout or stderr, switch directly to the
existing status and log-snapshot polling path.
- Ignore callbacks and delayed failures from the abandoned stream. Clear
its idle timer on every exit path.
- Preserve byte-offset deduplication, one command dispatch, and
independent timeouts for recovery reads.
- Add regressions for an initial stream stall, a stall after UTF-8
output, cancellation output, late callbacks, hung recovery reads, and
quiet commands that outlive an operation timeout.
- Update the provider documentation and keep hour-long healthy-stream
coverage active with periodic output.

## Verification

- `pnpm vitest run
packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 245
passed.
- `pnpm exec vitest run --project @paperclipai/plugin-daytona`: 339
passed; 14 gated live tests skipped.
- Both new stalled-stream regressions fail on the unchanged base driver
because it never starts snapshot recovery. Both pass with this change.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: the local run did not pass. It was stopped after
confirmed local skill-path and macOS skill-cache failures, once complete
PR CI was green. Four chat/email tests could not load connector skill
files from an ancestor directory outside the checkout. Three
company-skills tests hit macOS `EACCES` during cache publication. One
unrelated wakeup test timed out in the full run and passed on a focused
rerun (`1 passed`, `27 skipped`). No source or test assertions were
changed for these failures. This is not a complete local-suite pass.
- Complete PR CI on `11ac4e030b`: 53 successful checks, 2 expected
skips, no pending or failed checks. The clean CI run includes the full
test suite.
- Greptile reviewed `11ac4e030b` at 5/5 with no findings or unresolved
threads. The branch has no merge conflicts with `master`.
- `git diff --check` and a local scan for secrets and private
identifiers passed.

## Risks

- Quiet healthy commands also switch to polling. Full snapshots can
increase bandwidth as output grows; polling remains limited to one
snapshot per second.
- The SDK exposes no stream cancellation handle. The old socket remains
owned by session teardown, and its callbacks cannot publish after
fallback.
- This change recovers a stalled output stream. It does not claim to
identify every cause of an unanswered permission request or change the
requirement to verify termination before releasing work.
- No schema migration or command replay.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run local change-specific tests; the full suite passes in
CI, with local-suite limitations documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 21:04:41 -07:00
dependabot[bot] c9cde69299 build(deps): bump @anthropic-ai/sdk from 0.121.0 to 0.129.0 (#13386)
Bumps
[@anthropic-ai/sdk](https://github.com/anthropics/anthropic-sdk-typescript)
from 0.121.0 to 0.129.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/anthropics/anthropic-sdk-typescript/releases">@​anthropic-ai/sdk's
releases</a>.</em></p>
<blockquote>
<h2>sdk: v0.129.0</h2>
<h2>0.129.0 (2026-09-28)</h2>
<p>Full Changelog: <a
href="https://github.com/anthropics/anthropic-sdk-typescript/compare/sdk-v0.128.0...sdk-v0.129.0">sdk-v0.128.0...sdk-v0.129.0</a></p>
<h3>Features</h3>
<ul>
<li><strong>api:</strong> add between_tools thinking type (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/ab51075609a2d583dffdca24856f4214779d9e7f">ab51075</a>)</li>
<li><strong>api:</strong> add claude-sonnet-5-5 (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/4e7736625fd23a5607800e8de4aa3c37daa74a07">4e77366</a>)</li>
<li><strong>api:</strong> add ClientToolUnion type for client-executed
tools (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/2c1d2d712071a5f10d2e70ea02dc33425b087799">2c1d2d7</a>)</li>
<li><strong>api:</strong> add include_inherited and source to workspace
rate limits (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/66e925e54ff3435a56f51f687641728b94aff306">66e925e</a>)</li>
<li><strong>api:</strong> add typed event type values to the Managed
Agents events list filter (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/8ac4a26809418453b64fee79f8ec790d32d1f92c">8ac4a26</a>)</li>
<li><strong>api:</strong> cache diagnostics GA — diagnostics on Message
/ MessageCreateParams (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/5ba5f29696d520f68a794936eff7337dce83ac05">5ba5f29</a>)</li>
<li><strong>tools:</strong> optionally start tool calls while the reply
streams (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/083969beadb360b9c4b329ee2541f67bd78d9caf">083969b</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li><strong>client:</strong> also send X-Stainless-Timeout for
client-level timeouts (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/b84783b6ee03eed519d622ea1e477cf8eb3863bf">b84783b</a>)</li>
<li><strong>client:</strong> send upload filenames as given, with no
placeholder (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/f9d3e980b1b498c1c452300a205811f6a56de69b">f9d3e98</a>)</li>
<li><strong>helpers:</strong> degrade between_tools thinking to disabled
on fallback hops (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/841">#841</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/edbbf05a62ffe2c95c13810f6d7a0a2796786e5f">edbbf05</a>)</li>
<li><strong>internal:</strong> let bundlers drop unused classes with
more than ten private-member assignments (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/67c7adbe5ebae805631e104b341d932a1554bda8">67c7adb</a>)</li>
<li><strong>streaming:</strong> show every complete array item and hold
back unfinished numbers in partial tool input (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/781">#781</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/58065889d6b112db6da5127484e0a24025e3fa37">5806588</a>)</li>
</ul>
<h3>Performance Improvements</h3>
<ul>
<li><strong>streaming:</strong> drop the redundant iterSSEChunks layer
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/0357f812dab273c0780c558606663a03e93e34df">0357f81</a>)</li>
<li><strong>streaming:</strong> take each string token as one slice in
the partial JSON tokenizer (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/255">#255</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/cdfb1e55afebe1c3be50401ca56bebf0ff009a4e">cdfb1e5</a>)</li>
</ul>
<h3>Chores</h3>
<ul>
<li><strong>api:</strong> deprecate the betas param on GA models and
completions methods (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/5c74e4532bb69f22afa3e08013645ac059440aef">5c74e45</a>)</li>
<li><strong>api:</strong> list the known model ids first in the Model
types (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/94528226386311c3ade468cbdea0326c06a1e53d">9452822</a>)</li>
<li><strong>ci:</strong> choose the CI runner by repository (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/33db1ec2e99a415ff98619cf88ddddbf2b7ca7b1">33db1ec</a>)</li>
<li><strong>docs:</strong> clarify that stream: true returns the raw
event stream (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/4286c227c3a605bf04d4f59ffb496f44c102a6ec">4286c22</a>)</li>
<li><strong>docs:</strong> make Managed Agents actor descriptions
resource-neutral (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/15733f98deaae4aee2fba26c1bfe4ee1514c7080">15733f9</a>)</li>
<li><strong>docs:</strong> restore the research-preview notice on the
Dream type (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/b32f8baa27c40ad5901531024bae86e0ca046bb5">b32f8ba</a>)</li>
<li><strong>internal:</strong> move old constants around (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/0cd8edfdc6dd91af4ffee4ed3cc7fc8cc22d65e3">0cd8edf</a>)</li>
<li><strong>tests:</strong> add diagnostics to the parser test's Message
fixtures (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/eb5ca58d51ed3309b1f317b20f7d248b036ba76a">eb5ca58</a>)</li>
<li><strong>tools:</strong> remove client-side compaction control (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/802">#802</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/9c3e8a5ffa38c35dec32f0258becd360babd5fb1">9c3e8a5</a>)</li>
</ul>
<h3>Documentation</h3>
<ul>
<li><strong>api:</strong> prefer each field's own description over its
shared type's (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/51934782720725acca570edeb0a3a7501ca7dbd8">5193478</a>)</li>
<li>expand CLAUDE.md into a full contributor guide (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/98d2ddbce7c1eabbabe5a130d18d307c237fb75c">98d2ddb</a>)</li>
</ul>
<h2>sdk: v0.128.0</h2>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/anthropics/anthropic-sdk-typescript/blob/main/CHANGELOG.md">@​anthropic-ai/sdk's
changelog</a>.</em></p>
<blockquote>
<h2>0.129.0 (2026-09-28)</h2>
<p>Full Changelog: <a
href="https://github.com/anthropics/anthropic-sdk-typescript/compare/sdk-v0.128.0...sdk-v0.129.0">sdk-v0.128.0...sdk-v0.129.0</a></p>
<h3>Features</h3>
<ul>
<li><strong>api:</strong> add between_tools thinking type (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/ab51075609a2d583dffdca24856f4214779d9e7f">ab51075</a>)</li>
<li><strong>api:</strong> add claude-sonnet-5-5 (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/4e7736625fd23a5607800e8de4aa3c37daa74a07">4e77366</a>)</li>
<li><strong>api:</strong> add ClientToolUnion type for client-executed
tools (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/2c1d2d712071a5f10d2e70ea02dc33425b087799">2c1d2d7</a>)</li>
<li><strong>api:</strong> add include_inherited and source to workspace
rate limits (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/66e925e54ff3435a56f51f687641728b94aff306">66e925e</a>)</li>
<li><strong>api:</strong> add typed event type values to the Managed
Agents events list filter (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/8ac4a26809418453b64fee79f8ec790d32d1f92c">8ac4a26</a>)</li>
<li><strong>api:</strong> cache diagnostics GA — diagnostics on Message
/ MessageCreateParams (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/5ba5f29696d520f68a794936eff7337dce83ac05">5ba5f29</a>)</li>
<li><strong>tools:</strong> optionally start tool calls while the reply
streams (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/083969beadb360b9c4b329ee2541f67bd78d9caf">083969b</a>)</li>
</ul>
<h3>Bug Fixes</h3>
<ul>
<li><strong>client:</strong> also send X-Stainless-Timeout for
client-level timeouts (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/b84783b6ee03eed519d622ea1e477cf8eb3863bf">b84783b</a>)</li>
<li><strong>client:</strong> send upload filenames as given, with no
placeholder (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/f9d3e980b1b498c1c452300a205811f6a56de69b">f9d3e98</a>)</li>
<li><strong>helpers:</strong> degrade between_tools thinking to disabled
on fallback hops (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/841">#841</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/edbbf05a62ffe2c95c13810f6d7a0a2796786e5f">edbbf05</a>)</li>
<li><strong>internal:</strong> let bundlers drop unused classes with
more than ten private-member assignments (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/67c7adbe5ebae805631e104b341d932a1554bda8">67c7adb</a>)</li>
<li><strong>streaming:</strong> show every complete array item and hold
back unfinished numbers in partial tool input (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/781">#781</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/58065889d6b112db6da5127484e0a24025e3fa37">5806588</a>)</li>
</ul>
<h3>Performance Improvements</h3>
<ul>
<li><strong>streaming:</strong> drop the redundant iterSSEChunks layer
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/0357f812dab273c0780c558606663a03e93e34df">0357f81</a>)</li>
<li><strong>streaming:</strong> take each string token as one slice in
the partial JSON tokenizer (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/255">#255</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/cdfb1e55afebe1c3be50401ca56bebf0ff009a4e">cdfb1e5</a>)</li>
</ul>
<h3>Chores</h3>
<ul>
<li><strong>api:</strong> deprecate the betas param on GA models and
completions methods (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/5c74e4532bb69f22afa3e08013645ac059440aef">5c74e45</a>)</li>
<li><strong>api:</strong> list the known model ids first in the Model
types (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/94528226386311c3ade468cbdea0326c06a1e53d">9452822</a>)</li>
<li><strong>ci:</strong> choose the CI runner by repository (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/33db1ec2e99a415ff98619cf88ddddbf2b7ca7b1">33db1ec</a>)</li>
<li><strong>docs:</strong> clarify that stream: true returns the raw
event stream (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/4286c227c3a605bf04d4f59ffb496f44c102a6ec">4286c22</a>)</li>
<li><strong>docs:</strong> make Managed Agents actor descriptions
resource-neutral (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/15733f98deaae4aee2fba26c1bfe4ee1514c7080">15733f9</a>)</li>
<li><strong>docs:</strong> restore the research-preview notice on the
Dream type (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/b32f8baa27c40ad5901531024bae86e0ca046bb5">b32f8ba</a>)</li>
<li><strong>internal:</strong> move old constants around (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/0cd8edfdc6dd91af4ffee4ed3cc7fc8cc22d65e3">0cd8edf</a>)</li>
<li><strong>tests:</strong> add diagnostics to the parser test's Message
fixtures (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/eb5ca58d51ed3309b1f317b20f7d248b036ba76a">eb5ca58</a>)</li>
<li><strong>tools:</strong> remove client-side compaction control (<a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/802">#802</a>)
(<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/9c3e8a5ffa38c35dec32f0258becd360babd5fb1">9c3e8a5</a>)</li>
</ul>
<h3>Documentation</h3>
<ul>
<li><strong>api:</strong> prefer each field's own description over its
shared type's (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/51934782720725acca570edeb0a3a7501ca7dbd8">5193478</a>)</li>
<li>expand CLAUDE.md into a full contributor guide (<a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/98d2ddbce7c1eabbabe5a130d18d307c237fb75c">98d2ddb</a>)</li>
</ul>
<h2>0.128.0 (2026-09-22)</h2>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/bf2058689f845dfb10e59bd9ebeb5cb4e9318a9d"><code>bf20586</code></a>
Merge pull request <a
href="https://redirect.github.com/anthropics/anthropic-sdk-typescript/issues/1218">#1218</a>
from anthropics/release-please--branches--main--chan...</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/da41a5a0e5d6f498f1bb8b71beb5b1e0a2bd1e48"><code>da41a5a</code></a>
chore: release main</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/3a79b93f0210a83f8bf9fda91a42b577c9e6b93c"><code>3a79b93</code></a>
codegen metadata</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/4e7736625fd23a5607800e8de4aa3c37daa74a07"><code>4e77366</code></a>
feat(api): add claude-sonnet-5-5</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/2c1d2d712071a5f10d2e70ea02dc33425b087799"><code>2c1d2d7</code></a>
feat(api): add ClientToolUnion type for client-executed tools</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/f9d3e980b1b498c1c452300a205811f6a56de69b"><code>f9d3e98</code></a>
fix(client): send upload filenames as given, with no placeholder</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/8ac4a26809418453b64fee79f8ec790d32d1f92c"><code>8ac4a26</code></a>
feat(api): add typed event type values to the Managed Agents events list
filter</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/083969beadb360b9c4b329ee2541f67bd78d9caf"><code>083969b</code></a>
feat(tools): optionally start tool calls while the reply streams</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/9505bf37def4a4ada41c95e4e4fa7fb6b3b3f679"><code>9505bf3</code></a>
codegen metadata</li>
<li><a
href="https://github.com/anthropics/anthropic-sdk-typescript/commit/b84783b6ee03eed519d622ea1e477cf8eb3863bf"><code>b84783b</code></a>
fix(client): also send X-Stainless-Timeout for client-level
timeouts</li>
<li>Additional commits viewable in <a
href="https://github.com/anthropics/anthropic-sdk-typescript/compare/sdk-v0.121.0...sdk-v0.129.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 15:30:19 -07:00
Devin FoleyandPaperclip f2e0f19630 Defer agent directory cleanup until stop proof is available (#14866)
## Thinking Path

> - Paperclip manages agents and their persistent files.
> - Each run owns a temporary agent directory and a save receipt.
> - Cleanup needs independent proof that the owning process stopped.
> - A cleanup call without that proof currently waits for the directory
lock anyway.
> - A second lock failure can prevent environment release after the run
already reported a failed save.
> - This change skips cleanup that has no authority and retries
unavailable remote copies after exact destruction proof.
> - The save failure stays visible. Existing lock owners remain
protected.

## Linked Issues or Issue Description

Related work: Refs #14787 (lock diagnostics), #14695 (warm instruction
ownership), #9667 (stale lock proposal), and #9872 (control-plane
ownership proposal). I checked open PRs and issues. This change leaves
the shared filesystem lock protocol in place and does not duplicate the
warm-retention work in #14695.

**What happened?**

Heartbeat cleanup records an explicit unavailable instruction-save
warning, then calls directory release before releasing the environment
lease. Release can wait for a lock even though the copy has no
process-stop proof and cannot be removed. That secondary timeout
prevents the following lease-release step. If destruction proof arrives
later, the unavailable copy is excluded from both recovery queries.

**Expected behavior**

Skip a release that cannot remove anything. Preserve the failed-save
receipt and candidate fields. Once exact remote destruction is recorded,
recover remote cleanup without running a provider command. Unavailable
local copies retain their potentially uncollected edits even if local
stop proof arrives later. A blocked cleanup must not prevent cleanup for
other agents.

**Steps to reproduce**

1. Prepare an agent directory, report its save unavailable, and leave
process-stop proof absent.
2. Hold the shared directory lock and call release. Before this change,
release waits and fails although removal is not authorized.
3. Record destruction of the copy's exact remote lease. Before this
change, neither recovery sweep selects the unavailable copy.

**Paperclip version or commit**

Reproduced against `efc2e6810e9bc0dc8cb412b0e7647c0db9821caa`.

**Deployment mode**

Local and remote execution with persistent agent directories. Tests use
an isolated embedded PostgreSQL database and fixture transports.

## What Changed

- Re-read receipts and skip release before lock acquisition when stop
proof is absent, the copy is superseded, or cleanup is complete. Keep
the same checks inside the lock.
- Recover unavailable remote copies only after exact destruction proof.
Preserve their unavailable state, errors, candidate hash, and candidate
bytes. Keep unavailable local copies and their uncollected edits
unchanged.
- Store destruction-only cleanup authority with the stop proof. Later
cleanup honors it after a lost database response or restart, including
when a transport remains cached.
- Defer failed or unproven cleanup with bounded batches and a retry
delay. Keep failed cleanup visible in logs and its receipt.
- Serialize preparation of an existing run with cleanup. Fresh run
preparation keeps its existing admission path.
- Cover held locks, receipt scope, delayed proof, batch fairness, lost
update responses, cached transports, and concurrent same-run preparation
with database regressions.

## Verification

- Focused directory, legacy instruction-copy, shared lock, and bounded
diagnostic suites: 169 tests passed across four files.
- `pnpm -r typecheck`: passed on the final source.
- `pnpm build`: passed on the final source.
- Completed all selected local `pnpm test:run` groups: 733 general
server suites, 149 serialized suites, and 14 workspace projects. There
are 13 known macOS `EACCES` failures in the unchanged runtime skill
cache tests. Their exact signatures match earlier clean-base results,
and the cache source and test blobs match both that base and this PR
base (existing fix: #14290). One CLI import test timed out under
concurrent load; its full file passed separately (17 tests). Broad
coverage began before the review corrections; the final source has the
focused 169-test run, typecheck, and build. This is a local verification
limit, not a passing full local suite.
- `git diff --check` and local Gitleaks plus private-identifier/PII diff
scans passed.
- Independent review of the final source found no remaining actionable
issue. Its 17 targeted tests cover crash recovery, cached transports,
same-run preparation, real local edit preservation, proof scope, and
batch fairness. The main focused run also covers contained scheduling
failures.
- Final commit `35a24085f7`: Greptile 5/5 with no recommendations and
zero unresolved review threads.
- Final commit `35a24085f7`: all 53 checks passed, including Canary Dry
Run and the security scan; two visual checks were intentionally skipped.
The workspace shard passed on retry after GitHub reported that its first
runner lost communication. An earlier Canary runner shut down after the
release dry run passed. Neither interruption recorded an application
assertion failure; the exact final-head checks are now green.

## Risks

- This repairs cleanup ordering and recovery eligibility. It does not
repair an ambiguous legacy lock owner or restore unsaved files. Actual
collection still fails visibly when its lock cannot be acquired.
- An unavailable remote copy is recovered only after exact destruction
proof. A stopped but retained environment stays protected; recovery does
not execute a command that could restart it.
- Unavailable local copies with later stop proof still retain
potentially uncollected edits. A general local recollection or
reclamation policy remains outside this change.
- Existing-run preparation now waits for the same lock as cleanup. The
fresh-run path is unchanged.
- The cleanup mode is stored in the existing private receipt JSON. No
schema migration or public API change is required.
- No deployment, task replay, or runtime lock deletion was performed.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and test
execution. The runtime does not expose a more specific model suffix or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 14:28:47 -07:00
dependabot[bot] 0f9e9be408 build(deps): bump @codemirror/state from 6.7.2 to 6.7.6 (#13391)
Bumps [@codemirror/state](https://github.com/codemirror/state) from
6.7.2 to 6.7.6.
<details>
<summary>Commits</summary>
<ul>
<li>See full diff in <a
href="https://github.com/codemirror/state/commits">compare view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 13:45:48 -07:00
dependabot[bot] e2d4f07207 build(deps): bump googleapis from 176.0.0 to 182.0.0 (#13393)
Bumps
[googleapis](https://github.com/googleapis/google-api-nodejs-client)
from 176.0.0 to 182.0.0.
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/9eb464beb4310053dec5aef4661d6c6fd73051a8"><code>9eb464b</code></a>
chore: release main (<a
href="https://redirect.github.com/googleapis/google-api-nodejs-client/issues/4023">#4023</a>)</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/9e3b4655c95fde3c69d2a44496a586e906695734"><code>9e3b465</code></a>
feat: regenerate index files</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/279af87fcf2c4016321bde793a8cefd1a542615d"><code>279af87</code></a>
fix(youtubereporting): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/b37c22595087460ffc1be63471bb56e3e81675f5"><code>b37c225</code></a>
fix(youtubeAnalytics): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/512e921970a8556fc20aef5e7766826107018957"><code>512e921</code></a>
fix(youtube): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/6cd997a07c55c80a3f3070cea06affd0aa3274c9"><code>6cd997a</code></a>
fix(workstations): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/62d5f7629a9ddd6523b4f4050c955a2063007fce"><code>62d5f76</code></a>
fix(workspaceevents): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/0d71cd6b01ed41b1dff0a04d85567cfd2a511495"><code>0d71cd6</code></a>
feat(workloadmanager): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/2c46e13798ba24a8b4d2ad9e92f7995d7dcdd93e"><code>2c46e13</code></a>
fix(workflows): update the API</li>
<li><a
href="https://github.com/googleapis/google-api-nodejs-client/commit/c7b8641390370d9af9b12154e6bf455bc344feaa"><code>c7b8641</code></a>
fix(workflowexecutions): update the API</li>
<li>Additional commits viewable in <a
href="https://github.com/googleapis/google-api-nodejs-client/compare/googleapis-v176.0.0...googleapis-v182.0.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-01 13:16:41 -07:00
Michael NguyenandClaude Opus 5.5 b721d24cac fix(adapter-utils): retry GitHub broker transport failures before falling back (#14856)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run `git` and `gh` through a managed launcher. The launcher
gets a GitHub credential from the Paperclip control plane
> - The launcher sends one request to the credential broker for each
command
> - If that request fails at the transport level, for example after a
10-second timeout, the launcher continues without managed credentials
> - So a slow or restarting control plane removes the managed GitHub
identity from that command. Some agents then use other GitHub identities
that do not have the necessary permissions
> - This pull request retries a failed broker request two more times,
with a short backoff, before the launcher gives up
> - The benefit is that a short control-plane delay does not remove the
managed identity from an agent's GitHub operation

## Linked Issues or Issue Description

Refs #14175. That pull request changes the same broker request loop for
a different failure: sandbox network denials. The pull request that
merges second must rebase.

**What happened**
A Codex agent ran `git` and `gh` through the managed launcher while the
control plane was under heavy memory pressure. Each command printed
`Paperclip: GitHub broker_transport_unavailable; continuing without
managed credentials.` The agent then tried to open the pull request
through a different GitHub integration. GitHub rejected the request with
`403 Resource not accessible by integration`.

**Expected behavior**
A short broker delay or a short transport failure must not remove the
managed GitHub identity from the command. The launcher must try the
broker again before it continues without credentials.

**Steps to reproduce**
1. Set `PAPERCLIP_GITHUB_BROKER_URL` to a closed port.
2. Start a broker on that port after about 300 ms.
3. Run `gh` through the launcher.
4. Before this change, the launcher prints
`broker_transport_unavailable` and runs `gh` without the managed token.

**Version or commit**
`4ac374103` on master. Commit `3166e93a7` has the same code.

**Deployment mode**
Local trusted instance that runs as a launchd service, with
`codex_local` agents.

## What Changed

- `packages/adapter-utils/src/github-launcher.ts`: the broker request
loop now catches transport errors and retries up to two more times,
after 0.5 s and then after 1 s. The loop reads the response body inside
the retry, so a failed or slow body read is also retried. Busy (409)
responses keep their own budget of 30 attempts, separate from transport
retries. After the third transport failure, the launcher prints
`broker_transport_unavailable` as before.
- `packages/adapter-utils/src/github-launcher.test.ts`: two new tests
make the broker fail the first request and answer the second. In one,
the connection drops before the response. In the other, the connection
drops in the middle of the body. Each test checks that `gh` gets the
managed token, that the broker receives exactly two requests, and that
no `broker_transport_unavailable` message appears.
- The existing `broker-offline` test now has a 15-second timeout,
because each command now retries twice before it falls back.

## Verification

- `npx vitest run packages/adapter-utils/src/github-launcher.test.ts`: 9
of 9 tests pass.
- The body-read test fails on the first commit of this pull request and
passes with the second commit.
- `pnpm --filter @paperclipai/adapter-utils typecheck`: passes.
- The existing `broker-offline` test confirms that the launcher still
falls back after the retries, and that local Git still works.

## Risks

- When the broker is unreachable, each `git` or `gh` command now waits
about 1.5 s more before it continues without credentials. When the
broker times out, the worst case is about 31.5 s instead of 10 s.
- The change only adds retries. It does not change which credentials the
launcher accepts or which environment variables it copies.
- #14175 changes the same loop. The pull request that merges second
needs a small rebase.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Anthropic Claude Opus 5.5 (`claude-opus-5-5`), used through Claude
Code with tool use: shell commands, file edits and test runs. The model
wrote the change, the test and this description. The repository owner
approved the change before it was made.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — *targeted tests and the
package typecheck; see Verification*
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — *no
documentation describes the broker retry*
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — *CI has not run yet*
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
*Greptile has not reviewed yet*
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:57:41 -07:00
Devin FoleyandPaperclip dd9983b894 fix(adapter-utils): release restore locks when a process crashes (#14869)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent runs restore workspace files and collect instruction-file
changes.
> - Writers to the same target directory must wait for each other.
> - The current lock records a PID, which a new process can reuse after
a crash.
> - A reused PID can keep an orphaned lock alive and make each later run
fail.
> - This pull request makes a SQLite file lock decide ownership. The OS
releases it when the process exits.
> - Later runs can proceed after a crash, and concurrent live writers
remain protected.

## Linked Issues or Issue Description

Refs #10914. This addresses crash recovery. It does not cancel a stalled
operation in a process that is still alive.

Related work: #9667, #14787, and #12187. The earlier attempt in #9667
assumes one live server per lock root. This implementation uses an
OS-backed lock to support concurrent writers without treating a
different process token or an old timestamp as proof of a dead owner. It
retains the private lock root and bounded timeout diagnostics from the
merged changes.

After a process dies while holding a restore lock, a replacement process
can reuse its PID. The existing `process.kill(pid, 0)` check then
reports a live owner forever. Later runs can complete their model turn
but fail during file collection or restore.

## What Changed

- Hold a SQLite `BEGIN IMMEDIATE` transaction for each directory write.
Use the existing built-in `node:sqlite` dependency.
- Keep each lock database on a stable inode. Keep PID and time metadata
only for diagnostics.
- Retain the 30-second asynchronous wait and existing timeout error code
and diagnostic fields.
- Fail closed when an old directory lock exists. Document a
stopped-writer upgrade and rollback procedure.
- Add real child-process tests for crashes, PID reuse, live owners, and
connection cleanup. Cover callback failures, independent targets, stable
inodes, invalid lock files, and ambiguous legacy records.

## Verification

- Before the fix, the crash/PID-reuse test and the live-owner test both
failed. Both pass with this change.
- `pnpm exec vitest run
packages/adapter-utils/src/directory-merge-lock.test.ts
packages/adapter-utils/src/workspace-restore-merge.test.ts`: 56 tests
passed.
- Restore and agent-file working-copy integration tests: 118 tests
passed before the additional connection-cleanup test.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Full GitHub CI: all checks passed, including Linux workspace tests,
server test shards, build, typecheck, and browser tests.
- Greptile: 5/5, with no review threads or unresolved comments.
- `pnpm test:run`: started locally, then stopped with SIGINT (exit 130)
after full CI passed. The local serial run did not complete and is not
counted as a full local pass. The completed CI shards provide the
full-suite result.

## Risks

- **Upgrade and rollback require a drain.** Stop every old writer that
shares an instance root before switching protocols. Old and new versions
must not write concurrently.
- Existing legacy `.lock/` directories remain blocking. After all
writers stop, preserve run evidence and move those directories to an
operator scratch directory. The new code does not infer that they are
abandoned from PID or age.
- Never delete or replace a `.lock.sqlite` file while writers can run.
These small files remain after release.
- The shared filesystem must support reliable SQLite locking. Broken
network-filesystem locking is unsupported.
- This change prevents new orphaned ownership. It does not recover file
changes lost during earlier failed collections, or interrupt a live
operation that stalls.
- No application database migration or new native dependency is
required. See `doc/workspace-restore-locks.md` for the procedure.

## Model Used

OpenAI Codex based on GPT-6, with code execution and repository tools.
The exact model variant and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused regression and
integration suites; see the full-suite note above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 12:14:48 -07:00
DottaandPaperclip 4ac374103f fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.

## Linked Issues or Issue Description

Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.

**What happened?**

Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.

**Expected behavior**

Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.

**Steps to reproduce**

1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.

**Paperclip version or commit**

Reproduced from b54b2dc35c. Rebased onto
master at `0829d94af` after the single-screen setup change in #14811.

**Deployment mode**

Local development from source. The managed path also supports enrolled
self-hosted instances.

## What Changed

- Add the `asana.mcp` managed profile and the default Sign in with Asana
method.
- Use Asana's reviewed v2 protected-resource metadata before cached
endpoints.
- Require a custom MCP app's client secret and retain saved credentials
during setup or reconnect. Repair only the known Asana v1
issuer/resource binding, retaining company and callback checks.
- Expose a boolean for the acting user's saved client secret. The form
offers secret reuse only when that user has an active grant with the
required reference.
- Let users select their own app from the enrollment and
shared-app-unavailable screens, or from Advanced on the single-screen
setup page.
- Canonicalize HTTP loopback callbacks even when the request includes an
Origin header.
- Extend signed broker claims and provider URL validation for Asana.
Require refresh credentials on managed authorization.
- Document setup, distribution, and shared-app rollout requirements.
- Resolve permission-profile name collisions when finishing another
account. The live staging test found this after renaming the first Asana
connection; OAuth succeeded but profile finalization failed.

## Verification

- Final live staging proof used app commit
`c6053157c4e42ac017727117b754ae77fa5c45fa` and the real Cloud broker at
`767b63835170f664542afd0df99a76615e204b62`. In the embedded browser,
default shared sign-in required no client credentials, returned through
the central Cloud callback to the tenant, and discovered 39 actions. Get
me succeeded through the gateway as the selected QA agent (2.1 seconds).
The custom-app connection also returned a real result on this final
build (0.9 seconds).
- Retried the shared draft that failed during the first staging test. It
completed after the profile-name fix, retained the selected agent, and
kept the existing custom connection intact. Two database regressions
reproduced the collision before the fix and passed afterward. The
updated transaction rollback test also passes.
- Shared reconnect returned to the same staging connection with 39
actions. Earlier staging checks verified the custom-app fallback when
the shared profile was unavailable, saved-secret reuse on reconnect, and
Off blocking the action test. Allowed was restored after that check.
- Local live-provider checks also repaired a saved Asana v1
issuer/resource binding without reentering the secret and verified that
numeric-loopback setup uses the displayed localhost callback. Expiring
the local managed access-token timestamp triggered a real Asana refresh
and a successful Get me call. These early local broker tests used
enrollment/authentication and storage fixtures; the final staging proof
used deployed Cloud identity and persistent storage.
- The new production app is registered and configured, but production
sign-in has not been deployed or verified. Live provider revocation was
not run because the existing staging test app is shared with other
connections.

- After rebasing onto the single-screen setup flow, full `pnpm -r
typecheck`, `pnpm build`, and `pnpm check:token-gates` pass. Focused
verification passes 365 service and 36 broker-client tests. Broader
checks pass all 833 shared-package tests and all 530 connector-page
tests. The shared suite uses `TMPDIR=/private/tmp` to avoid macOS
temporary-directory symlinks in its canonical-path tests. The UI tests
verify the shared-app default and switching to a custom app with its
required client secret.
- Embedded-browser smoke on the current rebased build verified the
shared sign-in default, Advanced → custom app (client ID and secret
required), and switching back to Paperclip. Both existing Asana
connections remained connected after restart. No new provider
authorization was performed during this smoke.
- The full local `pnpm test:run` was interrupted when the execution
session restarted. Before interruption, it reported one runtime-slot
restart test failure. That test passed on an isolated retry after
clearing two unused PostgreSQL shared-memory segments. The full local
suite did not complete; CI must pass on the current head before merge.
- All 52 CI and security checks pass on
`9318fd4e9b616cdc3de12f40cdb9bd32d865af4c` (CI run `36882064080`),
including all eight browser shards, nine serialized-server shards,
build, typecheck, and canary dry run. Two optional Storybook checks were
skipped. Greptile review 4 reports 5/5 on this exact commit, with all
review threads resolved. Its updated summary identifies the current SHA;
this comment-triggered review did not publish a separate GitHub check
run.
- Provider revocation is unit-tested in the companion broker. Live
provider revocation was not run because the existing test app is shared
with other connections.

## Risks

- The shared Paperclip Asana MCP app has been registered with its
production callback and Any workspace distribution. Its secret is
provisioned in the production secret store, and the runtime client ID
and secret reference are configured. The production profile is enabled
in the saved deployment configuration. The companion broker has merged
and passed staging deployment; production sign-in still requires a
production deployment and live verification. Custom setup remains
available.
- Asana MCP uses the provider's fixed `default` grant. Paperclip action
policies limit agent tool use; they do not narrow provider consent.
- The callback correction affects HTTP loopback OAuth flows. Public
HTTPS callbacks retain their existing behavior.
- Reviewed discovery URLs now override stale cached endpoints. Tests
cover the Asana v1-to-v2 repair.
- No schema migration. Connection removal retains the existing
local-only revocation behavior.

## Model Used

OpenAI GPT-6 through Codex, with code execution, browser testing, and
GitHub tooling. The exact model variant and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:24:44 -05:00
467125fafb feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen

## Linked Issues or Issue Description

No public issue exists. This is the description, from the enhancement
template.

**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.

**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.

**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.

**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.

**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.

**Breaking changes**
None. No schema or API change. Existing connections keep their settings.

## What Changed

- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.

## Verification

- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
  - `action-permissions.test.ts` checks the group update.
  - `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".

## Risks

- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:13:47 -07:00
Devin FoleyandPaperclip 0d3e7bf6ac fix(daytona): keep commands alive after log stream closure (#14799)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox providers run agent processes and send their output to the
host.
> - The Daytona SDK can close a log socket while the remote command
still runs.
> - The driver treated a clean socket close as completion before it had
a command exit code.
> - This pull request recovers log observation for the same command and
waits for a recorded exit.
> - The host keeps receiving new output without a second command
dispatch.

## Linked Issues or Issue Description

**What happened?**

A clean close of the Daytona session log WebSocket resolves the SDK
callback promise. If the command still runs, the driver returned
`exitCode: null` with `timedOut: false` after a short status check. A
streamed ACP bridge can then report a process disconnect.

**Expected behavior**

A log socket close must not complete a running command. Recovery must
preserve new output, the caller's lifetime controls, and one command
dispatch.

**Steps to reproduce**

Use the callback form of `getSessionCommandLogs`. Let that promise
resolve while `getSessionCommand` still has no exit code. Keep the
command running, then expose its final logs and exit code. The new
regressions exercise this sequence, including streams that run longer
than the provider operation timeout.

Related: #11021, #11049. #14485 covers input delivery retries, which are
a separate transport path.

## What Changed

- Require a recorded command exit after a clean log-stream close.
Reconnect once, then read status and full log snapshots at most once per
second.
- Forward new output during recovery. Remove replayed prefixes and
reconcile a final snapshot so bytes written after socket close are
retained.
- Hold a trailing UTF-8 replacement suffix until replay or completion
resolves it. This handles the SDK decoder flush when a socket closes in
the middle of a character.
- Preserve healthy initial and reconnected stream lifetimes. Bound each
recovery read. Preserve the existing fallback timeout budget after
rejected stream attempts.
- Retain partial output on timeout. A log-observation timeout reports an
unconfirmed result and preserves any observed exit code in metadata; it
does not synthesize successful completion.

## Verification

- `pnpm exec vitest run --config
packages/plugins/sandbox-providers/daytona/vitest.config.ts`: 335
passed, 14 gated tests skipped.
- `pnpm exec vitest run --project @paperclipai/plugin-daytona`: 335
passed, 14 gated tests skipped.
- `pnpm exec tsc --noEmit -p
packages/plugins/sandbox-providers/daytona/tsconfig.json`: passed.
- `pnpm exec tsc -p
packages/plugins/sandbox-providers/daytona/tsconfig.json`: passed.
- `pnpm --workspace-concurrency=1 -r typecheck`: passed.
- `CARGO_BUILD_JOBS=2 pnpm --workspace-concurrency=1 -r build`: passed.
- `pnpm test:run`: exited with failure after 845.77 seconds. The server
phase reported 43 failed files, 529 passed, and 163 skipped; 14 failed
tests, 9,122 passed, and 5,599 skipped. All failure entries were traced
to local PostgreSQL startup/cleanup errors or ten macOS skill-cache
rename errors. The package helper restored 17 missing PostgreSQL library
links; the four sequencing/migration tests and fourteen
native-workspace-finalizer tests then passed. The ten cache failures
match clean-base evidence with identical source/test blobs. This is not
a full-suite pass; later local test groups did not run. PR CI supplies
the complete check result.
- Regressions cover clean and rejected stream recovery, two hour-long
streams with a five-minute operation budget, live fallback output, one
dispatch, delayed final output, stale snapshots, SDK UTF-8 decoding,
bounded observation, and timer cleanup.
- `git diff --check` and a local diff scan for secrets and private
identifiers passed.
- Greptile reviewed head `293c4dbe67` at 5/5. Both prior findings are
fixed, both threads are resolved, and no new actionable findings remain.

## Risks

- The SDK returns full snapshots with no offset API. After both stream
attempts end, polling bandwidth grows with retained output. The
one-second cadence limits request frequency.
- The SDK callback stream has no cancellation handle. Existing caller
stop logic and provider/session teardown still own its lifetime. Late
callbacks from a settled stream are ignored.
- A successful status read with no exit code keeps recovery active under
the existing caller guard. A failed read stops recovery; it is not
retried indefinitely.
- No command replay, provider API change, schema migration, or runtime
timeout policy change is included.

## Model Used

OpenAI GPT-6 through Codex, with tool use and independent code review.
The exact serving model identifier is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:37:54 -07:00
Devin FoleyandPaperclip 4b9a6000f7 Add bounded evidence for directory lock timeouts (#14787)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent files use directory locks during collection and cleanup.
> - A lock timeout can fail finalization after the model turn completes.
> - The timeout currently identifies no owner state or waiting
operation.
> - This pull request adds bounded evidence to the existing run failure
report.
> - Operators can distinguish a known local holder from a possible old
lock without changing lock safety.

## Linked Issues or Issue Description

**What happened?**

A directory lock timeout does not distinguish active local work from an
owner record left by an earlier process. The stored execution stage can
also precede the cleanup operation that failed.

**Expected behavior**

The failure report should identify the waiting operation and expose
bounded ownership clues. It must preserve the timeout and keep unknown
ownership protected.

**Steps to reproduce**

Hold a directory merge lock while a second caller reaches its
acquisition deadline. The regression tests exercise a live holder and an
older owner record with a live PID.

Related: #9667 proposes stale-lock recovery under a single-server
assumption. This change only adds evidence and does not adopt that
assumption. #14575 and #14665 add other run failure diagnostics.

## What Changed

- Record lock owner state, capped age and wait duration, same-process
and process-age comparisons, and whether this module holds the lock.
- Label agent-directory release, collection, checkpoint, and warm
handoff timeouts with a fixed operation code.
- Validate each field before the existing event-local Sentry report
accepts it. Exclude owner records, PIDs, paths, and absolute timestamps.
- Limit the extra diagnostic owner read to 100 ms with best-effort
abort; malformed JSON is `invalid` and unreadable owner records remain
`unknown`.
- Document the diagnostic limits and verify that contenders never
reclaim protected locks.

## Verification

- Focused lock, diagnostic, real Sentry SDK, and database-backed
agent-directory tests: 126 passed, including stalled-read and
malformed/missing/unreadable-owner regression coverage.
- Final revision `0691613dcc`: all 54 reported checks successful, with
two intentionally skipped Storybook checks. Greptile: 5/5, zero
unresolved review threads; no merge conflicts.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: complete suite coverage ran with the existing
repository shard flags: four general-server shards, four serialized
shards, two general-workspaces-a shards, and general-workspaces-b. The
full run is not green because of the base failures below.
- The broad run found 13 failures in the unchanged macOS skill-cache
tests. All 13 reproduce on the clean base revision. Open PR #14290
covers that existing failure.
- Two unchanged CLI archive tests hit their five-second limits during
the broad run; all 17 tests in that file pass on recheck. A CLI auth
socket error also cleared on recheck (19 tests), and its full serialized
shard passed on rerun.

## Risks

This is a diagnostic change, not a stale-lock fix. Owner observations
can race with release. Wall-clock shifts can affect the age comparison.
A local-holder flag covers only this module instance. None of these
fields authorizes reclamation or proves a file save. Lock acquisition,
release, retries, task status, and recovery guards retain their current
behavior. No schema change or deployment action is required.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact model build and context window were not exposed to this agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused checks pass;
existing base failures are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:35:12 -07:00
Devin FoleyandPaperclip 98d8a6ccac Stop replaying ambiguous database disconnects (#14773)
## Thinking Path

> - Paperclip stores agent work and control state in PostgreSQL.
> - Its database client must not repeat a mutation after an uncertain
result.
> - The global retry wrapper treated `write CONNECTION_CLOSED` as proof
that PostgreSQL never received a statement.
> - postgres.js also uses that message when the connection closes after
statement delivery.
> - This pull request removes that global replay and tests the actual
driver over a local wire connection.
> - Callers retain control of retries when they can prove the complete
operation is idempotent.

## Linked Issues or Issue Description

Follow-up to #13417. Preserve the transaction disconnect handling from
#13643 and the explicit actor synchronization retries introduced in
#12773. Searched open and closed issues and PRs for database retries,
disconnects, and `CONNECTION_CLOSED`. The open circuit-breaker proposal
#11142 addresses outage queue growth; it does not establish whether an
already-sent statement can be replayed.

**What happened?**
The database wrapper replayed an arbitrary statement up to three times
after `write CONNECTION_CLOSED`. The driver adds `write ` to
connection-close errors even after the peer receives the statement. A
local protocol peer receives the same submitted INSERT three times when
it drops each response. A committed write could therefore execute more
than once.

**Expected behavior**
An ambiguous statement result must fail without automatic replay. A
subsequent operation must be able to reconnect.

**Steps to reproduce**
Run the new wire regression against the parent commit. The six Simple
Query cases and the parameterized Drizzle case receive three executions
instead of one. The named prepared-client case was already safe and
stays covered. The peer reads the entire statement and then closes the
connection. This demonstrates repeated delivery with the real driver; it
does not claim that a historical incident duplicated a committed write.

**Paperclip version or commit**
Reproduced on source commit `018993140f` with the patched postgres.js
3.4.9 dependency.

**Deployment mode**
Built from source with a local PostgreSQL protocol peer. No live
provider or customer database is used.

## What Changed

- Pass the original postgres.js client to Drizzle and remove the global
statement replay wrapper.
- Add eight wire regressions: six Simple Query cases for INSERT,
side-effect-capable SELECT, and a data-changing CTE, plus parameterized
Drizzle and named prepared-client cases. The extended peer processes
Parse, Describe, Bind, and Execute, verifies bound parameters, and drops
the response only after Execute. Each case checks one delivery and
recovery on a fresh query.
- Document ambiguous outcomes and the retry compatibility tradeoff. Keep
explicit idempotent actor-sync retries and disconnected-transaction
handling unchanged.

## Verification

- Before the fix: the six Simple Query cases and the parameterized
Drizzle case failed with three executions instead of one. The named
prepared-client case was already safe. All eight wire cases pass on this
branch.
- Final focused client, pool teardown, configuration, and actor-sync
retry checks: 28 tests passed. `pnpm --filter @paperclipai/db typecheck`
also passed after the test-only follow-up.
- First implementation head, `pnpm exec vitest run --project
@paperclipai/db`: all 158 tests passed across 45 files, including real
PostgreSQL transaction/reserved-connection recovery. The local embedded
dependency's symlinks were hydrated before this run.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- The complete local `pnpm test:run` did not finish; no complete local
suite pass is claimed. All CI test, typecheck, and build gates passed on
the first implementation head `11f8b22d90`. Final-head CI is pending
after the test-only follow-up.
- `git diff --check`: passed. Reviewed the diff for secrets, personal
data, generated output, and run artifacts.

## Risks

Some transient statement failures that the global wrapper previously
replayed now reach the caller. Operation owners must retry only when
they have an idempotency guarantee or a durable receipt that prevents
duplicate effects. A connection error is not proof that a write failed
to commit. There is no SQL-text retry heuristic, new suppression, schema
change, or migration. This change prevents unsafe replay; it does not
prevent network disconnects.

## Model Used

OpenAI Codex / GPT-6, with reasoning, repository inspection, code
execution, and local protocol tests. The exact backend model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


Final verification (September30): every final-head CI check passed at
`3004c5bda39c985c3557547ec45e33870ce5d010`. Greptile scored5/5 on this
head, all review threads are resolved, and the branch is mergeable.
Eight real-wire regressions cover simple, parameterized Drizzle, and
named prepared queries. Full local suite did not produce a completed
result; the complete CI matrix passed. This public PR remains open for
maintainer merge.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:34:00 -07:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
Devin FoleyandPaperclip d6d67b00d3 Prevent background workspace scans from refreshing the Git index (#14666)
## Thinking Path

Paperclip runs background workspace scans alongside real Git writers.
`git status` can refresh the index as an optional side effect, taking a
lock that makes another operation fail. Disable optional locking in the
shared scan subprocess so background observation does not compete with
workspace updates.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Workspace Git scans used by changed-file browsing, cleanliness guards,
and sandbox snapshots.

**Current behavior**
The scan process inherits Git's default optional-lock behavior. Even a
clean `status` can rewrite stale stat-cache entries in the index and
contend with a concurrent writer.

**Proposed behavior**
Always set `GIT_OPTIONAL_LOCKS=0` for the shared scan subprocess while
preserving the selected environment and Git's required write locks.

**Reason and benefit**
Background reads stop creating avoidable index contention. Git documents
this behavior and recommends disabling optional locks for background
status: [background
refresh](https://git-scm.com/docs/git-status#_background_refresh).

**Breaking changes**
None to scan results or required write locking. Later scans may repeat
stat checks that would otherwise have been cached in the index.

Related scan implementation: #11572, #14253. This avoids one known
contention source; it does not identify every historical lock owner or
repair abandoned locks.

## What Changed

- Disable optional locking at the shared scan subprocess boundary,
including explicit caller environments.
- Test a clean status against a deliberately stale index and prove an
ordinary status would rewrite it.
- Test tracked/untracked results with an existing index lock,
preservation of that lock and working files, and continued rejection of
a mandatory-lock write.
- Document the scan behavior and performance tradeoff.

## Verification

- Focused stream, workspace-sync, and scheduler suites: 63 tests passed.
- Full `pnpm -r typecheck` and `pnpm build` passed locally. Final head
`effe6420c77bd18d36af6b093db3a564c04b8b38` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially had missing embedded-Postgres
library symlinks; the dependency setup was repaired. Duplicate local
full-suite runs were stopped after full CI completed. This PR does not
claim a completed full local suite.
- Review regression: real Git honors the supplied `GIT_CONFIG_*` setting
and the input environment remains unchanged; all four direct subprocess
cases passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk. Disabling optional index refresh can repeat filesystem stat
work on later scans. Required locks remain enforced; no lock is removed,
no failed reset is retried, and workspace mutation guards are unchanged.
No schema changes.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:28:07 -07:00
Devin FoleyandPaperclip 0ea6b10967 Record ACP activity and workspace restore failure evidence (#14665)
## Thinking Path

Paperclip records terminal run failures for operators. A timeout's own
log and cleanup output update the run's last-output timestamp, so that
timestamp can make a long-silent provider look active. Snapshot runtime
activity before finalization and include the saved workspace restore
classification to make the next failure actionable without copying tool
payloads.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Terminal run diagnostics in the existing opt-in Sentry integration.

**Current behavior**
Reports cannot distinguish runtime events from finalization logging and
omit the already-persisted workspace restore code. A later successful
run also does not establish that earlier workspace files were restored.

**Proposed behavior**
Record runtime-event age/count and pending-tool inventory at
finalization, before status reads and cleanup. Forward only finite
counts, a completeness boolean, and known restore codes through the
existing reporter.

**Reason and benefit**
Operators can distinguish a silent turn with unfinished tools from
recent runtime activity and see restore failures without retrieving
private run output. Neither signal certifies productive work or
successful recovery.

**Breaking changes**
None. Error grouping, execution deadlines, cancellation, recovery
policy, and the Sentry opt-in remain unchanged.

Related diagnostic work: #14573, #14575, #14639.

## What Changed

- Snapshot ACP activity before success/failure finalization, including
thrown relay failures.
- Forward bounded numeric/boolean evidence and shared workspace restore
codes; exclude commands, tool IDs, paths, and arbitrary result data.
- Document limitations and test silence, empty streams, timeout, cleanup
delay, incomplete tool inventory, and privacy.

## Verification

- `pnpm -r typecheck` passed after the final implementation.
- Changed suites: 252 tests passed; all 29 database reporter tests
subsequently passed after restoring the embedded-Postgres package
library symlinks. The migration test also passed (30 database cases
total).
- `pnpm build` passed during implementation. Final head
`fe78dba6f592b1abccac7cdbf341bd2e0b0d30cb` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially hit missing embedded-Postgres
library symlinks; the dependency setup was repaired and database tests
passed. Duplicate local full-suite runs were stopped after full CI
completed. This PR does not claim a completed full local suite.
- Review regression: completed, failed, and cancelled tools are excluded
from the pending count; focused activity/timeout tests and adapter-utils
typecheck passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk, diagnostic-only. The existing tool inventory is incomplete for
some runtime events, so the report carries its completeness flag. Event
age is measured at finalization and does not prove useful work or
identify the underlying provider failure. No schema changes or new
capture gate.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:27:53 -07:00
DottaandPaperclip ad55d0a281 fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use Apps through a gateway that checks identity, company
access, and action policies.
> - Personal pasted credentials can point to company secrets. Setup can
show success while the gateway rejects every call.
> - Several OAuth methods also omit the scopes needed for their
supported write actions.
> - This pull request gives setup, health checks, and invocation the
same credential rules. Owners repair existing connections by
reconnecting.
> - New connections request reviewed permissions for their supported
actions. Read-only choices remain available under Advanced.
> - Agents can use the connections people give them, while existing
consent, identity boundaries, and action restrictions remain enforced.

## Linked Issues or Issue Description

Refs #14009 and #14008. This addresses the personal-credential defect.
The separate GitHub organization-identity selection defect is outside
this change.

Related work: #13942 fixed part of new personal-key setup. #14200
independently fixes legacy personal reconnect and protects managed-agent
profile credentials during removal. This PR covers that ownership
invariant across key and secret-URL setup, reconnect, health, discovery,
and invocation, and keeps owner reconnect as the repair path. #14059
tracks requested versus provider-asserted OAuth scopes; it remains
separate work. I searched open PRs and issues for Zapier, Airtable
scopes, connector writes, and personal credential failures.

**What happened?**

A Zapier secret URL saved through personal setup can become a company
secret referenced by a user grant. Health checks bypass the gateway's
ownership check, so the connection appears healthy but calls fail with
`grant_credential_invalid`. Custom-header paths can also receive a
duplicate `credentials.` prefix. Omitted OAuth scopes make write access
depend on provider defaults.

**Expected behavior**

Personal invocation credentials belong to the selected user. Setup,
health, and actual calls enforce the same rule. New connections request
documented permissions for supported read and write actions. Existing
tokens gain no permissions without provider consent.

**Steps to reproduce**

1. Connect Zapier or a generic secret URL with the personal identity.
2. Allow an agent to use the connection and complete setup.
3. Invoke a tool through a run-scoped gateway. The legacy layout fails
ownership validation despite successful setup.

**Paperclip version or commit**

The implementation started from
`44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master
at `94e8dec56`.

**Deployment mode**

Built from source. Regression tests use isolated PostgreSQL fixtures and
controlled MCP transports.

## What Changed

- Share credential writing, ownership validation, and canonical paths
across initial setup, resume, reconnect, rotation, health, discovery,
and gateway calls. Keep OAuth client-registration secrets separate from
invocation credentials.
- Existing personal connections with company-scoped credentials require
owner reconnect with a fresh key or secret URL. Reconnect creates a
correctly owned value and updates the existing grant and declarations.
There is no automatic ownership backfill or new startup hook.
- Preserve PostgreSQL timestamp precision when reconnect checks whether
a grant changed. Previously, converting the timestamp to a JavaScript
Date could reject reconnect with a false concurrent-change error.
- Protect credentials used by other grants, connections, bindings,
managed-agent profiles, routine triggers, or secret proposals from
connection removal.
- Review all 117 tool methods, including 84 OAuth methods. Record
explicit scopes or documented provider-default exceptions with official
evidence. Add Airtable's seven scopes, Hugging Face repository/job
scopes, and other documented MCP permissions.
- Prefer available write-capable methods. Put explicit read-only choices
under Advanced. Explain pasted-key permissions and offer reconnect for
missing OAuth consent. Preserve existing grants, policies, Google
availability gates, and curated scope allowlists.
- Reconnect generic secret URLs and custom headers using their stored
credential fields. Refresh the catalog after setup, correct reconnect
feedback and error guidance, and let Cancel exit invalid setup while
Save & exit retains draft-saving behavior.
- Apply ownership checks to the new GitHub repository/skill connection
picker. Align the permission audit with the Google scope reductions
merged on master.
- Add run-scoped gateway, ownership, owner-reconnect, OAuth URL,
insufficient-scope, UI, and catalog-wide regression coverage. Update the
connector playbook and permission audit.

## Verification

Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all
CI/status gates (55 completed check runs, no failures or pending checks)
and has a completed Greptile review at **5/5 with no outstanding
findings**. GitHub reports the PR as mergeable/CLEAN.

- **Embedded browser:** used the actual server and built UI from this
worktree, a fresh isolated database, and local HTTP MCP fixtures.
Completed personal bearer-key, secret-URL, and custom-header setup;
reproduced the legacy ownership failure; reconnected through the owner’s
form; and completed writes afterward. Read-back was verified for
bearer-key and secret-URL connections. Public organization-wide setup
appeared immediately in Browse without reload. Zapier URL
validation/Cancel and Google’s enrollment gate were also exercised.
- **Persistence and invocation:** verified user ownership, canonical
`credentials.authorization` / `remote.url` / `headers.X-Api-Key`
declarations, and unchanged connection/grant identity. The old company
secrets retain their ownership. Separate HTTP calls through an actual
run-scoped gateway session completed a write and read-back.
- **Backend coverage:** the final gateway suite passes all 82 cases,
including catalog Zapier and generic inline reconnect. It checks
company/user isolation, canonical declarations, same-endpoint URL
validation, fresh credentials, retained restrictions, and real gateway
read/write execution using fixture transport. A timestamp with
PostgreSQL microseconds covers the former false reconnect conflict.
- **Local checks:** 368 catalog, gateway, repository, and UI tests
passed before the final extra Zapier case; 49 GitHub skill access tests
also passed. All three Apps browser regressions pass, including
reconnect through the actual form and catalog visibility without reload.
Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final
patch, and token gates passed. Full tool-access service runs hit varying
15-second Google fixture timeouts; both affected cases and the updated
reconnect assertion pass in isolation (3 tests). The complete test
matrix passes in CI on this head.
- **Verification limits:** no live provider account was available for
Zapier/Airtable/OAuth consent or account-bound write proof. Public
metadata and local fixtures do not establish provider consent. The
original development database clone failed on a pre-existing missing
`tool_connections_transport_check` constraint; browser acceptance used a
fresh isolated database created by the normal CLI onboarding flow.

## Risks

- Existing broken personal connections stay unusable until their owner
reconnects. Health, discovery, and invocation return an actionable
ownership error; startup does not rewrite credential ownership.
- Scope changes affect new authorization requests. Providers may still
require resource selection, account roles, paid plans, or app
verification. Existing consent and action restrictions remain unchanged.
- Shared credentials are retained rather than reassigned or revoked.
Provider-default exceptions and unavailable live checks are documented
in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`.
- No new endpoint, database table, lockfile change, or CI workflow
change is included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell
execution, web research, and browser tools. The exact deployment model
ID and context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:32:34 -05:00
DottaandPaperclip cbd278dc03 fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agents use saved questions to get human input and continue the same
task.
> - The standard question example recently told models to copy a user
ID.
> - A model can omit an identity prefix and create a question its
intended recipient cannot answer.
> - Agent Chat already knows the conversation owner, so the server can
supply that identity.
> - This pull request removes the blanket instruction and validates
explicit recipients before saving.
> - Ordinary questions stay simple, and explicit addressing remains
available for decisions that need a particular person.

## Linked Issues or Issue Description

Refs #14707, #14188. Related: #14238 handles legacy email recipients;
this change prevents invalid recipients in new cards and retains exact
ID matching.

**What happened?**

A model copied a Cloud user ID without its prefix into
`addresseeUserId`. Creation succeeded. The intended user's answer then
failed the exact recipient check.

**Expected behavior**

Ordinary chat questions use the saved conversation owner. A task may
optionally name a specific recipient. The API rejects an unknown or
unauthorized recipient before it creates a card.

**Steps to reproduce**

Create a chat question for a user whose ID is `paperclip-id:example`.
Supply `example` as the addressee. Before this change, creation accepts
the invalid recipient and the owner cannot answer. With this change,
creation returns 422. Omitting the field saves the full owner ID and
allows that owner to answer.

## What Changed

- Remove `addresseeUserId` from standard question examples and remove
the blanket requester-ID instruction.
- Derive the recipient of ordinary chat questions from the persisted
conversation owner. Reject conflicting explicit user IDs.
- Keep explicit task recipients optional. Validate supplied user IDs
with the existing board mutation policy, including company, viewer, and
Cloud restrictions.
- Preserve explicit agent routing, connector intents, confirmations,
exact recipient checks, idempotent retries, and no-login local-board
authority in local-trusted mode.
- Update the blocker grader to accept an omitted recipient and verify
the actual requester answered.
- Add database and HTTP tests for prefixed identities, denied
recipients, concurrent retries, saved answers, and response delivery.

## Verification

- Database interaction service suite: 90 tests passed, including
implicit local-board creation/answering and authenticated/Cloud denial.
- Interaction HTTP route suite: 84 tests passed.
- Affected interaction/native/connector/documentation suites: 231 tests
passed across six files after valid-user fixtures were updated.
- Resolver and interaction unit suites: 29 tests passed.
- Product E2E unit/calibration suite: 793 tests passed; Product E2E
typecheck and blocker catalog discovery passed.
- Generated API-reference and capability contract checks passed.
- `pnpm -r typecheck` and `pnpm build` passed.
- Full local `pnpm test:run` did not finish green: its initial
general-server pass had 14,416 passing assertions, one unrelated
native-resume assertion failure on macOS, and three teardowns from an
intermediate fixture cleanup fixed above. Separate broad local groups
also encountered timeout/live-port failures under host load. Local UI
(7,026), CLI (502), shared (817), and skills-catalog (20) tests passed;
the complete final-head CI matrix is the broad verification gate.
- After two CI cold-start readiness timeouts, a separate test-only
commit gives the first exposure lifecycle fixture the existing normal
30-second readiness budget. Its real HTTP, ordering, and cleanup
assertions remain intact; the targeted case and final Linux CI shard
passed. Production deadlines are unchanged.
- A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu
because its cached Node executable was group-writable; the same case
passed on AWS runners. The fixture now qualifies its own Linux copy with
mode `0500` and the actual copy digest. Host files and production
security checks are unchanged. The focused macOS case passed; the new
Linux-copy branch also passed on the final AWS-hosted Linux runner
(1,125 passing Runner tests, 3 skipped). The final run was not on a
GitHub-hosted runner.
- Final-head [CI run
36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176)
passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful
checks and two conditional Storybook skips, with no pending or failed
checks. The 27 general/serialized test jobs reported 28,635 passing
tests. Typecheck, build, Runner, browser E2E, and Canary gates passed.
Greptile reviewed that exact head at 5/5; both review threads are
resolved, with no open follow-ups.
- No live provider replay is claimed by this PR.

## Risks

- New explicitly addressed cards reject users who cannot mutate the
issue, including viewers, inactive members, and invalid IDs. Callers
that supplied invalid recipients must correct their request.
- Existing addressed cards are not rewritten. Existing authorization
checks remain strict.
- Chat inference applies only to questions without an agent addressee.
Connector intents and governed confirmations retain their own recipient
paths.
- No schema change or migration is required.

## Model Used

OpenAI Codex, GPT-6 (exact serving variant and context window are not
exposed in this environment). Used reasoning, tool use, code editing,
and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:16:15 -05:00
DottaandPaperclip b54b2dc35c fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path

> - Paperclip manages AI agents and keeps their instructions and files
durable.
> - Native Codex runners can keep a process alive between compatible
turns.
> - Managed file collection stopped that process after each turn, which
defeated warm reuse.
> - Agent folders can contain large images and other files, so full
copies on every turn are expensive.
> - This change keeps one managed directory for the live session and
saves only file changes after each turn.
> - Ownership, authorization, instruction changes, and process
retirement still control when reuse is safe.

## Linked Issues or Issue Description

Related: #13710 introduced native warm session reuse. This fixes managed
file collection that still forced those sessions to stop. No duplicate
open PR or issue was found.

**What happened?**

With managed instructions and warm native Codex enabled, consecutive
turns reused a Daytona sandbox but started a new runner process each
time. The managed directory collector required process termination
before saving files.

**Expected behavior**

Compatible turns keep the same process and managed `AGENT_HOME`. Each
completed turn saves added, changed, and deleted files before the next
turn starts. Unchanged large files do not transfer again.

**Steps to reproduce**

1. Use a native Codex agent with managed instructions and a reusable
Daytona environment.
2. Enable warm session reuse and run three turns on the same task.
3. Write a large binary on the first turn, edit a small note on each
turn, and delete a file on the second turn.
4. Compare process identity across turns and read the canonical files
through the public agent-files API.

**Paperclip version or commit**

Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`.

**Deployment mode**

Local server with remote Daytona execution; cloud native runner uses the
same path.

## What Changed

- Retain the managed directory only for the verified owner of a live
native Codex session.
- Checkpoint each completed turn before releasing the session for reuse.
Retry unstable captures, then stop and collect when a warm checkpoint
cannot be validated.
- Compare metadata and cached hashes, stream only changed file payloads,
record deletions, and validate path, content, quota, and authorization
before saving.
- Rotate sessions when canonical files, loaded instructions,
credentials, or launch policy change. Fence stale collection and cleanup
callbacks from later owners.
- Keep cleanup and recovery aware of the current session owner. Recheck
canonical files under the writer lock at handoff, attach the successor
collector before fallible bookkeeping, and emit one final save receipt
on checkpoint fallback. Preserve storage warnings across unchanged
checkpoints.
- Add regression coverage and a three-turn Daytona test with independent
public API file checks, an unchanged 8 MiB binary, deletion checks, and
strict process identity checks.
- Document checkpoint consistency, lifecycle behavior, and local run-log
counters.
- Replace a timing assumption in the Daytona teardown test with explicit
transfer-arrival gates after CI exposed an unset release callback.

## Verification

- Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks
were repeated after the final storage-warning fix.
- Runner E2E typecheck and 749 runner E2E unit tests passed.
- Focused file checkpoint, directory ownership, instruction collection,
native session, and merge tests passed. After review fixes, the
managed-directory and native-session suites passed 550 tests, including
intervening canonical edits, same-run fresh restore, failed handoff
collection, and one-call fallback collection. Server typecheck passed
again. The Daytona plugin suite passed 218 tests. The quota-warning
regression failed before the fix and passed afterward.
- Three real Daytona campaigns passed before the final handoff review
fixes. The latest kept PID 547 across all three turns. The first
checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes.
Public API reads verified the binary, note contents, and deletion after
every turn. Test cleanup deleted the sandbox.
- The final head was also deployed to an isolated cloud staging instance
and passed three UI-triggered native Codex turns with managed
instructions. All three retained the same process ID/start time, native
session, provider session, runner instance, and Daytona sandbox.
Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes
on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes.
Independent canonical API reads verified every byte of the unchanged 8
MiB binary and the exact note contents after every turn; the deleted
file returned 404 after turns 2 and 3. After restoring the original
lifecycle and agent-auth configuration, removing the temporary secret,
pausing the test agent, and deleting both test sandboxes, independent
canonical API reads still verified the entire binary, the final 78-byte
three-line note, and the deletion. The native runner flag remained
enabled and the final serving revision remained the PR head.
- Two earlier staging attempts are preserved as failures and are
excluded from the acceptance result: a saved ChatGPT login failed with a
provider routing 401, and its subsequent stopped-sandbox retry failed
before provider startup with a closed-lease admission error. The
successful campaign used a fresh sandbox and a temporary encrypted
API-key binding. The stopped-lease retry remains unexplained; this
campaign does not establish recovery of that failed sandbox.
- All [Paperclip CI
gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355)
pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test
partitions, build, typecheck, runner verification, E2E shards, and the
Canary clean public-npm install. Greptile reviewed that exact head at
5/5 with no unresolved review threads or outstanding findings.
- Full local repository coverage used the existing CI partitions, but
the 40,000-file Git streaming stress test timed out and its local retry
was interrupted by macOS thermal emergency sleep; this is not a green
full local suite claim. The exact stress test passed on the final head
in [CI server shard
2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290),
in 111.9 seconds.
- Repeat the live test with configured credentials and a Linux runner
artifact: `pnpm test:e2e:runner -- --id
daytona-warm-continuity.runner-codex.daytona.warm-three-turn`.

## Risks

- This is a file-level checkpoint, not an atomic snapshot of the whole
folder. Background writes after a capture are saved by the next
checkpoint or final stopped collection.
- Metadata scans still visit all paths. Modified files transfer in full;
unchanged files do not rehash or transfer.
- Incorrect ownership or reuse could collect the wrong directory. Run
ownership fences, current authorization, stable capture validation, and
stopped collection fallbacks are covered by tests.
- Warm reuse remains opt-in. No database migration or fleet default
changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and
test execution. The exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:33:37 -05:00
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00
DottaandPaperclip 25c422ba7e fix(apps): request minimal Google service scopes (#14740)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Google connections give agents service-specific tools through
governed credentials.
> - Each connection has a reviewed OAuth scope set.
> - Docs, Sheets, and Slides request Drive permissions in addition to
their own service scopes.
> - Google documents these permissions as alternatives, not combined
requirements.
> - This pull request removes those extra permissions and redundant
Calendar write free/busy access.
> - Users grant fewer permissions without adding tools or changing
connection access policy.

## Linked Issues or Issue Description

Refs #13820. Related #14739 changes other connector permissions; it does
not reduce these Google profiles.

Companion broker PR:
https://github.com/paperclipai/paperclip-cloud/pull/615. Ship the
matching changes together after fresh-grant validation.

**What happened?**

Seven Google profiles request redundant scopes. Docs, Sheets, and Slides
request Drive scopes. Calendar write requests free/busy even though
calendar.events authorizes its availability tool.

**Expected behavior**

Each profile requests only the scopes required for its reviewed tools.
Managed and customer-owned OAuth methods use the same set.

**Steps to reproduce**

Inspect the Google profile registry and the four app definitions on the
base commit. Compare their scope sets with Google's MCP authorization
alternatives linked in the updated documentation.

**Paperclip version or commit**

Base: 94e8dec56b.

**Deployment mode**

Hosted and self-hosted Google Workspace connections.

## What Changed

- Docs, Sheets, and Slides read profiles request only their service
read-only scope. Write profiles request only their service write scope.
- Calendar write keeps calendar-list read-only and event access.
Calendar read keeps free/busy.
- Apply these sets to managed and customer-owned OAuth methods and their
exact-scope tests.
- Document scope rationale, old-grant behavior, and the coordinated
broker rollout.
- Preserve the 21-scope integration union. Drive and Workspace Search
still need their own Drive permissions.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts`: 25
passed.
- `pnpm -r typecheck`: passed with repository-pinned pnpm 9.15.4 after
refreshing locked dependencies.
- `pnpm build`: passed with repository-pinned pnpm 9.15.4.
- `pnpm test:run`: attempted on the machine's Node 26 runtime, then
interrupted after unrelated suite/collection failures. This is NOT a
passing full-local-suite result. CI uses Node 24.
- Node 24 focused diagnostic rerun: `workspace-runtime-exposure.test.ts`
and `paperclip-control-plane-port.test.ts`: 43 passed, 3 skipped. The
first suite hit a local preview-readiness timeout in CI; the failed
shard passed its single retry without code changes.
- Node 24 rerun of the three local assertion-failure suites
(`chat-discord-adapter-patch`, `ai-connections`, and
`heartbeat-active-run-output-watchdog`): 133 passed. These files were
not changed.
- Node 24 rerun of the changed manifest contract suite: 25 passed.
- Latest-head GitHub checks are green at
`da85dcbc89c44d88f3cabd5ada142c922fa51a3c`: 54 passed, 2 intentionally
skipped; no failed or pending checks. Includes typecheck, build,
general/serialized tests, all 8 browser E2E shards, Runner verification,
and canary packaging. Greptile 5/5, no review threads.
- Cross-repository comparison: all 16 app and broker profiles match
exactly.
- `git diff --check`: passed.
- Google definitions are reviewed JSON source inputs preserved by the
ingestion script. The generated TypeScript registry imports them; no
generic provider regeneration is needed.
- Fresh minimal-grant provider testing is outstanding. No new video was
recorded, and existing broader credentials are not proof of minimal
authorization.

## Risks

The Cloud broker must ship the matching exact scope sets. Mixed versions
fail closed. Validate fresh grants in staging before production.
Existing provider tokens are not retroactively narrowed; affected
managed connections must reconnect if their grants retain extra scopes
or omit explicit scope evidence at refresh. Do not revoke the shared
Google project to migrate one connection. This change does not grant
access, add tools, or alter the Google Chat unread-filter block.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, code execution,
and browser tools. The exact runtime model ID and context-window size
were not exposed to the agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:10:42 -05:00
DottaandPaperclip 94e8dec56b fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner sends authorized tool calls to the server and saves their
results.
> - A provider turn can stop while a server write is still running.
> - The old shutdown path invented a failed result that could conflict
with the real result.
> - Truncated execution input and incomplete recovery records made the
failure harder to diagnose.
> - This pull request preserves exact inputs and actual outcomes through
shutdown and restart.
> - Tests force the race and crash boundaries so safe retries do not
repeat writes.

## Linked Issues or Issue Description

**What happened?**

Stopping a turn during a server tool call could record a false failure,
then reject the actual result as a conflict. The diagnostic input
formatter could truncate instruction content before execution. A crash
during saved-result delivery could leave that delivery permanently
indeterminate. Cleanup could hide the first failure, and a retry could
overwrite earlier run logs.

**Expected behavior**

Keep dispatched tools pending until their actual result is known.
Preserve accepted input bytes. Accept identical result delivery without
failing the task. Reject conflicting results with enough evidence to
diagnose them. Recover saved-result delivery without repeating the
business operation.

**Steps to reproduce**

1. Hold an instruction update at the filesystem commit barrier.
2. Stop its provider turn before the server returns the result.
3. Release the write, deliver its result, and replay the same result.
4. Repeat with a restart before and after the delivery receipt is saved.
5. Check that there is one write and one audit row, and that the exact
result survives.

**Paperclip version or commit**

The change was developed from `44736c9c7` and rebased onto `0e5830887`.

**Deployment mode**

Self-hosted server with the native runner. Tests use local runner
processes, scripted providers, and PostgreSQL.

Related work: #12353 added durable semantic tool receipts; #12384 added
durable Codex tool recovery; #12404 bound semantic tools to ACPX
sessions. #14633 covers separate native-provider cancellation and
qualification work. This PR addresses server semantic-tool outcomes and
their durable delivery. No duplicate fix was found. AgentMail discovery
is outside this PR.

## What Changed

- Close turn admission without inventing results for dispatched tools.
Keep pending calls and accept late actual results.
- Accept identical result replay with a diagnostic warning. Include call
identity and both result hashes in real conflict errors.
- Preserve exact execution arguments. Reject prohibited or oversized
input before dispatch. Keep diagnostic previews redacted and bounded.
- Commit instruction-attempt evidence before the filesystem write. Save
completed mutation receipts so concurrent and restarted duplicates
return the first result. Recheck authorization before replay. An attempt
without a completed result stays unknown and cannot execute again.
Definite pre-write failures save and replay their original error without
another write.
- Recover an interrupted saved-result delivery only for backends with
durable result receipts. Never replay an ordinary business operation
with an unknown outcome.
- Preserve the initiating error when cleanup also fails. Record
incomplete settlement evidence. Propagate typed unknown-outcome errors
through the native tool wrapper without creating a false completed tool
result.
- Append run-log attempts and restore the durable log before appending
after local file loss. Reject incomplete restores. Publish a restored
prefix only if the destination is absent so concurrent attempts cannot
overwrite new lines.
- Add deterministic race, crash, replay, authorization, exact-content,
and log-restoration tests. Document their assertions in
`packages/paperclip-runner/docs/durable-recovery.md`.

## Verification

- Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55
applicable checks pass; four conditional/manual checks are skipped. This
includes build, typecheck, Rust, both runner TypeScript shards, server
and workspace tests, all eight browser shards, isolated runner
compilation, and the clean-install release dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36746101110).
- Greptile reviewed this exact head at 5/5 with zero new findings. All
three earlier review threads are resolved.
- Focused local verification includes 11 instruction integration tests,
23 surrounding authority/tool tests, 26 run-log tests, and 169
controller/driver tests. The post-rebase controller/transport/runtime
selection passed 415 tests. The full Rust release suite passed 617 tests
with two ignored. The real-process SIGKILL recovery test passed three
consecutive runs.
- The fault matrix in
`packages/paperclip-runner/docs/durable-recovery.md` uses explicit
barriers, real PostgreSQL rollback, durable journal reloads, and killed
runner processes. It covers late results, identical and conflicting
replay, exact long content, concurrent log restoration, lost commit
acknowledgements, and definite failure replay after the original CAS
base becomes valid again. No paid model calls are needed.
- Full local recursive typecheck and build passed during implementation.
Server typecheck and the runner TypeScript build passed after the review
fixes. The broad local repository test run was stopped after repeated
database startup timeouts. Four timing/launch failures in an earlier
broad runner run passed focused reruns without changed assertions or
timeouts. These are local verification limitations; the complete
current-head CI suite is green. An earlier CI workspace job received an
infrastructure shutdown signal; its current-head replacement passed.

## Risks

- A stopped turn can remain blocked when a dispatched operation has no
proven result. The system does not guess its outcome or rerun its
effect.
- Conflicting results still fail settlement. Existing failed or
conflicting journals are not repaired automatically.
- Accepted semantic input is limited to 480 KiB of encoded JSON to fit
the encrypted transport. Larger input fails before execution.
- Instruction filesystem writes and database receipts are not one atomic
storage operation. A separately committed attempt and audit record
survive rollback. An attempt without a completed success or definite
pre-write failure receipt remains blocked as an unknown outcome. It is
not replayed or reported as success.
- Run-log restoration now reads the durable object before appending when
the local log is missing. Failed or incomplete reads reject the append.
- No schema migration, dependency change, workflow change, or AgentMail
change is included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; the broad
local run limitation is recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:54:56 -05:00
Nicky LeachandPaperclip af5c2d101c fix(paperclip-runner): deliver the shutdown settlement event past the terminal-turn gate (#14668)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The OpenCode driver maps provider events to runtime request events
> - A provider turn can end before a pending runtime request receives
its answer
> - The consumer reads one turn's events and stops at that turn's
terminal event
> - A settlement event that arrives after that terminal event never
reaches the consumer
> - This pull request settles the request inside its own turn, before
the terminal event
> - The benefit is reliable request settlement without weakening
late-frame protection

## Linked Issues or Issue Description

**What happened?**

A pending runtime request stayed open after an OpenCode turn failed
through `session.error`. Session shutdown then dropped its settlement
event as a late provider frame.

**Expected behavior**

The driver must deliver `runtime_request.expired` with the original
`turnId` and `itemId`, inside the same single read pass the consumer
performs on that turn.

**Steps to reproduce**

1. Start an OpenCode turn that creates a native runtime request.
2. Leave the request pending and fail the turn through `session.error`.
3. Close the session and inspect the emitted events.

**Paperclip version or commit**

`22d41c6081f05658e0d7c8485d0f22f35af4a79e`

**Deployment mode**

Built from source with the OpenCode driver test fixture.

## What Changed

- Settle a pending runtime request as soon as its own turn goes
terminal, before the terminal turn event.
- Add an optional `bypassTerminalTurnGate` parameter to the OpenCode
session emitter, and set it on the settlement emit.
- Keep a settlement loop in session close as a fallback for a request
whose turn never went terminal.
- Add a fixture trigger and a regression test that reads one turn in a
single pass.
- Keep the late provider frame gate unchanged for every provider event
path.

## Verification

- Run `pnpm vitest run
packages/paperclip-runner/src/drivers/opencode/opencode-server-driver.test.ts`.
- Run `tsc -p tsconfig.json --noEmit` in `packages/paperclip-runner`.
- Run `tsc -p tsconfig.surfaces.json --noEmit` in
`packages/paperclip-runner`.
- Confirm that the regression test receives `runtime_request.expired`
with the original identifiers.
- Confirm that the late-frame tests still report dropped provider
frames.

## Risks

Four call sites now reach the settlement path: session close and the
three terminal turn paths (completed, cancelled, and failed). At the
three terminal turn paths the turn is still the active turn, so the gate
admits the settlement event with or without the parameter. Session close
is the only place where the parameter changes the result of the gate,
and only for a request whose turn already ended. Every provider event
path keeps the existing terminal-turn gate. The known driver test
failure is pre-existing and does not touch this change.

## Model Used

Claude Sonnet 5, with code execution and test assistance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the changed tests locally; the known pre-existing
failure remains documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation, or no documentation change
applies
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I have addressed all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 09:48:14 -07:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip a36cbffa9e fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path

> - Paperclip manages agents and the services they need for work.
> - Agents use connection search to discover a setup path before they
request access.
> - Tool-only filtering hid channel and AI methods from this search.
> - Requiring every query word to match rejected normal task
descriptions.
> - A small aggregator index also omitted supported apps such as
Circleback.
> - This pull request broadens retrieval and returns purpose-specific
setup guidance.
> - Agents can choose a relevant result while existing access and
provider-choice checks still apply.

## Linked Issues or Issue Description

**What happened?**

A search such as “AgentMail create an email address and manage an agent
mailbox” returned no usable result. “Help me find tools for circle back”
also missed Composio's supported Circleback toolkit. Queries longer than
200 characters failed validation.

**Expected behavior**

Return useful native and verified aggregator matches from
natural-language queries. Include channel/email methods when Chat
connectors is enabled. Identify each method's purpose and the correct
setup path.

**Steps to reproduce**

Use the queries above with `connections_search` from an active task.
Enable Chat connectors for the AgentMail case. The regression suite
reproduces these misses before the change.

Related routing work: #13941. This change does not change the runner
failure path or add channel setup to tool-only connection cards.

## What Changed

- Rank name and capability matches. Accept extra words, split names,
small spelling errors, and queries up to 4,000 characters.
- Include tool, channel/email, and AI methods. Return company-prefix
setup links for channel and AI flows.
- Add a dated snapshot of 1,583 official Composio toolkit names and a
refresh script. Merge duplicate MCP variants for search and link each
support claim to official evidence.
- Find authorized indexed aggregator namespaces within longer queries.
Return multiple app matches when the agent needs to choose.
- Prefer exact app names over fuzzy matches for other apps; retain
existing AI readiness.
- Preserve native preference, company and identity boundaries,
administrative denials, and saved provider consent.
- Add relevance and database regressions, extend native tool-authority
coverage, and document search behavior.

## Verification

- Red: 15 new assertions failed against the previous implementation; the
existing baseline passed. Added red-green regressions for Motion versus
fuzzy Notion and existing AI access during review. A further regression
covers mixed ready/unconfigured AI results and their per-result setup
guidance.
- Green: all 103 tests in the eight focused shared, database,
runtime-tool, fixture, and route suites pass on the latest commit.
- `pnpm -r typecheck` and `pnpm build` passed.
- Latest-commit CI passed: 54 successful checks and two skipped checks,
including the full test matrix, browser E2E, typecheck, build, and
canary dry run.
- The long local `pnpm test:run` attempt began before the review fixes
and retained transformed pre-fix search code; it also hit an unrelated
timing failure. Fresh serial reruns of the affected search suites and
three timeout cases passed all 124 tests. Parallel local route shards
hit two additional database setup timeouts; both suites passed all 17
tests on a fresh serial rerun. The complete corresponding CI suites also
passed. Duplicate broad local runs were stopped after CI completed. The
local UI suite independently passed all 7,007 tests.
- Greptile: 5/5 on `cc6a0180d`; all review findings resolved.
- Browser verification passed in a disposable local instance through
real process-agent search requests: AgentMail opened its setup flow with
the requester selected; the saved Circleback choice produced the
Composio setup card; a paragraph-length Notion query produced its setup
card. No provider credentials or external accounts were created.
- The browser test caught an invalid UUID-based setup URL. The fix uses
the company prefix and has a regression assertion.
- A 3,971-character catalog query found Circleback first in a local 10
ms spot check after sharing query preparation across the catalog scan.
This is a single measurement, not a performance guarantee.

## Risks

- Broader retrieval can return extra candidates. Named services rank
first; agents must select the relevant method.
- The public support snapshot can age. It proves catalog support, not
account authorization or the availability of every requested action.
- Channel and AI methods use existing setup links. The tool connection
card still accepts tool methods only.
- No schema, migration, credential, or runner lifecycle changes.

## Model Used

OpenAI Codex (GPT-6). The session does not expose a more specific model
identifier or context-window size. Used reasoning, repository search,
code execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:20:50 -05:00
DottaandPaperclip dd7fc1f90a fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native sessions persist control-plane state so they can resume
safely.
> - The state includes committed provider history needed for recovery.
> - The server, runnerd recovery, and durable control plane validate
this state before trusting its identity.
> - Their differing 64 MiB and 192 MiB limits can reject a valid journal
before recovery.
> - This pull request aligns all three local state limits at 256 MiB.
> - Larger files remain bounded, while recovery can read larger valid
histories.

## Linked Issues or Issue Description

Refs #13882
Refs #14312

## What Changed

- Raise the server and runnerd recovery limits from 64 MiB to 256 MiB,
and align the durable control-plane limit from 192 MiB to 256 MiB.
- Add coverage for a valid history above 64 MiB and rejection above 256
MiB.

## Verification

- Matching server recovery passed with more than 64 MiB of actual
committed event payloads (128 events with 512 KiB deltas).
- The real runnerd exact-authority resume regression with the test Codex
provider passed with 193 MiB of valid JSON whitespace appended. It
crosses the former 192 MiB core limit and confirms the same provider
identity. This exercises the runner process and durable control plane
with a simulated provider, not a live OpenAI API call. This test used
approximately 1.15 GiB peak RSS.
- The actual runnerd reader accepted valid 256 MiB JSON and rejected
valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process
peaked at 1,244 MiB RSS.
- Server tests reject mismatched identity above 64 MiB and files above
256 MiB.
- `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed.
- Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed
across 6 files. All failures were embedded PostgreSQL startup errors
after five attempts. They affected agent hiring, instruction revisions,
environment images, reviewed chat bindings, issue monitoring, and legacy
continuation authority. The focused journal tests passed; the latest
pushed head passed all ordinary CI checks. Superagent is the only
blocking check.

## Risks

- **Open review concern:** Greptile is 5/5, but Superagent is
`ACTION_REQUIRED` with two P2 findings on the server and runnerd
readers. Both flag the increased synchronous parsing and memory cost.
This PR keeps the requested fixed-limit change small. It does not add a
worker parser or a process-wide memory budget. This resource tradeoff
needs review before merge.

- Large state parsing is synchronous and can consume several times the
file size in memory.
- Remote checkpoint archive and expanded-size limits remain 64 MiB, so
this change alone does not make larger remote checkpoint transfers
portable.

## Model Used

- OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at
high reasoning effort; tool use and code execution. The GPT-6 context
window is not exposed in this task runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:08:25 -05:00
DottaandPaperclip d30b03bd8c test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Agents use the coordination skill when work needs human authority or
a scope decision.
> - PR #14188 replaced automatic manager escalation with direct blocker
handling.
> - This behavior needs real browser, server, database, and provider
tests.
> - The test must verify saved human input, task ownership, and resumed
work.
> - This pull request adds six reusable Product E2E cases and improves
the skill examples that they exercise.

## Linked Issues or Issue Description

Refs #14188.

The merged change needs repeatable behavior coverage. The new suite
tests missing administrator access, missing hiring permission, and
requester scope questions. Searches found no duplicate blocker-guidance
suite. This extends the existing eval system described in ROADMAP.md.

## What Changed

- Add the explicit-only `blocker-guidance` Product E2E suite. It has
three local scenarios on legacy Codex and legacy Claude.
- Use the production UI and public APIs to create work, save a
human-only question or confirmation, answer it after reload, and resume
the same task.
- Check requester identity, ownership history, manager activity, hiring,
saved answers, and completion. Keep direct text input as a separate UX
result.
- Save pending and final screenshots, API checkpoints, skill hashes,
provider evidence, and billing data through the existing report
pipeline.
- Isolate the Claude fixture home. Verify the served skill bytes before
dispatch so an old installed skill cannot silently replace the evaluated
skill.
- Improve the coordination and hiring skill examples. Include the
human-only policy, requester address, wake behavior, and handling of
authorized scope changes.
- Grader v5 requires the exact approved public welcome note as a new
worker comment. Browser input checks reject unwritable scope cards
before clicking, and confirmation direction must be saved in the
resolution before the worker wakes.
- Add grader calibration and browser-input tests. Update the fixture
guide and generated capability inventories.

## Verification

- `pnpm build`: passed after rebasing onto current master.
- `pnpm -r typecheck`: passed.
- `pnpm test:e2e:runner:typecheck`: passed.
- `pnpm test:e2e:runner:unit`: 742 tests passed.
- `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests
passed.
- `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells
found.
- Capability inventory and generated-contract checks: passed.
- `pnpm exec vitest run
server/src/__tests__/hiring-operational-examples.test.ts`: four tests
passed after synchronizing the generated API reference and section
anchor.
- Full general and serialized test suites: passed in CI on
`6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm
test:run` was interrupted after complete CI coverage passed; it is not
claimed as a completed local full-suite run.
- Final GitHub checks: 54 passed, two optional Storybook checks skipped.
The runtime-exposure startup test hit a 10-second readiness timeout
once, passed a targeted local reproduction, and its CI shard passed the
single retry without code changes.
- Current-head Greptile: 5/5, clean check, zero unresolved threads.
- Historical live measurement on September 29 at
`4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell
runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex
`gpt-5.6-sol` passed 8/9. These runs predate this rebase.
- Version 5 changes the scope answer to an exact approved publication
draft. The historical runs do not qualify that new requirement; the
two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec`
passed Codex and failed Claude. Claude posted the correct salary-free
sentence but omitted its required reference line from that comment,
placing the reference in a separate completion message. The
`public-welcome-note` check correctly failed. An earlier Claude
database-startup failure was retained separately; its fresh-instance
retry reached the model. This pilot is not a six-cell qualification.
- The failed Codex scope case omitted `addresseeUserId`. The strict
routing check remains. All 18 attempts had clean evidence manifests and
passed cleanup.
- To repeat with provider credentials: `pnpm test:e2e:runner -- --suite
blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is
excluded from `--all`.

## Risks

- The live suite measures variable model behavior. The retained 17/18
historical result and the current 1/2 scope pilot are not all-pass
qualifications. These paid cases are opt-in; their observed model
failures remain visible independently of deterministic CI checks.
- A separate generic task-replacement diagnostic still exposed a Claude
refusal. The ordinary cases use specific business decisions. The
diagnostic is not a standalone catalog case in this change.
- Earlier measurements included an old installed Claude skill and test
defects. Their grades remain retained and are not combined with the
three final repetitions.
- Skill examples can affect when agents ask for human input. Downstream
permission checks still apply.
- Native runners, Daytona, agent-requester routing, and real external
connection authorization are outside this suite.
- Raw provider traces and credentials remain private. No screenshots,
raw reports, secrets, workflow changes, or lockfile changes are
committed.

## Model Used

OpenAI GPT-6 through Codex assisted with this change. The exact deployed
variant and context window size are not exposed in this session. The
assistant used reasoning, repository edits, tool use, and shell
execution. The evaluated models were `gpt-5.6-sol` and
`claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:56:54 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip f4f9a7c613 test(runner): guard continuation after journals exceed 2 MiB (#14312)
Add an actual runner resume regression above the former 2 MiB journal boundary and an explicit-only three-turn Daytona workflow that grows real execution history. Verify journal size and distinct completed tool calls without exporting private payloads.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 07:01:45 -05:00
DottaandPaperclip 2f6fa3b6dc fix: recover provider authentication inside tasks (#14629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a working model provider connection to run a task.
> - A provider can reject a stored credential after the task starts.
> - The failed run must ask the responsible user to repair that
connection.
> - This pull request adds that request directly to the task and reuses
Connections sign-in.
> - The user can choose an API key or subscription, then continue the
task with a fresh session.

## Linked Issues or Issue Description

**What happened?**

A run that ended with `acpx_auth_required` or another known provider
authentication error did not immediately offer an inline way to connect
the provider. A repair form could also lock the user to the failed
account's sign-in method.

**Expected behavior**

Show a provider connection card in the task as soon as the
authentication failure is saved. Allow the responsible user to connect
or repair the provider with any supported sign-in method. Keep the
connection name automatic and resume the task after successful setup.

**Steps to reproduce**

1. Run a task with a supported provider and an expired or invalid
credential.
2. Let the run fail with a provider authentication error.
3. Open the task and attempt to repair the connection.

Related work: Refs #13724 and #13726. This change adds the inline task
repair flow and method choice.

## What Changed

- Classify provider authentication failures and create one connection
request for the current task. A persisted blocked classification
suppresses automatic retries only after the repair card is created;
unsupported providers retain their existing recovery path.
- Mark only the attributed, unchanged credential as needing sign-in.
Preserve credentials that were refreshed after the failed run started.
- Reuse the provider sign-in controls inside the task. Allow API key and
subscription choices for Claude, Codex, and Grok. Keep names hidden and
generate a default from the user, provider, and method.
- Keep the existing account when reconnecting with the same method.
Create and select another account when the method changes. Validate
updates to explicit agent bindings through the normal agent save path.
- Require explicit adoption for legacy agent authentication. Validate in
the agent environment, then commit the binding, connection install,
audit, and card completion in one transaction. Keep failed setup and
account selection visible and retryable.
- Add regression tests and update the specification and Connections
documentation.

## Verification

- Fresh local verification: 199 tests passed across the inline form,
provider method selector, default naming, authentication and recovery
classifiers, run liveness, OpenAPI routes, database adoption/rollback,
and Cursor execution suites. The adoption database suite also passed
against disposable Docker PostgreSQL.
- Full repository `pnpm build` and `pnpm -r typecheck` passed on the
latest commit. Token gates are clean.
- Embedded browser: opened real task cards from seeded authentication
failures; switched Claude from API key to subscription and back;
switched Codex from subscription to API key; confirmed the name field
stays hidden. Provider sign-in was not completed with real credentials.
- The broad local `pnpm test:run` started before review fixes and was
interrupted after the working tree changed; it is not counted as a
passing full run. Fresh focused tests passed. CI supplies the full test
and browser suite results for the current commit.
- CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54
checks passed and two Storybook checks were skipped by their path rules.
The workspace preview job passed on one rerun after a local-server
startup timeout; its rerun passed 835 tests.
- Greptile is 5/5 on the same commit with no actionable findings and no
unresolved review threads.

## Risks

- Incorrect authentication classification could prompt for a connection
unnecessarily. Tests exclude tool authorization, quota, and unrelated
runtime failures.
- A method change selects the new personal provider default, which also
applies to other agents that use that user's default. Explicit account
bindings use the existing permission and runtime validation path.
- Credential invalidation must not race with refresh or reconnect. The
code compares the saved credential generation and grant update time
under locks.
- No database migration or new credential storage format is required.

## Model Used

OpenAI GPT-6 through Codex. The exact model ID and context window size
were not exposed in this session. Capabilities used: reasoning,
repository editing, shell commands, database tests, and embedded-browser
interaction.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests listed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 06:39:37 -05:00
Nicky LeachandPaperclip eb31b926a1 fix(runner): keep the OpenCode session event stream open across turns (#14582)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip Runner keeps provider sessions and their event streams for
agent runs
> - The OpenCode driver closed its event queue after each terminal turn
> - A second turn on the same session then lost its response and
completion events
> - This pull request keeps the queue open between turns and rejects
late events for sealed turns
> - The benefit is reliable multi-turn OpenCode sessions with visible
diagnostics for late provider events

## Linked Issues or Issue Description

**What happened?**

The OpenCode driver closed its event queue when a turn completed, was
cancelled, or failed. A second turn on the same session then lost its
response and completion events.

**Expected behavior**

The session must keep its event stream open between turns. Each turn
must deliver its response and one terminal event. The session must close
the stream only during session shutdown or an unrecoverable pump error.

**Steps to reproduce**

1. Start one OpenCode session.
2. Run one turn and wait for its terminal event.
3. Run a second turn on the same session.
4. Confirm that the second turn delivers its response and terminal
event.

**Paperclip version or commit**

`c0e1d87ddc181329471fa80a2061b2c538bb6618`

**Deployment mode**

Built from source with the Paperclip Runner package test suite.

## What Changed

- Keep the OpenCode event queue open across completed, cancelled, and
failed turns.
- Track sealed turn ids and reject later events for those turns with a
diagnostic event.
- Preserve queue shutdown on session close and unrecoverable pump
errors.
- Give each simulated fixture turn unique provider event ids.
- Add regression tests for completed, cancelled, failed, and
closed-session paths.

## Verification

- Type check: `cd packages/paperclip-runner && node
./node_modules/typescript/bin/tsc -p tsconfig.json --noEmit` passed.
- Driver tests: `cd packages/paperclip-runner && npx vitest run
src/drivers/opencode/opencode-server-driver.test.ts` passed except for
the known pre-existing flaky test described below.
- Consumer tests: `cd packages/paperclip-runner && npx vitest run
src/native-session-runtime.test.ts
src/backends/harness-driver-backend.test.ts
src/cli/opencode-app-server-proxy.test.ts
src/conformance/harness-driver.test.ts` passed.
- The known flaky test reproduced on unmodified `master` because fixture
event order depends on a local MCP HTTP round-trip.
- CI must run the full pull request suite.

## Risks

- The queue now retains sealed turn ids for the session lifetime.
OpenCode does not reuse turn ids, so this set grows with the session.
- A late provider event cannot reach a later turn. The driver emits a
diagnostic event so the rejection remains visible.
- The change does not alter session shutdown or unrecoverable pump error
handling.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. This pull request fixes an
OpenCode Runner bug and does not add a roadmap feature.

## Model Used

OpenAI Codex, GPT-5, tool use and code execution; Anthropic Claude
Sonnet 5 also assisted with the implementation. The runtime did not
provide a context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 22:05:41 -07:00
Devin Foley f38b5693f6 fix: always enable keyboard shortcuts (#14643)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI has keyboard shortcuts for the inbox, task lists, cases,
and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and
`]`
> - Shortcut enablement was an instance-wide General setting until
#14141 moved it to a per-user preference that defaults to off
> - The move did not carry the old instance value over, so every
existing user lost shortcuts on upgrade and had to find a new toggle
under Profile settings
> - A toggle that only turns off a standard, input-safe feature costs a
setting, a database column, two API routes, and a React context for
little benefit
> - This pull request removes both the instance setting and the personal
preference and enables keyboard shortcuts for every signed-in user
> - The benefit is one less thing to configure, no silent loss of
shortcuts on upgrade, and less code to maintain

## Linked Issues or Issue Description

Refs #14141 (the change that introduced the personal preference).

**What existing behavior does this improve?**

Keyboard shortcuts in the web UI stay off unless each user turns them on
in Profile settings.

**Subsystem affected**

Web UI shortcuts, Profile settings, instance general settings, the
`/api/auth/preferences` routes, and the `user` table.

**Current behavior**

Shortcuts default to off per user. #14141 moved the toggle from Instance
settings → General to Profile settings and did not carry the old
instance value over. Users who had shortcuts on lost them after the
upgrade and had to find the new toggle.

**Proposed behavior**

Keyboard shortcuts are always enabled for every signed-in user. There is
no instance setting and no personal preference. Shortcuts already ignore
key presses inside text inputs and modal dialogs, so an opt-out is not
needed.

**Reason and benefit**

Fewer settings, no silent loss of shortcuts on upgrade, and removal of a
database column, two API routes, a query hook, and a React context that
existed only to gate this feature.

**Breaking changes**

`GET` and `PATCH /api/auth/preferences` are removed. `PATCH
/api/instance/settings/general` no longer accepts `keyboardShortcuts`;
that schema is strict, so the key now returns 400.
`instance.general.keyboardShortcuts` is no longer a valid
`PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a
warning.

## What Changed

- Removed the Keyboard shortcuts section from Profile settings, the
`useUserPreferences` hook, `queryKeys.auth.preferences`, and
`authApi.getPreferences` / `authApi.updatePreferences`.
- Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list,
legacy task list, cases, and task detail pages no longer gate their key
handlers.
- Removed the `enabled` option from `useKeyboardShortcuts`. The app
shell always registers the global shortcuts.
- Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI
entries, and the `currentUserPreferencesSchema` /
`updateCurrentUserPreferencesSchema` validators.
- Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the
general settings zod schema, the settings service defaults, and
`HIDEABLE_GENERAL_SECTIONS`.
- Added migration `0289_drop_user_keyboard_shortcuts`, which drops
`user.keyboard_shortcuts`.
- Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and
`docs/deploy/environment-variables.md`.
- Parsed the stored general settings row with
`instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a
retired key left in the row cannot reset the sharing preference to
`prompt` and overwrite the stored choice.
- Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open
modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before.
- Updated the affected tests and added a Profile settings test that
asserts the toggle is gone, a hook test for the modal dialog guard, and
a feedback service regression test for the retired-key case.

## Verification

- Typecheck passes for `@paperclipai/shared`, `@paperclipai/db`
(including the migration numbering and safety checks),
`@paperclipai/server`, and `ui`.
- `pnpm exec vitest run
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/openapi-routes.test.ts
server/src/__tests__/auth-routes.test.ts
server/src/__tests__/sentry.test.ts` → 119 passed.
- `pnpm exec vitest run ui/src/components/Layout.test.tsx
ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx
ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx
ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx
ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed.
- `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts`
→ 16 passed.
- `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7
passed.
- `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts`
(embedded Postgres) → the new retired-key test passes with the fix and
fails without it.
- Manual: sign in with no settings changed, open the inbox, press `j`
and `k` to move the selection, press `?` to open the cheatsheet. Open
Settings → Profile and confirm there is no Keyboard shortcuts section.

## Risks

- The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the
column has no readers after this change. If you roll back to a build
from before this PR after the migration has run, re-add the column
first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean
DEFAULT false NOT NULL;`. The older build's ORM selects that column when
it loads users.
- Any external client that still sends `keyboardShortcuts` to `PATCH
/api/instance/settings/general` receives a 400. No in-repo client does.
- Stored `instance_settings.general.keyboardShortcuts` values are
stripped on read and ignored.
- Users who never turned the toggle on now get shortcuts. The handlers
skip text inputs, contenteditable regions, and modal dialogs, so typing
is unaffected.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended
thinking and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-29 21:28:06 -07:00
Devin FoleyandPaperclip 35a4448c02 fix(cursor): select failure diagnostics after trace notices (#14636)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - The Cursor CLI adapter turns process output into run results.
> - Cursor can print a trace-file notice before a real error.
> - The adapter used the first stderr line as the failure summary.
> - This could hide the error behind an informational file path.
> - This change selects the first diagnostic after that known notice and
preserves the full logs.

## Linked Issues or Issue Description

**What happened?**

When Cursor exits with a nonzero code, a leading `cursor-retrieval:
tracing to ...` notice can become the error summary. A later error
remains in stderr but is absent from the summary. If the notice is the
only output, the summary does not explain that the process exited
unsuccessfully.

**Expected behavior**

Prefer the structured error, then a stderr diagnostic, then the exit
code. Keep the run failed and preserve the original logs.

**Steps to reproduce**

1. Use a fixture Cursor executable that writes the trace-file notice to
stderr.
2. Write `Authentication failed` on the next line, then exit with code
7.
3. The old adapter reports the trace-file notice. This change reports
the authentication error.
4. Repeat with only the notice. This change reports `Cursor exited with
code 7`.

**Paperclip version or commit**

Reproduced against master commit
`17780751551b3bc1c2521f7694026c34534c46c9` with local process fixtures.

**Deployment mode**

Local CLI adapter. The diagnostic helper is also used by environment
probes.

Searched open Cursor PRs and issues. PRs #14631 and #14435 concern
native ACP support; #11106 concerns MCP configuration. None changes this
legacy CLI diagnostic selection.

## What Changed

- Skip only the exact trace-location notice when choosing a diagnostic
line.
- Remove terminal control codes from summary candidates.
- Use the same selection for execution and environment probes.
- Preserve structured-error priority, exit status, retry decisions, and
raw stdout/stderr.
- Add child-process regression tests and narrow parsing cases. Document
the behavior.

## Verification

- Two execution regression cases failed on the previous implementation;
structured-error priority already passed.
- `pnpm exec vitest run packages/adapters/cursor-local`: all 16 tests
passed across five files.
- `pnpm --filter @paperclipai/adapter-cursor-local typecheck` passed.
- `pnpm -r typecheck` passed.
- All GitHub CI checks passed on the PR head. One preview-runtime
readiness test failed on the first attempt; its full local suite passed
(28 tests, three skips) and the failed CI shard passed on retry. No
unrelated source change was needed.
- The broad local `pnpm test:run` command did not complete in the
available verification window and was stopped; no full local-suite pass
is claimed. The full sharded GitHub CI suite passed. `pnpm build`
passed.
- Tests use local fixture processes. They make no Cursor provider
requests.

## Risks

A future Cursor notice format may no longer match and will remain
visible. Retrieval error lines and unknown diagnostics remain visible.
This improves diagnosis; it does not claim to fix an unknown provider or
machine failure. There are no schema, authentication, cancellation, or
retry-policy changes.

## Model Used

OpenAI Codex (GPT-6), with reasoning, repository inspection, and command
execution. The session does not expose a more specific model revision or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run the targeted tests locally and they pass; full checks
are in progress
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:36:36 -07:00
DottaandPaperclip 2de43fc909 fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Each task has one assignee. Explicit assignment and review requests
select who should act.
> - An agent mention started another agent on a task it did not own.
Native attachment staging then rejected that run.
> - Allowing that run through startup could also let two agents work on
the same task.
> - Mentions should identify relevant context. They should not start
work or forward comments to other tasks.
> - A personal app installed on a shared agent must also wait until tool
use to resolve the current user's grant.
> - This pull request removes mention dispatch and keeps missing
personal app credentials from blocking startup.

## Linked Issues or Issue Description

**What happened?**

A native agent mentioned on another agent's task failed with
`paperclip_runner_attachment_staging_not_authorized`. The source task
could already be complete. A nearby optional-app warning was a separate
problem: personal app tools were excluded when their shared health state
required attention.

**Expected behavior**

An agent mention is context only. It does not wake the agent, take
ownership, or copy a comment onto another task. Normal feedback still
reaches the assignee. Assignment and explicit review requests still
dispatch work. An unavailable personal app does not block startup or
produce a startup warning. Tool use requests the current user's
authorization and never uses another user's grant.

**Steps to reproduce**

1. Assign a task to agent A. Post a comment that mentions agent B,
including a comment that closes A's task or references B's child task.
2. Confirm the comment retains its agent link and B receives no run or
deferred wake. A can still receive normal feedback.
3. Install an active personal MCP connection on B. Give only Alice a
grant and leave shared health at `error`.
4. Explicitly assign work to B for another user. Confirm it can finish
without using the app.
5. Ask B to use the app. Confirm its tool call shows an inline
connection request for the current user.

Related work: Refs #11144. This change uses the existing execution-time
personal grant resolution.

## What Changed

- Remove mention dispatch from standalone comments and issue updates.
Remove implicit forwarding of parent comments to a mentioned worker's
child task.
- Ignore new requests with the legacy mention wake reason before
creating a run or deferred request. Preserve already accepted queue
entries, which can combine assignments and feedback with a later
mention.
- Remove the native mention admission, staging, and finalization
exceptions from this PR. Native task ownership checks remain intact.
- Keep active, installed personal app tools available despite shared
health errors. Remove optional-app startup warnings. Tool execution
retains the current user's grant and policy checks.
- Update agent instructions and product/API docs. Refresh generated
capability source anchors.

## Verification

- Red: comment-route regressions reproduced extra agent wakes and child
comment forwarding. A separate regression proved that cancelling by the
last coalesced reason could drop an accepted assignment.
- Green: the targeted route, wake queue, heartbeat, workspace,
responsible-user, MCP discovery, and HTTP gateway suites passed. The
final queue and heartbeat rerun passed 104 tests, the restored queue
adapter passed 56, and both comment-route suites passed 135. These
include accepted assignment preservation, rejection of new mention
requests, and normal assignee feedback.
- `pnpm -r typecheck` and `pnpm build` passed locally. The full local
`pnpm test:run` attempt was interrupted for review/CI fixes, so it is
not claimed as a completed local pass. It exposed a cleanup timing race
in the concurrent-mention assertion, now fixed and verified across 10
repetitions. CI also exposed an obsolete test waiting for the removed
mention lookup; it was reproduced and fixed, then both comment suites
passed. Final full-suite verification is through CI.
- Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks
passed, 2 Storybook checks intentionally skipped; no pending or failing
checks. Full CI includes general and serialized suites, all 8 browser
shards, runner verification, typecheck, build, and canary dry run.
Greptile is 5/5 on this exact commit, with no unresolved findings.
- One unchanged Cursor adapter test hit its 10-second CI timeout. All 5
tests in that file passed locally; one retry of its CI shard passed all
674 tests (3 skipped). The aggregate verification gate then passed. No
code or timeout was changed for that retry.
- Live browser check: inserted a structured mention with the picker on a
human-owned task. The saved link remained visible. Database checks found
zero new runs and zero wake requests.
- Live Codex runner check: explicitly assigned that task with the
unavailable personal app attached. The run succeeded and committed
completion without using the app or creating a connection card.
- Live browser follow-up: asked the assignee to call PostHog and
mentioned another enabled agent as context. Only the assignee ran. It
succeeded and displayed the existing inline connection card. Only
Alice's grant existed; the run belonged to a different user.
- The HTTP regression covers tool discovery with no provider calls or
connection cards, first use returning the current user's authorization
request, and successful retry after that user's grant exists.
- App checks use an isolated local fixture and a fake MCP provider. They
do not use production app credentials.

## Risks

- Intentional behavior change: workflows that used mentions to wake
agents must use assignment, a bounded child task, or an explicit review
request.
- Already accepted queue entries retain their prior rules. An old entry
can combine assignment or feedback with a later mention; its last reason
cannot safely identify mention-only work. New mention requests create no
run or deferred wake.
- Personal apps with a shared health error remain discoverable. Actual
tool use still requires the responsible user's grant and existing policy
gates.
- No database migration or public API schema change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact serving model ID and
context-window size are not exposed in this session.
- Live native-run verification used `gpt-6-astra` through the Codex
provider.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:49:01 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
da887ea3e9 fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path

> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.

## Linked Issues or Issue Description

Refs #14322

**What happened?**

The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.

**Expected behavior**

The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.

**Steps to reproduce**

1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.

## What Changed

- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.

## Risks

- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.

> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.

## Model Used

OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
2026-09-29 09:41:05 -05:00
DottaandPaperclip 7636966452 fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Mine inbox shows work that needs the current user.
> - Failed-run rows used the latest run for every agent in the company.
> - A failure from another user therefore appeared in Mine and its
badge.
> - Run list responses also omitted the responsible user needed to
filter these rows.
> - This pull request uses run ownership for personal failure routing.
> - Users see their own failures and can still inspect company failures
in All.

## Linked Issues or Issue Description

**What happened?**

An agent run started for one user failed. Its row and failure badge
appeared in another user's Mine inbox.

**Expected behavior**

Mine and its badge include failed runs for the current responsible user.
Other users' failures remain available in All and run details.

**Steps to reproduce**

1. Use a company with two human users.
2. Create a failed or timed-out run attributed to the first user.
3. Open Mine as the second user. Before this fix, the failed run appears
there and increases the badge.

**Paperclip version or commit**

Reproduced in regression tests on master at `24beb0057`.

**Deployment mode**

Authenticated deployment with multiple users. Tests also cover the local
single-user board.

Related prior work: #933 addressed inbox dismissal and badge
consistency. No duplicate ownership fix was found.

## What Changed

- Return `responsibleUserId` in normal and summary run lists.
- Share one ownership rule across both inbox versions and client/server
badges.
- Select the latest run per agent before applying the ownership filter.
This prevents old failures from resurfacing on shared agents.
- Keep unattributed historical failures in the local board's Mine view.
Hide them from authenticated users with no matching owner.
- Keep company health alerts outside the personal badge, consistent with
the client.
- Document the routing contract and add page, badge, and database
regression coverage.

## Verification

- Red: the new badge cases failed with three company failures instead of
one personal failure; eight Mine page cases failed across both inbox
versions.
- Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`,
`ui/src/pages/Inbox.test.tsx`,
`server/src/__tests__/heartbeat-list.test.ts`, and
`server/src/__tests__/inbox-dismissals.test.ts`.
- `pnpm check:token-gates` passes.
- Agent calls on behalf of a user have two additional red-to-green API
regressions.
- Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also
passes after the agent-call fix.
- All CI test shards and browser tests pass on `243bfa681`. The
duplicate local `pnpm test:run` was stopped after the CI test lanes
completed; it did not finish locally.

## Risks

- Authenticated users no longer receive unattributed legacy failures in
Mine. Those failures remain visible in All.
- The server badge no longer counts company health alerts, matching the
existing client badge.
- No migration, run state, retry behavior, or company access rules
change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, terminal execution, and
browser tools. The exact deployment variant and context window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:21 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip d172197117 feat(storage): add plain directory sync with conflict preflight (#14416)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent runs use workspace transport to restore and collect files.
> - Some directories belong to the agent across tasks.
> - Those directories need plain file transport without task Git state.
> - A concurrent file edit must be detected before a merge changes any
file.
> - This pull request adds optional plain-directory sync and conflict
preflight.
> - Existing task workspace sync keeps its defaults.

## Linked Issues or Issue Description

Refs #14325. This is the transport prerequisite for a replacement of its
instruction revision design with current agent files.

## What Changed

- Add an opt-in plain-directory transport mode to command and sandbox
runtimes.
- Add strict merge preflight for file edits, deletions, and directory
changes.
- Accept identical replay after an interrupted merge. Preserve competing
changes.
- Set the compiled OpenCode test executable to 0755, independent of the
CI host’s file-creation mask. Preserve the original startup error if
cleanup also fails.

## Verification

- At `ced53ae532ce6966cad5a83575d83db61af98126`, all 212 targeted
transport tests passed across workspace restore, remote managed runtime,
SSH fixture, and execution-target sandbox suites. Adapter-utils
typecheck passed.
- Real isolated SSH retry fixture previously passed with
`PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1`; stale deleted files remain
absent while gitignored binary bytes survive.
- Dependent PR #14420 passed real native local, legacy local, and native
Daytona persistence E2E at `169fab46d5af21caa2269b4c1b29b69c933a6951`,
which includes all production transport changes through `ced53ae53`; the
subsequent two commits only fix the OpenCode test fixture. Nine tasks,
three server restarts, and all cleanup checks passed.
- A hosted OpenCode fixture failed twice at `ced53ae53`. Reproduced the
failure locally and in Linux with `umask 0002`: the compiler created a
group-writable executable, correctly rejected by the qualified launch
boundary. Explicit 0755 permissions fix the test without weakening the
production guard. The focused test and non-root Linux reproduction now
pass under that same mask.
- Before rebase, head `69e97de0475d34aac5d532e559a405eaf015fd2b`
includes the deterministic fixture permission fix and preserves original
bootstrap diagnostics. All production transport code is unchanged since
the 212-test validation. Fresh Greptile review is 5/5 on this exact head
with no unresolved findings. All 54 current-head checks passed, with two
conditional skips. The full CI run completed successfully, including the
previously failing OpenCode runner shard.

- Merge validation on rebased head
`c509d79dd190c5cb00dc65edfde209097ff21465`: all five commits are
patch-identical to the reviewed branch. All 54 checks passed with two
conditional skips, and fresh Greptile review is 5/5 with no findings.
One retry cleared an npm archive 404 and a Cursor fixture timeout.

## Risks

- New behavior is opt-in. Existing task snapshot behavior retains its
defaults.
- Generic strict merge preflight remains opt-in. The dependent
agent-folder feature rebases changed paths before applying them to
provide per-file last-sync-wins; it does not create a conflict-review
queue.
- This change adds no database migration, dependency, or UI.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and tool use
assisted this change. Live provider validation used `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 07:59:11 -05:00
Devin FoleyandPaperclip 53aad90b9e fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path

> - Paperclip coordinates agent work through execution adapters.
> - Sandbox ACP sessions send ordered input through a remote file queue.
> - A temporary provider 502 currently closes the session during an
input upload.
> - A lost response can occur after the sandbox has consumed the
message, so a blind retry can duplicate input.
> - This pull request retries gateway failures with the same sequence
and drops consumed sequences at the receiver.
> - The session can continue through a brief provider failure without
repeating a tool call.

## Linked Issues or Issue Description

**What happened?**

A sandbox ACP run can fail with `ACP agent disconnected during request
(connection_close, exit=null, signal=null)` when a provider input upload
returns HTTP 502. The bridge destroys its local socket on the first
failure and can discard the diagnostic before the proxy reads it.

**Expected behavior**

A temporary gateway failure should get a bounded retry. A lost response
after successful delivery must not duplicate input or reorder later
messages. Permanent failures must still close the session.

**Steps to reproduce**

1. Run the real sandbox process bridge with an echo child and a local
test runner.
2. Inject a provider 502 before preparation, after a chunk upload, or
after final publication and consumption.
3. Send the next input message. Before this change, the connection
closes instead of delivering it.

Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and
`502 sandbox`. Related work: #13287 covers shutdown after bridge loss;
#13793 covers large launch envelopes. This change covers ordered input
delivery within a running legacy ACP session.

## What Changed

- Retry input uploads up to three times for recognized Daytona and
Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms
delays.
- Give each upload separate temporary paths and discard already-consumed
input sequences, including late publication from an earlier attempt.
Clean failed attempts in the background without removing a published
message or another attempt’s files. Cleanup cannot delay retries or
shutdown.
- Keep later input behind the retry. Stop queued input on permanent
failure and flush a fixed diagnostic before closing the socket. Neither
failure-diagnostic persistence nor shutdown-warning persistence can
block teardown.
- Add real-process regression tests for lost responses, late
publication, retry exhaustion, immediate permanent failure, and
diagnostic redaction.
- Give accepted run-log file appends up to three seconds to drain before
finalization computes the size, hash, and durable copy. Close the run
handle to later appends. This waits only for file writes, independently
of later DB progress or live-event persistence. If writes remain
stalled, return null size/hash metadata and skip the final durable copy
so the run can settle. Late writes cannot restart mirroring.
- Preserve legacy comment attribution when final log size is unknown by
reading existing entries within the unchanged 2 MB scan limit. Storage
errors or a three-second read deadline return the evidence already read
instead of failing the comment listing; pagination stops at the
deadline. The deadline requests cancellation of the underlying local
stream or S3 HEAD, GET, and response stream. A separate response timeout
returns partial evidence even when filesystem I/O delays cancellation;
late reads cannot append evidence or start another page. Each listing
retains its existing batches of eight reads, without a shared admission
cap that skips readable logs under contention.
- Document the retry and log-finalization boundaries in the development
guide.

## Verification

- Final commit `347daa564b`: [Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2)
passed. Greptile Apex review 13 scored this commit 5/5 with no new
findings; all 12 review threads are resolved.
- The final CI run initially hit a Cursor test timeout and four Discord
credential-lock contention failures. All five cases passed in isolation.
The two failed shards and their aggregate gate passed on retry without a
code change. Those intermittent failures are not claimed fixed by this
PR.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passed.
- `pnpm exec vitest run
packages/adapter-utils/src/execution-target-stdin-race.test.ts
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests
passed on the final implementation, including 21 new regressions. The
original three fault-injection cases failed before the fix.
- The regressions cover failed and indefinitely stalled cleanup,
Cloudflare gateway responses and retry exhaustion, permanent errors that
must not retry, and teardown while failure logging remains indefinitely
stalled. Seven Apex regression cases failed before the review fixes.
Adapter-utils typecheck and build passed again after the final review
change.
- `pnpm exec vitest run server/src/services/run-log-store.test.ts
server/src/services/run-log-store-cancellation.test.ts`: all 25 tests
passed, including four new regressions that failed before the
finalization fix. They cover delayed and failed appends, late-write
admission, agreement between the local bytes/summary/durable copy, and a
stalled append that exhausts the three-second budget. The timeout case
verifies unknown metadata, no final upload, and no mirror restart after
late completion. New cancellation tests use the real AWS SDK against a
local HTTP server. They verify that stalled HEAD, GET, and response-body
connections close on abort and that a subsequent read succeeds. Local
range and already-aborted read cases also pass.
- `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t
'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14
targeted tests passed. The null-size reader case, both storage-error
cases, the stalled-read case, and the cancellation/concurrent-listing
cases failed before their fixes. The new regressions verify that
timed-out reads are cancelled, subsequent listings recover, and two
concurrent listings both retain their attribution markers. A read that
ignores cancellation still returns partial evidence at three seconds and
cannot resume pagination when it finishes; this regression failed before
the response-timeout fix.
- `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter
@paperclipai/server build` passed after the response-timeout change.
- Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR;
the affected packages were rechecked after review fixes.
- Full local `pnpm test:run` failed in the general-server group: 511
files passed, 40 failed, and 158 were skipped. Failures include embedded
PostgreSQL initialization, read-only cache directory renames, a macOS
long-path fixture, and a workspace exposure assertion. The PostgreSQL,
cache-permission, and long-path failures also reproduce with both
changed implementation files restored to baseline commit `24c58e479a`.
The exposure suite passes in isolation both on baseline and the fixed
branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later
local test groups were not reached.
- An earlier CI run hit the Telegram retry-timing failure fixed upstream
in #14501. The branch includes that master fix. The selected recovery
test passed against a fresh, migrated PostgreSQL 16 database. The
embedded PostgreSQL runner is unavailable on this Mac; the isolated
database was stopped and removed afterward.
- No live agent turn was replayed. The tests use local child processes
and injected provider failures.

## Risks

Retries are restricted to recognized Daytona SDK and Cloudflare bridge
gateway-error messages, which survive plugin RPC serialization. Other
errors fail immediately. Temporary upload paths are now unique for all
command-managed queue writes. Receiver sequence checks prevent duplicate
input; retries do not restart an agent turn. Cleanup and failure logging
are nonblocking and best effort; session teardown remains the final
cleanup boundary. Log finalization now drains accepted local file writes
for at most three seconds and ignores later appends on the closed run
handle. A timeout leaves final size/hash unknown and skips the final
durable upload; an existing partial mirror may remain available, but it
is not claimed as a verified final snapshot. It does not wait for later
DB progress or live-event persistence. Optional attribution keeps
partial evidence when a read fails or times out. Cancellation closes S3
requests and response streams. Local filesystem I/O may finish after the
caller deadline, but a late read cannot change the returned evidence or
continue pagination. Later listings can retry after storage recovers.
There is no schema, authentication, or permission change. Revert this
commit to restore the previous behavior.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and local test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally; targeted tests pass and full-suite
limitations are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 19:15:33 -07:00