mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
b721d24cacd33faab588b389118f5524502e8af7
1519
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b721d24cac |
fix(adapter-utils): retry GitHub broker transport failures before falling back (#14856)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run `git` and `gh` through a managed launcher. The launcher gets a GitHub credential from the Paperclip control plane > - The launcher sends one request to the credential broker for each command > - If that request fails at the transport level, for example after a 10-second timeout, the launcher continues without managed credentials > - So a slow or restarting control plane removes the managed GitHub identity from that command. Some agents then use other GitHub identities that do not have the necessary permissions > - This pull request retries a failed broker request two more times, with a short backoff, before the launcher gives up > - The benefit is that a short control-plane delay does not remove the managed identity from an agent's GitHub operation ## Linked Issues or Issue Description Refs #14175. That pull request changes the same broker request loop for a different failure: sandbox network denials. The pull request that merges second must rebase. **What happened** A Codex agent ran `git` and `gh` through the managed launcher while the control plane was under heavy memory pressure. Each command printed `Paperclip: GitHub broker_transport_unavailable; continuing without managed credentials.` The agent then tried to open the pull request through a different GitHub integration. GitHub rejected the request with `403 Resource not accessible by integration`. **Expected behavior** A short broker delay or a short transport failure must not remove the managed GitHub identity from the command. The launcher must try the broker again before it continues without credentials. **Steps to reproduce** 1. Set `PAPERCLIP_GITHUB_BROKER_URL` to a closed port. 2. Start a broker on that port after about 300 ms. 3. Run `gh` through the launcher. 4. Before this change, the launcher prints `broker_transport_unavailable` and runs `gh` without the managed token. **Version or commit** `4ac374103` on master. Commit `3166e93a7` has the same code. **Deployment mode** Local trusted instance that runs as a launchd service, with `codex_local` agents. ## What Changed - `packages/adapter-utils/src/github-launcher.ts`: the broker request loop now catches transport errors and retries up to two more times, after 0.5 s and then after 1 s. The loop reads the response body inside the retry, so a failed or slow body read is also retried. Busy (409) responses keep their own budget of 30 attempts, separate from transport retries. After the third transport failure, the launcher prints `broker_transport_unavailable` as before. - `packages/adapter-utils/src/github-launcher.test.ts`: two new tests make the broker fail the first request and answer the second. In one, the connection drops before the response. In the other, the connection drops in the middle of the body. Each test checks that `gh` gets the managed token, that the broker receives exactly two requests, and that no `broker_transport_unavailable` message appears. - The existing `broker-offline` test now has a 15-second timeout, because each command now retries twice before it falls back. ## Verification - `npx vitest run packages/adapter-utils/src/github-launcher.test.ts`: 9 of 9 tests pass. - The body-read test fails on the first commit of this pull request and passes with the second commit. - `pnpm --filter @paperclipai/adapter-utils typecheck`: passes. - The existing `broker-offline` test confirms that the launcher still falls back after the retries, and that local Git still works. ## Risks - When the broker is unreachable, each `git` or `gh` command now waits about 1.5 s more before it continues without credentials. When the broker times out, the worst case is about 31.5 s instead of 10 s. - The change only adds retries. It does not change which credentials the launcher accepts or which environment variables it copies. - #14175 changes the same loop. The pull request that merges second needs a small rebase. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Anthropic Claude Opus 5.5 (`claude-opus-5-5`), used through Claude Code with tool use: shell commands, file edits and test runs. The model wrote the change, the test and this description. The repository owner approved the change before it was made. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — *targeted tests and the package typecheck; see Verification* - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — *no documentation describes the broker retry* - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — *CI has not run yet* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — *Greptile has not reviewed yet* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
dd9983b894 |
fix(adapter-utils): release restore locks when a process crashes (#14869)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs restore workspace files and collect instruction-file changes. > - Writers to the same target directory must wait for each other. > - The current lock records a PID, which a new process can reuse after a crash. > - A reused PID can keep an orphaned lock alive and make each later run fail. > - This pull request makes a SQLite file lock decide ownership. The OS releases it when the process exits. > - Later runs can proceed after a crash, and concurrent live writers remain protected. ## Linked Issues or Issue Description Refs #10914. This addresses crash recovery. It does not cancel a stalled operation in a process that is still alive. Related work: #9667, #14787, and #12187. The earlier attempt in #9667 assumes one live server per lock root. This implementation uses an OS-backed lock to support concurrent writers without treating a different process token or an old timestamp as proof of a dead owner. It retains the private lock root and bounded timeout diagnostics from the merged changes. After a process dies while holding a restore lock, a replacement process can reuse its PID. The existing `process.kill(pid, 0)` check then reports a live owner forever. Later runs can complete their model turn but fail during file collection or restore. ## What Changed - Hold a SQLite `BEGIN IMMEDIATE` transaction for each directory write. Use the existing built-in `node:sqlite` dependency. - Keep each lock database on a stable inode. Keep PID and time metadata only for diagnostics. - Retain the 30-second asynchronous wait and existing timeout error code and diagnostic fields. - Fail closed when an old directory lock exists. Document a stopped-writer upgrade and rollback procedure. - Add real child-process tests for crashes, PID reuse, live owners, and connection cleanup. Cover callback failures, independent targets, stable inodes, invalid lock files, and ambiguous legacy records. ## Verification - Before the fix, the crash/PID-reuse test and the live-owner test both failed. Both pass with this change. - `pnpm exec vitest run packages/adapter-utils/src/directory-merge-lock.test.ts packages/adapter-utils/src/workspace-restore-merge.test.ts`: 56 tests passed. - Restore and agent-file working-copy integration tests: 118 tests passed before the additional connection-cleanup test. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Full GitHub CI: all checks passed, including Linux workspace tests, server test shards, build, typecheck, and browser tests. - Greptile: 5/5, with no review threads or unresolved comments. - `pnpm test:run`: started locally, then stopped with SIGINT (exit 130) after full CI passed. The local serial run did not complete and is not counted as a full local pass. The completed CI shards provide the full-suite result. ## Risks - **Upgrade and rollback require a drain.** Stop every old writer that shares an instance root before switching protocols. Old and new versions must not write concurrently. - Existing legacy `.lock/` directories remain blocking. After all writers stop, preserve run evidence and move those directories to an operator scratch directory. The new code does not infer that they are abandoned from PID or age. - Never delete or replace a `.lock.sqlite` file while writers can run. These small files remain after release. - The shared filesystem must support reliable SQLite locking. Broken network-filesystem locking is unsupported. - This change prevents new orphaned ownership. It does not recover file changes lost during earlier failed collections, or interrupt a live operation that stalls. - No application database migration or new native dependency is required. See `doc/workspace-restore-locks.md` for the procedure. ## Model Used OpenAI Codex based on GPT-6, with code execution and repository tools. The exact model variant and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused regression and integration suites; see the full-suite note above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ac374103f |
fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.
## Linked Issues or Issue Description
Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.
**What happened?**
Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.
**Expected behavior**
Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.
**Steps to reproduce**
1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.
**Paperclip version or commit**
Reproduced from
|
||
|
|
467125fafb |
feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen
## Linked Issues or Issue Description
No public issue exists. This is the description, from the enhancement
template.
**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.
**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.
**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.
**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.
**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.
**Breaking changes**
None. No schema or API change. Existing connections keep their settings.
## What Changed
- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.
## Verification
- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
- `action-permissions.test.ts` checks the group update.
- `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".
## Risks
- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.
> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.
## Model Used
- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
0d3e7bf6ac |
fix(daytona): keep commands alive after log stream closure (#14799)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox providers run agent processes and send their output to the host. > - The Daytona SDK can close a log socket while the remote command still runs. > - The driver treated a clean socket close as completion before it had a command exit code. > - This pull request recovers log observation for the same command and waits for a recorded exit. > - The host keeps receiving new output without a second command dispatch. ## Linked Issues or Issue Description **What happened?** A clean close of the Daytona session log WebSocket resolves the SDK callback promise. If the command still runs, the driver returned `exitCode: null` with `timedOut: false` after a short status check. A streamed ACP bridge can then report a process disconnect. **Expected behavior** A log socket close must not complete a running command. Recovery must preserve new output, the caller's lifetime controls, and one command dispatch. **Steps to reproduce** Use the callback form of `getSessionCommandLogs`. Let that promise resolve while `getSessionCommand` still has no exit code. Keep the command running, then expose its final logs and exit code. The new regressions exercise this sequence, including streams that run longer than the provider operation timeout. Related: #11021, #11049. #14485 covers input delivery retries, which are a separate transport path. ## What Changed - Require a recorded command exit after a clean log-stream close. Reconnect once, then read status and full log snapshots at most once per second. - Forward new output during recovery. Remove replayed prefixes and reconcile a final snapshot so bytes written after socket close are retained. - Hold a trailing UTF-8 replacement suffix until replay or completion resolves it. This handles the SDK decoder flush when a socket closes in the middle of a character. - Preserve healthy initial and reconnected stream lifetimes. Bound each recovery read. Preserve the existing fallback timeout budget after rejected stream attempts. - Retain partial output on timeout. A log-observation timeout reports an unconfirmed result and preserves any observed exit code in metadata; it does not synthesize successful completion. ## Verification - `pnpm exec vitest run --config packages/plugins/sandbox-providers/daytona/vitest.config.ts`: 335 passed, 14 gated tests skipped. - `pnpm exec vitest run --project @paperclipai/plugin-daytona`: 335 passed, 14 gated tests skipped. - `pnpm exec tsc --noEmit -p packages/plugins/sandbox-providers/daytona/tsconfig.json`: passed. - `pnpm exec tsc -p packages/plugins/sandbox-providers/daytona/tsconfig.json`: passed. - `pnpm --workspace-concurrency=1 -r typecheck`: passed. - `CARGO_BUILD_JOBS=2 pnpm --workspace-concurrency=1 -r build`: passed. - `pnpm test:run`: exited with failure after 845.77 seconds. The server phase reported 43 failed files, 529 passed, and 163 skipped; 14 failed tests, 9,122 passed, and 5,599 skipped. All failure entries were traced to local PostgreSQL startup/cleanup errors or ten macOS skill-cache rename errors. The package helper restored 17 missing PostgreSQL library links; the four sequencing/migration tests and fourteen native-workspace-finalizer tests then passed. The ten cache failures match clean-base evidence with identical source/test blobs. This is not a full-suite pass; later local test groups did not run. PR CI supplies the complete check result. - Regressions cover clean and rejected stream recovery, two hour-long streams with a five-minute operation budget, live fallback output, one dispatch, delayed final output, stale snapshots, SDK UTF-8 decoding, bounded observation, and timer cleanup. - `git diff --check` and a local diff scan for secrets and private identifiers passed. - Greptile reviewed head `293c4dbe67` at 5/5. Both prior findings are fixed, both threads are resolved, and no new actionable findings remain. ## Risks - The SDK returns full snapshots with no offset API. After both stream attempts end, polling bandwidth grows with retained output. The one-second cadence limits request frequency. - The SDK callback stream has no cancellation handle. Existing caller stop logic and provider/session teardown still own its lifetime. Late callbacks from a settled stream are ignored. - A successful status read with no exit code keeps recovery active under the existing caller guard. A failed read stops recovery; it is not retried indefinitely. - No command replay, provider API change, schema migration, or runtime timeout policy change is included. ## Model Used OpenAI GPT-6 through Codex, with tool use and independent code review. The exact serving model identifier is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b9a6000f7 |
Add bounded evidence for directory lock timeouts (#14787)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent files use directory locks during collection and cleanup. > - A lock timeout can fail finalization after the model turn completes. > - The timeout currently identifies no owner state or waiting operation. > - This pull request adds bounded evidence to the existing run failure report. > - Operators can distinguish a known local holder from a possible old lock without changing lock safety. ## Linked Issues or Issue Description **What happened?** A directory lock timeout does not distinguish active local work from an owner record left by an earlier process. The stored execution stage can also precede the cleanup operation that failed. **Expected behavior** The failure report should identify the waiting operation and expose bounded ownership clues. It must preserve the timeout and keep unknown ownership protected. **Steps to reproduce** Hold a directory merge lock while a second caller reaches its acquisition deadline. The regression tests exercise a live holder and an older owner record with a live PID. Related: #9667 proposes stale-lock recovery under a single-server assumption. This change only adds evidence and does not adopt that assumption. #14575 and #14665 add other run failure diagnostics. ## What Changed - Record lock owner state, capped age and wait duration, same-process and process-age comparisons, and whether this module holds the lock. - Label agent-directory release, collection, checkpoint, and warm handoff timeouts with a fixed operation code. - Validate each field before the existing event-local Sentry report accepts it. Exclude owner records, PIDs, paths, and absolute timestamps. - Limit the extra diagnostic owner read to 100 ms with best-effort abort; malformed JSON is `invalid` and unreadable owner records remain `unknown`. - Document the diagnostic limits and verify that contenders never reclaim protected locks. ## Verification - Focused lock, diagnostic, real Sentry SDK, and database-backed agent-directory tests: 126 passed, including stalled-read and malformed/missing/unreadable-owner regression coverage. - Final revision `0691613dcc`: all 54 reported checks successful, with two intentionally skipped Storybook checks. Greptile: 5/5, zero unresolved review threads; no merge conflicts. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm test:run`: complete suite coverage ran with the existing repository shard flags: four general-server shards, four serialized shards, two general-workspaces-a shards, and general-workspaces-b. The full run is not green because of the base failures below. - The broad run found 13 failures in the unchanged macOS skill-cache tests. All 13 reproduce on the clean base revision. Open PR #14290 covers that existing failure. - Two unchanged CLI archive tests hit their five-second limits during the broad run; all 17 tests in that file pass on recheck. A CLI auth socket error also cleared on recheck (19 tests), and its full serialized shard passed on rerun. ## Risks This is a diagnostic change, not a stale-lock fix. Owner observations can race with release. Wall-clock shifts can affect the age comparison. A local-holder flag covers only this module instance. None of these fields authorizes reclamation or proves a file save. Lock acquisition, release, retries, task status, and recovery guards retain their current behavior. No schema change or deployment action is required. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact model build and context window were not exposed to this agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass (focused checks pass; existing base failures are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
98d8a6ccac |
Stop replaying ambiguous database disconnects (#14773)
## Thinking Path > - Paperclip stores agent work and control state in PostgreSQL. > - Its database client must not repeat a mutation after an uncertain result. > - The global retry wrapper treated `write CONNECTION_CLOSED` as proof that PostgreSQL never received a statement. > - postgres.js also uses that message when the connection closes after statement delivery. > - This pull request removes that global replay and tests the actual driver over a local wire connection. > - Callers retain control of retries when they can prove the complete operation is idempotent. ## Linked Issues or Issue Description Follow-up to #13417. Preserve the transaction disconnect handling from #13643 and the explicit actor synchronization retries introduced in #12773. Searched open and closed issues and PRs for database retries, disconnects, and `CONNECTION_CLOSED`. The open circuit-breaker proposal #11142 addresses outage queue growth; it does not establish whether an already-sent statement can be replayed. **What happened?** The database wrapper replayed an arbitrary statement up to three times after `write CONNECTION_CLOSED`. The driver adds `write ` to connection-close errors even after the peer receives the statement. A local protocol peer receives the same submitted INSERT three times when it drops each response. A committed write could therefore execute more than once. **Expected behavior** An ambiguous statement result must fail without automatic replay. A subsequent operation must be able to reconnect. **Steps to reproduce** Run the new wire regression against the parent commit. The six Simple Query cases and the parameterized Drizzle case receive three executions instead of one. The named prepared-client case was already safe and stays covered. The peer reads the entire statement and then closes the connection. This demonstrates repeated delivery with the real driver; it does not claim that a historical incident duplicated a committed write. **Paperclip version or commit** Reproduced on source commit `018993140f` with the patched postgres.js 3.4.9 dependency. **Deployment mode** Built from source with a local PostgreSQL protocol peer. No live provider or customer database is used. ## What Changed - Pass the original postgres.js client to Drizzle and remove the global statement replay wrapper. - Add eight wire regressions: six Simple Query cases for INSERT, side-effect-capable SELECT, and a data-changing CTE, plus parameterized Drizzle and named prepared-client cases. The extended peer processes Parse, Describe, Bind, and Execute, verifies bound parameters, and drops the response only after Execute. Each case checks one delivery and recovery on a fresh query. - Document ambiguous outcomes and the retry compatibility tradeoff. Keep explicit idempotent actor-sync retries and disconnected-transaction handling unchanged. ## Verification - Before the fix: the six Simple Query cases and the parameterized Drizzle case failed with three executions instead of one. The named prepared-client case was already safe. All eight wire cases pass on this branch. - Final focused client, pool teardown, configuration, and actor-sync retry checks: 28 tests passed. `pnpm --filter @paperclipai/db typecheck` also passed after the test-only follow-up. - First implementation head, `pnpm exec vitest run --project @paperclipai/db`: all 158 tests passed across 45 files, including real PostgreSQL transaction/reserved-connection recovery. The local embedded dependency's symlinks were hydrated before this run. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - The complete local `pnpm test:run` did not finish; no complete local suite pass is claimed. All CI test, typecheck, and build gates passed on the first implementation head `11f8b22d90`. Final-head CI is pending after the test-only follow-up. - `git diff --check`: passed. Reviewed the diff for secrets, personal data, generated output, and run artifacts. ## Risks Some transient statement failures that the global wrapper previously replayed now reach the caller. Operation owners must retry only when they have an idempotency guarantee or a durable receipt that prevents duplicate effects. A connection error is not proof that a write failed to commit. There is no SQL-text retry heuristic, new suppression, schema change, or migration. This change prevents unsafe replay; it does not prevent network disconnects. ## Model Used OpenAI Codex / GPT-6, with reasoning, repository inspection, code execution, and local protocol tests. The exact backend model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Final verification (September30): every final-head CI check passed at `3004c5bda39c985c3557547ec45e33870ce5d010`. Greptile scored5/5 on this head, all review threads are resolved, and the branch is mergeable. Eight real-wire regressions cover simple, parameterized Drizzle, and named prepared queries. Full local suite did not produce a completed result; the complete CI matrix passed. This public PR remains open for maintainer merge. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
33f2b3a159 |
fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Connectors catalog lets people give agents tools or connect agents to conversations. > - GitHub put these two uses behind one card and an extra choice. > - People should choose the connection they need from the catalog. > - This pull request keeps GitHub for tools and adds GitHub Code Review Bot as a separate card. > - Each card opens its setup directly. Both use the existing connection code. ## Linked Issues or Issue Description **What existing behavior does this improve?** GitHub connector discovery and setup. **Current behavior** With chat connectors enabled, GitHub opens a menu that asks whether to use tools or create a bot. Saved tools and bots share the same catalog entry. **Proposed behavior** GitHub opens tool account access. GitHub Code Review Bot opens agent selection. Saved bots and drafts appear under the bot card. Chat-disabled instances show only GitHub tools. **Reason and benefit** The catalog names the two uses and removes an extra setup choice. The bot keeps the existing GitHub provider, credentials, endpoint IDs, setup steps, and runtime. **Additional context** Related work: https://github.com/paperclipai/paperclip/pull/12843 and https://github.com/paperclipai/paperclip/pull/14594 established GitHub account identity. This change preserves that tool flow. No duplicate catalog split was found. ## What Changed - Split the generated app definitions into GitHub tools and GitHub Code Review Bot. Reuse the existing GitHub logo and channel method. - Open bot setup directly, including old resume and reconnect links. - Put existing bot endpoints and drafts under the bot card. Hide duplicate internal chat applications. - Keep pasted GitHub URLs mapped to the tool connection. - Add seven Storybook states for the catalog, saved connections, disabled chat, both setup paths, mobile, and light mode. - Fix narrow-screen bot rows so the label cannot overlap status and setup actions. - Update catalog, route, browser, and API tests, plus the GitHub connector guide. ## Verification - [Hosted Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog): seven states built from this branch. The deployment passed its public-file verification. - All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`. Two optional Storybook jobs skip under their normal trigger rules; the manual Storybook deployment passes. The branch has no merge conflicts. - Greptile: 5/5 on the current head, with no review comments or unresolved threads. - `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm build-storybook` passed. The final Storybook fixture also passed UI typecheck and the hosted build. - Targeted catalog, URL matching, routing, grouping, brand, and chat UI contract tests passed. - GitHub provider browser tests: 2 passed. These cover direct tool setup and the bot setup and management lifecycle with provider responses mocked. - Embedded-browser test on an isolated local instance: opened both cards, selected an agent, saved a bot draft, and resumed the same endpoint under the bot card after a reload. - Storybook Tool Setup and Bot Setup assertions pass in the published preview. Chat Disabled assertions pass locally. Inspected mobile and light mode, including the draft-row layout and official GitHub marks. - Local full-suite limitation: `pnpm test:run` was not clean. A cross-company route assertion failed in the aggregate run and passed in isolation; a workspace-runtime test reached its 30-second hook timeout. Some isolated database reruns skipped when the embedded-PostgreSQL availability probe failed. The local aggregate was stopped after CI completed. The corresponding full CI suites pass all 360 tool-access tests and all 162 workspace-runtime tests. - No live GitHub authorization or installation was performed. The isolated instance correctly stopped at the cloud enrollment or public HTTPS prerequisites. ## Risks - Low scope: catalog presentation and routing change. There is no database migration or provider credential change. - Existing GitHub bot URLs now open bot setup directly. The tool route remains `/apps/connect?source=github`. - The bot remains behind the existing chat-connectors feature flag. Existing endpoints retain `provider: github`. - Channel applications are represented by endpoint rows. Regression tests cover legacy bot applications, tools, active bots, and drafts together. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and embedded-browser tools. The exact deployed model ID, context window size, and reasoning setting are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
018993140f |
feat: let agents name prompt-only tasks (#14761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d6d67b00d3 |
Prevent background workspace scans from refreshing the Git index (#14666)
## Thinking Path Paperclip runs background workspace scans alongside real Git writers. `git status` can refresh the index as an optional side effect, taking a lock that makes another operation fail. Disable optional locking in the shared scan subprocess so background observation does not compete with workspace updates. ## Linked Issues or Issue Description **What existing behavior does this improve?** Workspace Git scans used by changed-file browsing, cleanliness guards, and sandbox snapshots. **Current behavior** The scan process inherits Git's default optional-lock behavior. Even a clean `status` can rewrite stale stat-cache entries in the index and contend with a concurrent writer. **Proposed behavior** Always set `GIT_OPTIONAL_LOCKS=0` for the shared scan subprocess while preserving the selected environment and Git's required write locks. **Reason and benefit** Background reads stop creating avoidable index contention. Git documents this behavior and recommends disabling optional locks for background status: [background refresh](https://git-scm.com/docs/git-status#_background_refresh). **Breaking changes** None to scan results or required write locking. Later scans may repeat stat checks that would otherwise have been cached in the index. Related scan implementation: #11572, #14253. This avoids one known contention source; it does not identify every historical lock owner or repair abandoned locks. ## What Changed - Disable optional locking at the shared scan subprocess boundary, including explicit caller environments. - Test a clean status against a deliberately stale index and prove an ordinary status would rewrite it. - Test tracked/untracked results with an existing index lock, preservation of that lock and working files, and continued rejection of a mandatory-lock write. - Document the scan behavior and performance tradeoff. ## Verification - Focused stream, workspace-sync, and scheduler suites: 63 tests passed. - Full `pnpm -r typecheck` and `pnpm build` passed locally. Final head `effe6420c77bd18d36af6b093db3a564c04b8b38` passed all 53 CI checks, including complete test coverage, typecheck, build, and browser/runner gates; two unrelated checks intentionally skipped. - Full local test attempts initially had missing embedded-Postgres library symlinks; the dependency setup was repaired. Duplicate local full-suite runs were stopped after full CI completed. This PR does not claim a completed full local suite. - Review regression: real Git honors the supplied `GIT_CONFIG_*` setting and the input environment remains unchanged; all four direct subprocess cases passed. - Reviewed the diff for secrets, customer data, and internal references. ## Risks Low risk. Disabling optional index refresh can repeat filesystem stat work on later scans. Required locks remain enforced; no lock is removed, no failed reset is retried, and workspace mutation guards are unchanged. No schema changes. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code execution, and tests. Exact model build identifier is not exposed by this session. ## Checklist - [x] Thinking path and model are specified - [x] Checked ROADMAP.md; this is a maintenance correction, not planned feature work - [x] Searched for duplicate and related PRs - [x] Described the issue using the enhancement template - [x] No internal issue references, customer data, or private instance links - [x] Descriptive branch name - [x] Focused regression tests pass - [x] Added tests and updated documentation - [x] Risks documented - [x] Required validation and CI gates are green (full suite validated in CI; local scope documented above) - [x] Greptile is 5/5 with no unresolved findings - [x] I will address review comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0ea6b10967 |
Record ACP activity and workspace restore failure evidence (#14665)
## Thinking Path Paperclip records terminal run failures for operators. A timeout's own log and cleanup output update the run's last-output timestamp, so that timestamp can make a long-silent provider look active. Snapshot runtime activity before finalization and include the saved workspace restore classification to make the next failure actionable without copying tool payloads. ## Linked Issues or Issue Description **What existing behavior does this improve?** Terminal run diagnostics in the existing opt-in Sentry integration. **Current behavior** Reports cannot distinguish runtime events from finalization logging and omit the already-persisted workspace restore code. A later successful run also does not establish that earlier workspace files were restored. **Proposed behavior** Record runtime-event age/count and pending-tool inventory at finalization, before status reads and cleanup. Forward only finite counts, a completeness boolean, and known restore codes through the existing reporter. **Reason and benefit** Operators can distinguish a silent turn with unfinished tools from recent runtime activity and see restore failures without retrieving private run output. Neither signal certifies productive work or successful recovery. **Breaking changes** None. Error grouping, execution deadlines, cancellation, recovery policy, and the Sentry opt-in remain unchanged. Related diagnostic work: #14573, #14575, #14639. ## What Changed - Snapshot ACP activity before success/failure finalization, including thrown relay failures. - Forward bounded numeric/boolean evidence and shared workspace restore codes; exclude commands, tool IDs, paths, and arbitrary result data. - Document limitations and test silence, empty streams, timeout, cleanup delay, incomplete tool inventory, and privacy. ## Verification - `pnpm -r typecheck` passed after the final implementation. - Changed suites: 252 tests passed; all 29 database reporter tests subsequently passed after restoring the embedded-Postgres package library symlinks. The migration test also passed (30 database cases total). - `pnpm build` passed during implementation. Final head `fe78dba6f592b1abccac7cdbf341bd2e0b0d30cb` passed all 53 CI checks, including complete test coverage, typecheck, build, and browser/runner gates; two unrelated checks intentionally skipped. - Full local test attempts initially hit missing embedded-Postgres library symlinks; the dependency setup was repaired and database tests passed. Duplicate local full-suite runs were stopped after full CI completed. This PR does not claim a completed full local suite. - Review regression: completed, failed, and cancelled tools are excluded from the pending count; focused activity/timeout tests and adapter-utils typecheck passed. - Reviewed the diff for secrets, customer data, and internal references. ## Risks Low risk, diagnostic-only. The existing tool inventory is incomplete for some runtime events, so the report carries its completeness flag. Event age is measured at finalization and does not prove useful work or identify the underlying provider failure. No schema changes or new capture gate. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code execution, and tests. Exact model build identifier is not exposed by this session. ## Checklist - [x] Thinking path and model are specified - [x] Checked ROADMAP.md; this is a maintenance correction, not planned feature work - [x] Searched for duplicate and related PRs - [x] Described the issue using the enhancement template - [x] No internal issue references, customer data, or private instance links - [x] Descriptive branch name - [x] Focused regression tests pass - [x] Added tests and updated documentation - [x] Risks documented - [x] Required validation and CI gates are green (full suite validated in CI; local scope documented above) - [x] Greptile is 5/5 with no unresolved findings - [x] I will address review comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad55d0a281 |
fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use Apps through a gateway that checks identity, company access, and action policies. > - Personal pasted credentials can point to company secrets. Setup can show success while the gateway rejects every call. > - Several OAuth methods also omit the scopes needed for their supported write actions. > - This pull request gives setup, health checks, and invocation the same credential rules. Owners repair existing connections by reconnecting. > - New connections request reviewed permissions for their supported actions. Read-only choices remain available under Advanced. > - Agents can use the connections people give them, while existing consent, identity boundaries, and action restrictions remain enforced. ## Linked Issues or Issue Description Refs #14009 and #14008. This addresses the personal-credential defect. The separate GitHub organization-identity selection defect is outside this change. Related work: #13942 fixed part of new personal-key setup. #14200 independently fixes legacy personal reconnect and protects managed-agent profile credentials during removal. This PR covers that ownership invariant across key and secret-URL setup, reconnect, health, discovery, and invocation, and keeps owner reconnect as the repair path. #14059 tracks requested versus provider-asserted OAuth scopes; it remains separate work. I searched open PRs and issues for Zapier, Airtable scopes, connector writes, and personal credential failures. **What happened?** A Zapier secret URL saved through personal setup can become a company secret referenced by a user grant. Health checks bypass the gateway's ownership check, so the connection appears healthy but calls fail with `grant_credential_invalid`. Custom-header paths can also receive a duplicate `credentials.` prefix. Omitted OAuth scopes make write access depend on provider defaults. **Expected behavior** Personal invocation credentials belong to the selected user. Setup, health, and actual calls enforce the same rule. New connections request documented permissions for supported read and write actions. Existing tokens gain no permissions without provider consent. **Steps to reproduce** 1. Connect Zapier or a generic secret URL with the personal identity. 2. Allow an agent to use the connection and complete setup. 3. Invoke a tool through a run-scoped gateway. The legacy layout fails ownership validation despite successful setup. **Paperclip version or commit** The implementation started from `44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master at `94e8dec56`. **Deployment mode** Built from source. Regression tests use isolated PostgreSQL fixtures and controlled MCP transports. ## What Changed - Share credential writing, ownership validation, and canonical paths across initial setup, resume, reconnect, rotation, health, discovery, and gateway calls. Keep OAuth client-registration secrets separate from invocation credentials. - Existing personal connections with company-scoped credentials require owner reconnect with a fresh key or secret URL. Reconnect creates a correctly owned value and updates the existing grant and declarations. There is no automatic ownership backfill or new startup hook. - Preserve PostgreSQL timestamp precision when reconnect checks whether a grant changed. Previously, converting the timestamp to a JavaScript Date could reject reconnect with a false concurrent-change error. - Protect credentials used by other grants, connections, bindings, managed-agent profiles, routine triggers, or secret proposals from connection removal. - Review all 117 tool methods, including 84 OAuth methods. Record explicit scopes or documented provider-default exceptions with official evidence. Add Airtable's seven scopes, Hugging Face repository/job scopes, and other documented MCP permissions. - Prefer available write-capable methods. Put explicit read-only choices under Advanced. Explain pasted-key permissions and offer reconnect for missing OAuth consent. Preserve existing grants, policies, Google availability gates, and curated scope allowlists. - Reconnect generic secret URLs and custom headers using their stored credential fields. Refresh the catalog after setup, correct reconnect feedback and error guidance, and let Cancel exit invalid setup while Save & exit retains draft-saving behavior. - Apply ownership checks to the new GitHub repository/skill connection picker. Align the permission audit with the Google scope reductions merged on master. - Add run-scoped gateway, ownership, owner-reconnect, OAuth URL, insufficient-scope, UI, and catalog-wide regression coverage. Update the connector playbook and permission audit. ## Verification Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all CI/status gates (55 completed check runs, no failures or pending checks) and has a completed Greptile review at **5/5 with no outstanding findings**. GitHub reports the PR as mergeable/CLEAN. - **Embedded browser:** used the actual server and built UI from this worktree, a fresh isolated database, and local HTTP MCP fixtures. Completed personal bearer-key, secret-URL, and custom-header setup; reproduced the legacy ownership failure; reconnected through the owner’s form; and completed writes afterward. Read-back was verified for bearer-key and secret-URL connections. Public organization-wide setup appeared immediately in Browse without reload. Zapier URL validation/Cancel and Google’s enrollment gate were also exercised. - **Persistence and invocation:** verified user ownership, canonical `credentials.authorization` / `remote.url` / `headers.X-Api-Key` declarations, and unchanged connection/grant identity. The old company secrets retain their ownership. Separate HTTP calls through an actual run-scoped gateway session completed a write and read-back. - **Backend coverage:** the final gateway suite passes all 82 cases, including catalog Zapier and generic inline reconnect. It checks company/user isolation, canonical declarations, same-endpoint URL validation, fresh credentials, retained restrictions, and real gateway read/write execution using fixture transport. A timestamp with PostgreSQL microseconds covers the former false reconnect conflict. - **Local checks:** 368 catalog, gateway, repository, and UI tests passed before the final extra Zapier case; 49 GitHub skill access tests also passed. All three Apps browser regressions pass, including reconnect through the actual form and catalog visibility without reload. Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final patch, and token gates passed. Full tool-access service runs hit varying 15-second Google fixture timeouts; both affected cases and the updated reconnect assertion pass in isolation (3 tests). The complete test matrix passes in CI on this head. - **Verification limits:** no live provider account was available for Zapier/Airtable/OAuth consent or account-bound write proof. Public metadata and local fixtures do not establish provider consent. The original development database clone failed on a pre-existing missing `tool_connections_transport_check` constraint; browser acceptance used a fresh isolated database created by the normal CLI onboarding flow. ## Risks - Existing broken personal connections stay unusable until their owner reconnects. Health, discovery, and invocation return an actionable ownership error; startup does not rewrite credential ownership. - Scope changes affect new authorization requests. Providers may still require resource selection, account roles, paid plans, or app verification. Existing consent and action restrictions remain unchanged. - Shared credentials are retained rather than reassigned or revoked. Provider-default exceptions and unavailable live checks are documented in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`. - No new endpoint, database table, lockfile change, or CI workflow change is included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell execution, web research, and browser tools. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbd278dc03 |
fix(interactions): derive chat recipients and validate explicit users (#14742)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agents use saved questions to get human input and continue the same task. > - The standard question example recently told models to copy a user ID. > - A model can omit an identity prefix and create a question its intended recipient cannot answer. > - Agent Chat already knows the conversation owner, so the server can supply that identity. > - This pull request removes the blanket instruction and validates explicit recipients before saving. > - Ordinary questions stay simple, and explicit addressing remains available for decisions that need a particular person. ## Linked Issues or Issue Description Refs #14707, #14188. Related: #14238 handles legacy email recipients; this change prevents invalid recipients in new cards and retains exact ID matching. **What happened?** A model copied a Cloud user ID without its prefix into `addresseeUserId`. Creation succeeded. The intended user's answer then failed the exact recipient check. **Expected behavior** Ordinary chat questions use the saved conversation owner. A task may optionally name a specific recipient. The API rejects an unknown or unauthorized recipient before it creates a card. **Steps to reproduce** Create a chat question for a user whose ID is `paperclip-id:example`. Supply `example` as the addressee. Before this change, creation accepts the invalid recipient and the owner cannot answer. With this change, creation returns 422. Omitting the field saves the full owner ID and allows that owner to answer. ## What Changed - Remove `addresseeUserId` from standard question examples and remove the blanket requester-ID instruction. - Derive the recipient of ordinary chat questions from the persisted conversation owner. Reject conflicting explicit user IDs. - Keep explicit task recipients optional. Validate supplied user IDs with the existing board mutation policy, including company, viewer, and Cloud restrictions. - Preserve explicit agent routing, connector intents, confirmations, exact recipient checks, idempotent retries, and no-login local-board authority in local-trusted mode. - Update the blocker grader to accept an omitted recipient and verify the actual requester answered. - Add database and HTTP tests for prefixed identities, denied recipients, concurrent retries, saved answers, and response delivery. ## Verification - Database interaction service suite: 90 tests passed, including implicit local-board creation/answering and authenticated/Cloud denial. - Interaction HTTP route suite: 84 tests passed. - Affected interaction/native/connector/documentation suites: 231 tests passed across six files after valid-user fixtures were updated. - Resolver and interaction unit suites: 29 tests passed. - Product E2E unit/calibration suite: 793 tests passed; Product E2E typecheck and blocker catalog discovery passed. - Generated API-reference and capability contract checks passed. - `pnpm -r typecheck` and `pnpm build` passed. - Full local `pnpm test:run` did not finish green: its initial general-server pass had 14,416 passing assertions, one unrelated native-resume assertion failure on macOS, and three teardowns from an intermediate fixture cleanup fixed above. Separate broad local groups also encountered timeout/live-port failures under host load. Local UI (7,026), CLI (502), shared (817), and skills-catalog (20) tests passed; the complete final-head CI matrix is the broad verification gate. - After two CI cold-start readiness timeouts, a separate test-only commit gives the first exposure lifecycle fixture the existing normal 30-second readiness budget. Its real HTTP, ordering, and cleanup assertions remain intact; the targeted case and final Linux CI shard passed. Production deadlines are unchanged. - A separate OpenCode fixture failed twice on GitHub-hosted Ubuntu because its cached Node executable was group-writable; the same case passed on AWS runners. The fixture now qualifies its own Linux copy with mode `0500` and the actual copy digest. Host files and production security checks are unchanged. The focused macOS case passed; the new Linux-copy branch also passed on the final AWS-hosted Linux runner (1,125 passing Runner tests, 3 skipped). The final run was not on a GitHub-hosted runner. - Final-head [CI run 36762078176](https://github.com/paperclipai/paperclip/actions/runs/36762078176) passed for `116b968b24fa0a8c5724a7bf96e73a8dda5f0425`: 54 successful checks and two conditional Storybook skips, with no pending or failed checks. The 27 general/serialized test jobs reported 28,635 passing tests. Typecheck, build, Runner, browser E2E, and Canary gates passed. Greptile reviewed that exact head at 5/5; both review threads are resolved, with no open follow-ups. - No live provider replay is claimed by this PR. ## Risks - New explicitly addressed cards reject users who cannot mutate the issue, including viewers, inactive members, and invalid IDs. Callers that supplied invalid recipients must correct their request. - Existing addressed cards are not rewritten. Existing authorization checks remain strict. - Chat inference applies only to questions without an agent addressee. Connector intents and governed confirmations retain their own recipient paths. - No schema change or migration is required. ## Model Used OpenAI Codex, GPT-6 (exact serving variant and context window are not exposed in this environment). Used reasoning, tool use, code editing, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b54b2dc35c |
fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path > - Paperclip manages AI agents and keeps their instructions and files durable. > - Native Codex runners can keep a process alive between compatible turns. > - Managed file collection stopped that process after each turn, which defeated warm reuse. > - Agent folders can contain large images and other files, so full copies on every turn are expensive. > - This change keeps one managed directory for the live session and saves only file changes after each turn. > - Ownership, authorization, instruction changes, and process retirement still control when reuse is safe. ## Linked Issues or Issue Description Related: #13710 introduced native warm session reuse. This fixes managed file collection that still forced those sessions to stop. No duplicate open PR or issue was found. **What happened?** With managed instructions and warm native Codex enabled, consecutive turns reused a Daytona sandbox but started a new runner process each time. The managed directory collector required process termination before saving files. **Expected behavior** Compatible turns keep the same process and managed `AGENT_HOME`. Each completed turn saves added, changed, and deleted files before the next turn starts. Unchanged large files do not transfer again. **Steps to reproduce** 1. Use a native Codex agent with managed instructions and a reusable Daytona environment. 2. Enable warm session reuse and run three turns on the same task. 3. Write a large binary on the first turn, edit a small note on each turn, and delete a file on the second turn. 4. Compare process identity across turns and read the canonical files through the public agent-files API. **Paperclip version or commit** Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`. **Deployment mode** Local server with remote Daytona execution; cloud native runner uses the same path. ## What Changed - Retain the managed directory only for the verified owner of a live native Codex session. - Checkpoint each completed turn before releasing the session for reuse. Retry unstable captures, then stop and collect when a warm checkpoint cannot be validated. - Compare metadata and cached hashes, stream only changed file payloads, record deletions, and validate path, content, quota, and authorization before saving. - Rotate sessions when canonical files, loaded instructions, credentials, or launch policy change. Fence stale collection and cleanup callbacks from later owners. - Keep cleanup and recovery aware of the current session owner. Recheck canonical files under the writer lock at handoff, attach the successor collector before fallible bookkeeping, and emit one final save receipt on checkpoint fallback. Preserve storage warnings across unchanged checkpoints. - Add regression coverage and a three-turn Daytona test with independent public API file checks, an unchanged 8 MiB binary, deletion checks, and strict process identity checks. - Document checkpoint consistency, lifecycle behavior, and local run-log counters. - Replace a timing assumption in the Daytona teardown test with explicit transfer-arrival gates after CI exposed an unset release callback. ## Verification - Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks were repeated after the final storage-warning fix. - Runner E2E typecheck and 749 runner E2E unit tests passed. - Focused file checkpoint, directory ownership, instruction collection, native session, and merge tests passed. After review fixes, the managed-directory and native-session suites passed 550 tests, including intervening canonical edits, same-run fresh restore, failed handoff collection, and one-call fallback collection. Server typecheck passed again. The Daytona plugin suite passed 218 tests. The quota-warning regression failed before the fix and passed afterward. - Three real Daytona campaigns passed before the final handoff review fixes. The latest kept PID 547 across all three turns. The first checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes. Public API reads verified the binary, note contents, and deletion after every turn. Test cleanup deleted the sandbox. - The final head was also deployed to an isolated cloud staging instance and passed three UI-triggered native Codex turns with managed instructions. All three retained the same process ID/start time, native session, provider session, runner instance, and Daytona sandbox. Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes. Independent canonical API reads verified every byte of the unchanged 8 MiB binary and the exact note contents after every turn; the deleted file returned 404 after turns 2 and 3. After restoring the original lifecycle and agent-auth configuration, removing the temporary secret, pausing the test agent, and deleting both test sandboxes, independent canonical API reads still verified the entire binary, the final 78-byte three-line note, and the deletion. The native runner flag remained enabled and the final serving revision remained the PR head. - Two earlier staging attempts are preserved as failures and are excluded from the acceptance result: a saved ChatGPT login failed with a provider routing 401, and its subsequent stopped-sandbox retry failed before provider startup with a closed-lease admission error. The successful campaign used a fresh sandbox and a temporary encrypted API-key binding. The stopped-lease retry remains unexplained; this campaign does not establish recovery of that failed sandbox. - All [Paperclip CI gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355) pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test partitions, build, typecheck, runner verification, E2E shards, and the Canary clean public-npm install. Greptile reviewed that exact head at 5/5 with no unresolved review threads or outstanding findings. - Full local repository coverage used the existing CI partitions, but the 40,000-file Git streaming stress test timed out and its local retry was interrupted by macOS thermal emergency sleep; this is not a green full local suite claim. The exact stress test passed on the final head in [CI server shard 2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290), in 111.9 seconds. - Repeat the live test with configured credentials and a Linux runner artifact: `pnpm test:e2e:runner -- --id daytona-warm-continuity.runner-codex.daytona.warm-three-turn`. ## Risks - This is a file-level checkpoint, not an atomic snapshot of the whole folder. Background writes after a capture are saved by the next checkpoint or final stopped collection. - Metadata scans still visit all paths. Modified files transfer in full; unchanged files do not rehash or transfer. - Incorrect ownership or reuse could collect the wrong directory. Run ownership fences, current authorization, stable capture validation, and stopped collection fallbacks are covered by tests. - Warm reuse remains opt-in. No database migration or fleet default changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d432dc7fa3 |
Add GitHub-synced skill sources (#14713)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Company skills supply instructions and files to those agents. > - GitHub imports already exist, but users cannot manage repositories as skill sources. > - Repository refresh also needs caller-authorized access and complete local packages. > - This pull request adds Sources inside Skills and reuses GitHub connections from Apps. > - Installed snapshots let agents use skills without fetching GitHub during a run. > - Manual refresh preserves skill identity and leaves failed imports on their last good version. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: skills UI, server, database, shared contracts, and runtime materialization. **Problem or motivation** Users keep skills in GitHub repositories. They need a clear way to select, import, and refresh those skills. Existing imports do not expose repository management or consistently preserve supporting files. **Proposed solution** Add company-scoped skill sources. Browse repositories from all accessible GitHub connections, or paste a public repository or branch URL. Select whole skill packages, inspect included files and reference warnings, and install complete, immutable snapshots. Refresh each source manually. **Alternatives considered** Project repository settings hide the workflow from Skills. A second GitHub connector would duplicate credentials and grants. Upstream editing and PR creation are separate work. **Roadmap alignment** This implements the Skills Manager direction in ROADMAP.md. The maintainer requested this scope and reviewed the component and full-app journey stories before implementation. Related reports: Refs #10285, Refs #10949, Refs #13464. Related work: #14356, #13656, #9268. ## What Changed - Add source and entry records, an idempotent migration, company-scoped APIs, and legacy GitHub import adoption. - Reuse current caller grants and credential refresh. Combine and deduplicate repository inventories across accessible connections. Pasted public URLs also prefer the active user’s authorized connections. Tokens stay in the Git child environment, never argv or disk. - Fetch a shallow Git snapshot at one immutable commit. Scan the full local tree, including hidden and nested folders. Read Git objects without checkout or archive transformations and enforce nested package boundaries. - Bound Git downloads to 128 MiB and three minutes. Cancel active process groups and remove incomplete downloads. Preserve cancellation and deadlines while progress drains; close stalled HTTP progress streams after 30 seconds. Reuse caller-scoped temporary snapshots for preview/import after reauthorization. - Index repository package boundaries once and cap expanded work at 1,000 packages, 10,000 files, and 100 MiB, including repeated copies of shared blobs. Bound path depth and the shared path index. Discovery keeps audited manifests without retaining all package bodies. - Resolve moving branches before fetching so unchanged discovery reuses caller-scoped snapshots. Limit active scans, scan frequency, and new downloads per caller and company; quotas apply before metadata reads and across connections, and cached scans do not consume the download quota. - Store complete versions with script content, binary bytes, and executable modes. Preserve these through copies, runtime caches, and runner packaging. - Stage downloads before publication. Use source leases, revision checks, and transactional activity records. Keep installed versions after failures, upstream deletion, deselection, and disconnect. - Add the approved import flow, Sources page, selection tree, provenance, read-only Studio behavior, and saved return from GitHub setup. - Add package manifests, commit-pinned file previews, and separate runtime requirements and reference warnings. Supporting files are included together; nested skills remain independently selectable. Preview requests reauthorize the caller and re-audit package content. - Show installed skills as compact links beneath each source. Repository titles open GitHub. Keep Refresh, Select skills, and Disconnect source in a three-dot menu. Source rows omit the branch, imported count, and refresh timestamp; action alignment and repository titles work at narrow widths. - Stream discovery metadata over an opt-in NDJSON response. Show measured Git download progress and real package/file counts, animate newly checked skills, support cancellation, and require a complete scan before selection. Keep the existing JSON API. - Retain component stories and add a separate full-app journey story group. Include fixed progress states and interactive scan, large-repository, interruption, and saving stories. - Update Skills documentation and product contracts. Suppress private GitHub skill references in telemetry. Privacy review requested for the telemetry changes. ## Verification - Local repository typecheck, full build, token gates, and Storybook build passed during this work. Focused transport, authorization, scanner, persistence, route, and UI tests pass. The final UI refinement passes all eight focused UI tests, UI typecheck/build, and token gates. The scanner resource and repeated-discovery fixes pass 132 focused scanner, transport, authorization, source-service, route, and rate-limit tests, plus server typecheck/build. Full-suite verification comes from CI; the older full local Vitest run was stopped after unrelated chat failures and a font-test failure, all of which passed in fresh focused runs. At commit `1098d5996`, all 54 active checks pass; two optional Storybook jobs are skipped. CI covers repository typecheck, build, the full test suites, browser shards, and the canary dry run. Greptile is 5/5 with no open findings; the security scan also passes. - Adversarial scanner tests verify repeated-blob byte accounting with and without declared sizes, package/file/path caps, one-time repository indexing, metadata-only discovery audits, and nested package boundaries. Additional tests cover branch movement, snapshot reuse, caller/company quotas, isolation across connections, active-lease cleanup, quota recovery, and rejection before any metadata API call. - Real Git tests verify hidden paths, exact binary bytes, executable modes, export-ignore preservation, symlink/submodule reporting, pinned commits, caller-scoped cache reuse, cancellation, cleanup, and credential isolation. Regression tests hold both download slots with permanently blocked progress callbacks, verify timeout/cancellation cleanup and retry, and exercise HTTP backpressure cancellation. Access tests cover automatic public-URL connection selection and revoked grants. Database tests verify company and grant audiences. - Live isolated browser test: the public `anthropics/skills` scan now completes and discovers all 20 skills without connecting an account. Imported canvas-design with all 83 files, opened it from Sources, and verified the installed binary-font preview/download control. Package previews also expose the complete file inventory before import. Cancelled an active Git download and retried successfully to all 20 discovered skills; the browser displayed measured download progress. The current audits reject four other packages; eligible selections remain importable. - Browser checks verify the simplified source rows at desktop and narrow widths, keyboard navigation into the actions menu, Refresh from the menu, selection, and fixture disconnect with installed skills retained. Storybook includes a menu-open checkpoint and a 320px layout. - Storybook includes receiving/preparing download checkpoints and a timed full-app import journey, plus cancellation, retry, large-repository, and saving states. Streaming tests cover split UTF-8 frames, incomplete streams, late responses, cross-company requests, HTTP errors, and JSON compatibility. - Earlier live acceptance on this PR imported `stitch-skill` with `DESIGN.md`, assigned it to an agent, disconnected its source, and ran a successful Studio test that read both installed files. An editable copy changed independently. Both Skills variants, mobile selection, and return from GitHub setup were exercised. - Private access, revoked credentials, OAuth success return, binary/script preservation, concurrent refresh, transaction rollback, version pins, and legacy adoption have automated coverage. A real private-repository OAuth grant was not created during this test. ## Risks - The migration groups recognizable legacy imports without provider calls. Their first successful refresh completes the local package snapshot. - Reference checks are advisory. They cover Markdown links and explicit relative resource paths, not arbitrary runtime dependency graphs. Preview text is capped at 64 KiB; imported bytes remain complete. - Git must be installed on the server. Shallow fetches still download the branch snapshot, including files outside selected packages. Downloads have size/time/concurrency limits. Temporary caches are bounded and caller-scoped. GitHub API quota still applies to repository metadata and the connection picker; content no longer uses per-file API requests. Failed scans retain installed content. - Sources depend on the current caller's GitHub access. A saved connection does not grant access to another person's token. - GitHub script support and immediate manual refresh are explicit maintainer-approved requirements. The operator trusts the selected repository and accepts upstream script and executable-mode changes on refresh. Static audits are not a sandbox or a guarantee of safe code; agents may later invoke installed helpers under their runtime permissions. Import and refresh do not execute scripts, hooks, package installation, or builds. Raw URL and skills.sh imports keep their prior script restrictions. - Source originals remain read-only. Refresh affects subsequent unpinned runs; explicit pins and active runs retain their versions. - The telemetry change removes source-managed GitHub identifiers from skill-reference events. It introduces no event or field. Please review the privacy boundary. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
25c422ba7e |
fix(apps): request minimal Google service scopes (#14740)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Google connections give agents service-specific tools through
governed credentials.
> - Each connection has a reviewed OAuth scope set.
> - Docs, Sheets, and Slides request Drive permissions in addition to
their own service scopes.
> - Google documents these permissions as alternatives, not combined
requirements.
> - This pull request removes those extra permissions and redundant
Calendar write free/busy access.
> - Users grant fewer permissions without adding tools or changing
connection access policy.
## Linked Issues or Issue Description
Refs #13820. Related #14739 changes other connector permissions; it does
not reduce these Google profiles.
Companion broker PR:
https://github.com/paperclipai/paperclip-cloud/pull/615. Ship the
matching changes together after fresh-grant validation.
**What happened?**
Seven Google profiles request redundant scopes. Docs, Sheets, and Slides
request Drive scopes. Calendar write requests free/busy even though
calendar.events authorizes its availability tool.
**Expected behavior**
Each profile requests only the scopes required for its reviewed tools.
Managed and customer-owned OAuth methods use the same set.
**Steps to reproduce**
Inspect the Google profile registry and the four app definitions on the
base commit. Compare their scope sets with Google's MCP authorization
alternatives linked in the updated documentation.
**Paperclip version or commit**
Base:
|
||
|
|
94e8dec56b |
fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The runner sends authorized tool calls to the server and saves their results. > - A provider turn can stop while a server write is still running. > - The old shutdown path invented a failed result that could conflict with the real result. > - Truncated execution input and incomplete recovery records made the failure harder to diagnose. > - This pull request preserves exact inputs and actual outcomes through shutdown and restart. > - Tests force the race and crash boundaries so safe retries do not repeat writes. ## Linked Issues or Issue Description **What happened?** Stopping a turn during a server tool call could record a false failure, then reject the actual result as a conflict. The diagnostic input formatter could truncate instruction content before execution. A crash during saved-result delivery could leave that delivery permanently indeterminate. Cleanup could hide the first failure, and a retry could overwrite earlier run logs. **Expected behavior** Keep dispatched tools pending until their actual result is known. Preserve accepted input bytes. Accept identical result delivery without failing the task. Reject conflicting results with enough evidence to diagnose them. Recover saved-result delivery without repeating the business operation. **Steps to reproduce** 1. Hold an instruction update at the filesystem commit barrier. 2. Stop its provider turn before the server returns the result. 3. Release the write, deliver its result, and replay the same result. 4. Repeat with a restart before and after the delivery receipt is saved. 5. Check that there is one write and one audit row, and that the exact result survives. **Paperclip version or commit** The change was developed from `44736c9c7` and rebased onto `0e5830887`. **Deployment mode** Self-hosted server with the native runner. Tests use local runner processes, scripted providers, and PostgreSQL. Related work: #12353 added durable semantic tool receipts; #12384 added durable Codex tool recovery; #12404 bound semantic tools to ACPX sessions. #14633 covers separate native-provider cancellation and qualification work. This PR addresses server semantic-tool outcomes and their durable delivery. No duplicate fix was found. AgentMail discovery is outside this PR. ## What Changed - Close turn admission without inventing results for dispatched tools. Keep pending calls and accept late actual results. - Accept identical result replay with a diagnostic warning. Include call identity and both result hashes in real conflict errors. - Preserve exact execution arguments. Reject prohibited or oversized input before dispatch. Keep diagnostic previews redacted and bounded. - Commit instruction-attempt evidence before the filesystem write. Save completed mutation receipts so concurrent and restarted duplicates return the first result. Recheck authorization before replay. An attempt without a completed result stays unknown and cannot execute again. Definite pre-write failures save and replay their original error without another write. - Recover an interrupted saved-result delivery only for backends with durable result receipts. Never replay an ordinary business operation with an unknown outcome. - Preserve the initiating error when cleanup also fails. Record incomplete settlement evidence. Propagate typed unknown-outcome errors through the native tool wrapper without creating a false completed tool result. - Append run-log attempts and restore the durable log before appending after local file loss. Reject incomplete restores. Publish a restored prefix only if the destination is absent so concurrent attempts cannot overwrite new lines. - Add deterministic race, crash, replay, authorization, exact-content, and log-restoration tests. Document their assertions in `packages/paperclip-runner/docs/durable-recovery.md`. ## Verification - Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55 applicable checks pass; four conditional/manual checks are skipped. This includes build, typecheck, Rust, both runner TypeScript shards, server and workspace tests, all eight browser shards, isolated runner compilation, and the clean-install release dry run. [CI run](https://github.com/paperclipai/paperclip/actions/runs/36746101110). - Greptile reviewed this exact head at 5/5 with zero new findings. All three earlier review threads are resolved. - Focused local verification includes 11 instruction integration tests, 23 surrounding authority/tool tests, 26 run-log tests, and 169 controller/driver tests. The post-rebase controller/transport/runtime selection passed 415 tests. The full Rust release suite passed 617 tests with two ignored. The real-process SIGKILL recovery test passed three consecutive runs. - The fault matrix in `packages/paperclip-runner/docs/durable-recovery.md` uses explicit barriers, real PostgreSQL rollback, durable journal reloads, and killed runner processes. It covers late results, identical and conflicting replay, exact long content, concurrent log restoration, lost commit acknowledgements, and definite failure replay after the original CAS base becomes valid again. No paid model calls are needed. - Full local recursive typecheck and build passed during implementation. Server typecheck and the runner TypeScript build passed after the review fixes. The broad local repository test run was stopped after repeated database startup timeouts. Four timing/launch failures in an earlier broad runner run passed focused reruns without changed assertions or timeouts. These are local verification limitations; the complete current-head CI suite is green. An earlier CI workspace job received an infrastructure shutdown signal; its current-head replacement passed. ## Risks - A stopped turn can remain blocked when a dispatched operation has no proven result. The system does not guess its outcome or rerun its effect. - Conflicting results still fail settlement. Existing failed or conflicting journals are not repaired automatically. - Accepted semantic input is limited to 480 KiB of encoded JSON to fit the encrypted transport. Larger input fails before execution. - Instruction filesystem writes and database receipts are not one atomic storage operation. A separately committed attempt and audit record survive rollback. An attempt without a completed success or definite pre-write failure receipt remains blocked as an unknown outcome. It is not replayed or reported as success. - Run-log restoration now reads the durable object before appending when the local log is missing. Failed or incomplete reads reject the append. - No schema migration, dependency change, workflow change, or AgentMail change is included. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test analysis. The exact served model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; the broad local run limitation is recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
af5c2d101c |
fix(paperclip-runner): deliver the shutdown settlement event past the terminal-turn gate (#14668)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The OpenCode driver maps provider events to runtime request events > - A provider turn can end before a pending runtime request receives its answer > - The consumer reads one turn's events and stops at that turn's terminal event > - A settlement event that arrives after that terminal event never reaches the consumer > - This pull request settles the request inside its own turn, before the terminal event > - The benefit is reliable request settlement without weakening late-frame protection ## Linked Issues or Issue Description **What happened?** A pending runtime request stayed open after an OpenCode turn failed through `session.error`. Session shutdown then dropped its settlement event as a late provider frame. **Expected behavior** The driver must deliver `runtime_request.expired` with the original `turnId` and `itemId`, inside the same single read pass the consumer performs on that turn. **Steps to reproduce** 1. Start an OpenCode turn that creates a native runtime request. 2. Leave the request pending and fail the turn through `session.error`. 3. Close the session and inspect the emitted events. **Paperclip version or commit** `22d41c6081f05658e0d7c8485d0f22f35af4a79e` **Deployment mode** Built from source with the OpenCode driver test fixture. ## What Changed - Settle a pending runtime request as soon as its own turn goes terminal, before the terminal turn event. - Add an optional `bypassTerminalTurnGate` parameter to the OpenCode session emitter, and set it on the settlement emit. - Keep a settlement loop in session close as a fallback for a request whose turn never went terminal. - Add a fixture trigger and a regression test that reads one turn in a single pass. - Keep the late provider frame gate unchanged for every provider event path. ## Verification - Run `pnpm vitest run packages/paperclip-runner/src/drivers/opencode/opencode-server-driver.test.ts`. - Run `tsc -p tsconfig.json --noEmit` in `packages/paperclip-runner`. - Run `tsc -p tsconfig.surfaces.json --noEmit` in `packages/paperclip-runner`. - Confirm that the regression test receives `runtime_request.expired` with the original identifiers. - Confirm that the late-frame tests still report dropped provider frames. ## Risks Four call sites now reach the settlement path: session close and the three terminal turn paths (completed, cancelled, and failed). At the three terminal turn paths the turn is still the active turn, so the gate admits the settlement event with or without the parameter. Session close is the only place where the parameter changes the result of the gate, and only for a request whose turn already ended. Every provider event path keeps the existing terminal-turn gate. The known driver test failure is pre-existing and does not touch this change. ## Model Used Claude Sonnet 5, with code execution and test assistance. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the changed tests locally; the known pre-existing failure remains documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation, or no documentation change applies - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I have addressed all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c561642b4 |
fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents ask for decisions and optional details through cards in chat. > - A clear approval in a message can leave the matching card pending. > - An unanswered question can also block an unrelated later reply. > - Decisions need a saved source message, while optional questions need to remain answerable in history. > - This pull request records conversational decisions and lets users move on from questions and answer them later. ## Linked Issues or Issue Description **What happened?** Native Claude and Codex could act on approval in chat while the original approval card stayed pending. Pending question forms stayed above the composer, were absent from history, and could suppress later chat replies. A late native question answer could wait for a finished run to reconnect. **Expected behavior** The active agent records a clear approval or refusal against the exact card and user message. Ambiguous replies do not grant consent. Users can send another message without answering a question. The question remains pending in history and can be reopened and answered later. The saved answer reaches the agent. **Steps to reproduce** 1. Ask an agent to propose work with a confirmation card, then approve it in chat. 2. Check that the original card records that approval before work starts. 3. Ask an interactive question, send an unrelated message, and reload. 4. Open the unanswered question from history and submit an answer. Related work: #14408 added completion delivery. #14607 tests completion reporting turns. Neither records conversational answers on approval cards. ## What Changed - Add a confirmation endpoint backed by a user comment, with schema validation, OpenAPI discovery, and native Plan-mode access. Ask mode remains read-only. - Check company, active run, actor, current session, message provenance, revision, and resolver policy. Save the decision and audit in one transaction. Retries do not repeat effects. Emit resolution telemetry after commit. - Give fresh and resumed chat turns the actual pending confirmation identities. Teach agents to save clear conversational decisions before acting and to clarify ambiguity. - Keep unanswered Agent Chat questions as compact history entries. A newer user message closes the old form. Question cards never contribute to composer pending counts or navigation, including after dismissing a fresh form. The history card is the sole reminder; clicking it restores that exact form and draft. - Preserve Agent Chat questions when later messages or questions arrive. Historical ordinary inputs no longer gate later chat replies. Current-run requests, task execution, and governed approvals keep their gates. Remove the special acknowledgement-publication proof helpers that this rule replaces. - Route answers to finished native runs through durable fresh-wake delivery, with existing idempotency and source-question context. Settle late replies against contiguous completed conversation turns and freeze their history replay; failed, unhandled, and newly arriving messages remain actionable. - Add real-component Storybook scenarios, database and UI regressions, and a three-turn native Claude/Codex E2E case. Capture distinct, UI-ready screenshots and report the individual assertions. ## Verification - Focused decision/publication/UI regressions after merging master: 288 passed; subsequent UI draft, failed-send, and conversation checks: 199 passed. - Native question and durable delivery regressions: 106 passed, including all four terminal run states and exactly-once late delivery. Seven targeted regressions fail against the original implementation and pass with the fix. - Latest conversation/decision/native-delivery regressions after the master merge: 121 passed. Covers completed progress, missing or failed intervening turns, new messages during a late reply, stale sessions, and frozen retry/replay boundaries. Four new assertions fail before the ordering fix. - E2E support suite after the master merge: 792 passed. Negative controls reject expired cards, wrong questions/answers, stale or missing replies, unrelated clarification forms, and unexpected tasks. - The embedded-browser walkthrough caught one additional defect: dismissing a fresh question still showed a composer badge. Both Cancel and close-button regressions failed before the fix. The fix at `65f2ade12` passes 170 chat-thread tests and 792 E2E support tests. After merging master, 232 chat-thread/confirmation tests, server/UI typechecks, and token gates pass. The preview and two-provider live E2E pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at that commit. All 55 checks are now successful at `e5512a206` (four conditional checks skipped), including the aggregate verification gate and clean-install canary test. The first attempt was interrupted by simultaneous CI worker shutdowns; one failed-job rerun passed without code changes. - [Published Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on): nine real-component scenarios. Manually exercised move on, reopen, preserve draft, answer later, answer one of multiple questions, and a custom mobile answer in the embedded browser. Retested fresh Cancel and close-button dismissal in the updated build, then reopened and submitted the preserved Green selection and inspected its answered receipt. Static preview has no live model/backend; its callbacks are fixture responses. - [First live campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/) reproduced the late-answer completion-state defect on both providers despite correct saved answers and acknowledgements. It also exposed a valid imperative clarification rejected by the old oracle. Both issues are fixed with regression controls; this failing run is retained as evidence. - [Four-cell qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/) passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous confirmation, each on native Claude and Codex. Inspected saved state, source-message decisions, visible cards, and agent replies. Both late-answer chats settled to waiting; no unrequested tasks were created. [Final branch rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/) passed 2/2 at `142630720`: the same unanswered-question journey after merging master, plus an additional screenshot and browser assertion for the actual late-answer acknowledgement. - [Composer-reminder E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/) passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each, with explicit no-badge assertions before and after reload. Inspected saved pending/answered state, both screenshots with a clear composer, and actual Blue acknowledgements; all five behavioral matchers passed per provider and neither created tasks. Cost coverage is partial; this is bounded workflow qualification. - [Fresh-dismissal E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/) passed 2/2 at `e5512a206`: native Claude and Codex, including fresh Cancel, clear composer, reopen, unrelated message, reload, late Blue answer, and actual agent acknowledgement. All five behavioral matchers pass per provider. Inspected the fresh-dismissal screenshots and saved pending/answered identity; neither created tasks. Cost coverage is partial (4/6 runs). - Prior evidence remains available in [the earlier campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/). Its early loading screenshot and overwritten final capture prompted the UI-ready, distinct screenshot fixes. ## Risks - The model interprets intent. The server verifies permission and provenance; it does not infer consent from text. Ambiguous and unrelated replies are not approvals. - Historical questions can accumulate. They remain visible, pending, and answerable; no automatic answer or expiry is invented. - The change to completion gates is scoped to Agent Chat and ordinary historical inputs. Current-turn and governed approvals retain their existing controls. - Live qualification is limited to the selected stories. Broader native onboarding finalization remains separate work. - No database migration. Telemetry adds no fields or values; the contract and README document the commit boundary. Privacy review was requested on the PR. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser-test orchestration. The exact model ID and context-window size are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a36cbffa9e |
fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path > - Paperclip manages agents and the services they need for work. > - Agents use connection search to discover a setup path before they request access. > - Tool-only filtering hid channel and AI methods from this search. > - Requiring every query word to match rejected normal task descriptions. > - A small aggregator index also omitted supported apps such as Circleback. > - This pull request broadens retrieval and returns purpose-specific setup guidance. > - Agents can choose a relevant result while existing access and provider-choice checks still apply. ## Linked Issues or Issue Description **What happened?** A search such as “AgentMail create an email address and manage an agent mailbox” returned no usable result. “Help me find tools for circle back” also missed Composio's supported Circleback toolkit. Queries longer than 200 characters failed validation. **Expected behavior** Return useful native and verified aggregator matches from natural-language queries. Include channel/email methods when Chat connectors is enabled. Identify each method's purpose and the correct setup path. **Steps to reproduce** Use the queries above with `connections_search` from an active task. Enable Chat connectors for the AgentMail case. The regression suite reproduces these misses before the change. Related routing work: #13941. This change does not change the runner failure path or add channel setup to tool-only connection cards. ## What Changed - Rank name and capability matches. Accept extra words, split names, small spelling errors, and queries up to 4,000 characters. - Include tool, channel/email, and AI methods. Return company-prefix setup links for channel and AI flows. - Add a dated snapshot of 1,583 official Composio toolkit names and a refresh script. Merge duplicate MCP variants for search and link each support claim to official evidence. - Find authorized indexed aggregator namespaces within longer queries. Return multiple app matches when the agent needs to choose. - Prefer exact app names over fuzzy matches for other apps; retain existing AI readiness. - Preserve native preference, company and identity boundaries, administrative denials, and saved provider consent. - Add relevance and database regressions, extend native tool-authority coverage, and document search behavior. ## Verification - Red: 15 new assertions failed against the previous implementation; the existing baseline passed. Added red-green regressions for Motion versus fuzzy Notion and existing AI access during review. A further regression covers mixed ready/unconfigured AI results and their per-result setup guidance. - Green: all 103 tests in the eight focused shared, database, runtime-tool, fixture, and route suites pass on the latest commit. - `pnpm -r typecheck` and `pnpm build` passed. - Latest-commit CI passed: 54 successful checks and two skipped checks, including the full test matrix, browser E2E, typecheck, build, and canary dry run. - The long local `pnpm test:run` attempt began before the review fixes and retained transformed pre-fix search code; it also hit an unrelated timing failure. Fresh serial reruns of the affected search suites and three timeout cases passed all 124 tests. Parallel local route shards hit two additional database setup timeouts; both suites passed all 17 tests on a fresh serial rerun. The complete corresponding CI suites also passed. Duplicate broad local runs were stopped after CI completed. The local UI suite independently passed all 7,007 tests. - Greptile: 5/5 on `cc6a0180d`; all review findings resolved. - Browser verification passed in a disposable local instance through real process-agent search requests: AgentMail opened its setup flow with the requester selected; the saved Circleback choice produced the Composio setup card; a paragraph-length Notion query produced its setup card. No provider credentials or external accounts were created. - The browser test caught an invalid UUID-based setup URL. The fix uses the company prefix and has a regression assertion. - A 3,971-character catalog query found Circleback first in a local 10 ms spot check after sharing query preparation across the catalog scan. This is a single measurement, not a performance guarantee. ## Risks - Broader retrieval can return extra candidates. Named services rank first; agents must select the relevant method. - The public support snapshot can age. It proves catalog support, not account authorization or the availability of every requested action. - Channel and AI methods use existing setup links. The tool connection card still accepts tool methods only. - No schema, migration, credential, or runner lifecycle changes. ## Model Used OpenAI Codex (GPT-6). The session does not expose a more specific model identifier or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
dd7fc1f90a |
fix: raise the native journal read limit to 256 MiB (#14711)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions persist control-plane state so they can resume safely. > - The state includes committed provider history needed for recovery. > - The server, runnerd recovery, and durable control plane validate this state before trusting its identity. > - Their differing 64 MiB and 192 MiB limits can reject a valid journal before recovery. > - This pull request aligns all three local state limits at 256 MiB. > - Larger files remain bounded, while recovery can read larger valid histories. ## Linked Issues or Issue Description Refs #13882 Refs #14312 ## What Changed - Raise the server and runnerd recovery limits from 64 MiB to 256 MiB, and align the durable control-plane limit from 192 MiB to 256 MiB. - Add coverage for a valid history above 64 MiB and rejection above 256 MiB. ## Verification - Matching server recovery passed with more than 64 MiB of actual committed event payloads (128 events with 512 KiB deltas). - The real runnerd exact-authority resume regression with the test Codex provider passed with 193 MiB of valid JSON whitespace appended. It crosses the former 192 MiB core limit and confirms the same provider identity. This exercises the runner process and durable control plane with a simulated provider, not a live OpenAI API call. This test used approximately 1.15 GiB peak RSS. - The actual runnerd reader accepted valid 256 MiB JSON and rejected valid 256 MiB + 1 byte. The reader call took 231 ms; the fresh process peaked at 1,244 MiB RSS. - Server tests reject mismatched identity above 64 MiB and files above 256 MiB. - `pnpm -r typecheck`, `pnpm build`, and `git diff --check` passed. - Full local `pnpm test:run`: 13,730 passed, 575 skipped, 7 failed across 6 files. All failures were embedded PostgreSQL startup errors after five attempts. They affected agent hiring, instruction revisions, environment images, reviewed chat bindings, issue monitoring, and legacy continuation authority. The focused journal tests passed; the latest pushed head passed all ordinary CI checks. Superagent is the only blocking check. ## Risks - **Open review concern:** Greptile is 5/5, but Superagent is `ACTION_REQUIRED` with two P2 findings on the server and runnerd readers. Both flag the increased synchronous parsing and memory cost. This PR keeps the requested fixed-limit change small. It does not add a worker parser or a process-wide memory budget. This resource tradeoff needs review before merge. - Large state parsing is synchronous and can consume several times the file size in memory. - Remote checkpoint archive and expanded-size limits remain 64 MiB, so this change alone does not make larger remote checkpoint transfers portable. ## Model Used - OpenAI Codex, GPT-6, with delegated assistance from `gpt-6-luna` at high reasoning effort; tool use and code execution. The GPT-6 context window is not exposed in this task runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d30b03bd8c |
test: add persistent E2E coverage for human blocker decisions (#14707)
## Thinking Path > - Paperclip manages work for AI agents. > - Agents use the coordination skill when work needs human authority or a scope decision. > - PR #14188 replaced automatic manager escalation with direct blocker handling. > - This behavior needs real browser, server, database, and provider tests. > - The test must verify saved human input, task ownership, and resumed work. > - This pull request adds six reusable Product E2E cases and improves the skill examples that they exercise. ## Linked Issues or Issue Description Refs #14188. The merged change needs repeatable behavior coverage. The new suite tests missing administrator access, missing hiring permission, and requester scope questions. Searches found no duplicate blocker-guidance suite. This extends the existing eval system described in ROADMAP.md. ## What Changed - Add the explicit-only `blocker-guidance` Product E2E suite. It has three local scenarios on legacy Codex and legacy Claude. - Use the production UI and public APIs to create work, save a human-only question or confirmation, answer it after reload, and resume the same task. - Check requester identity, ownership history, manager activity, hiring, saved answers, and completion. Keep direct text input as a separate UX result. - Save pending and final screenshots, API checkpoints, skill hashes, provider evidence, and billing data through the existing report pipeline. - Isolate the Claude fixture home. Verify the served skill bytes before dispatch so an old installed skill cannot silently replace the evaluated skill. - Improve the coordination and hiring skill examples. Include the human-only policy, requester address, wake behavior, and handling of authorized scope changes. - Grader v5 requires the exact approved public welcome note as a new worker comment. Browser input checks reject unwritable scope cards before clicking, and confirmation direction must be saved in the resolution before the worker wakes. - Add grader calibration and browser-input tests. Update the fixture guide and generated capability inventories. ## Verification - `pnpm build`: passed after rebasing onto current master. - `pnpm -r typecheck`: passed. - `pnpm test:e2e:runner:typecheck`: passed. - `pnpm test:e2e:runner:unit`: 742 tests passed. - `pnpm test:e2e:runner:browser-support blocker-input.spec.ts`: 10 tests passed. - `pnpm test:e2e:runner -- --list --suite blocker-guidance`: six cells found. - Capability inventory and generated-contract checks: passed. - `pnpm exec vitest run server/src/__tests__/hiring-operational-examples.test.ts`: four tests passed after synchronizing the generated API reference and section anchor. - Full general and serialized test suites: passed in CI on `6652cee74517039676bad6a720f213625d265acd`. The redundant local `pnpm test:run` was interrupted after complete CI coverage passed; it is not claimed as a completed local full-suite run. - Final GitHub checks: 54 passed, two optional Storybook checks skipped. The runtime-exposure startup test hit a 10-second readiness timeout once, passed a targeted local reproduction, and its CI shard passed the single retry without code changes. - Current-head Greptile: 5/5, clean check, zero unresolved threads. - Historical live measurement on September 29 at `4edc77ae2b95b10dd61426ce3f042bac00527ad9`: three independent six-cell runs scored 5/6, 6/6, and 6/6. Claude Sonnet 4.6 passed 9/9. Codex `gpt-5.6-sol` passed 8/9. These runs predate this rebase. - Version 5 changes the scope answer to an exact approved publication draft. The historical runs do not qualify that new requirement; the two-provider scope pilot at `49a1f4eab369948b9e3b34a6ce436489e875e4ec` passed Codex and failed Claude. Claude posted the correct salary-free sentence but omitted its required reference line from that comment, placing the reference in a separate completion message. The `public-welcome-note` check correctly failed. An earlier Claude database-startup failure was retained separately; its fresh-instance retry reached the model. This pilot is not a six-cell qualification. - The failed Codex scope case omitted `addresseeUserId`. The strict routing check remains. All 18 attempts had clean evidence manifests and passed cleanup. - To repeat with provider credentials: `pnpm test:e2e:runner -- --suite blocker-guidance --max-parallel 1`. This is a paid, opt-in suite and is excluded from `--all`. ## Risks - The live suite measures variable model behavior. The retained 17/18 historical result and the current 1/2 scope pilot are not all-pass qualifications. These paid cases are opt-in; their observed model failures remain visible independently of deterministic CI checks. - A separate generic task-replacement diagnostic still exposed a Claude refusal. The ordinary cases use specific business decisions. The diagnostic is not a standalone catalog case in this change. - Earlier measurements included an old installed Claude skill and test defects. Their grades remain retained and are not combined with the three final repetitions. - Skill examples can affect when agents ask for human input. Downstream permission checks still apply. - Native runners, Daytona, agent-requester routing, and real external connection authorization are outside this suite. - Raw provider traces and credentials remain private. No screenshots, raw reports, secrets, workflow changes, or lockfile changes are committed. ## Model Used OpenAI GPT-6 through Codex assisted with this change. The exact deployed variant and context window size are not exposed in this session. The assistant used reasoning, repository edits, tool use, and shell execution. The evaluated models were `gpt-5.6-sol` and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4f9a7c613 |
test(runner): guard continuation after journals exceed 2 MiB (#14312)
Add an actual runner resume regression above the former 2 MiB journal boundary and an explicit-only three-turn Daytona workflow that grows real execution history. Verify journal size and distinct completed tool calls without exporting private payloads. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2f6fa3b6dc |
fix: recover provider authentication inside tasks (#14629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need a working model provider connection to run a task. > - A provider can reject a stored credential after the task starts. > - The failed run must ask the responsible user to repair that connection. > - This pull request adds that request directly to the task and reuses Connections sign-in. > - The user can choose an API key or subscription, then continue the task with a fresh session. ## Linked Issues or Issue Description **What happened?** A run that ended with `acpx_auth_required` or another known provider authentication error did not immediately offer an inline way to connect the provider. A repair form could also lock the user to the failed account's sign-in method. **Expected behavior** Show a provider connection card in the task as soon as the authentication failure is saved. Allow the responsible user to connect or repair the provider with any supported sign-in method. Keep the connection name automatic and resume the task after successful setup. **Steps to reproduce** 1. Run a task with a supported provider and an expired or invalid credential. 2. Let the run fail with a provider authentication error. 3. Open the task and attempt to repair the connection. Related work: Refs #13724 and #13726. This change adds the inline task repair flow and method choice. ## What Changed - Classify provider authentication failures and create one connection request for the current task. A persisted blocked classification suppresses automatic retries only after the repair card is created; unsupported providers retain their existing recovery path. - Mark only the attributed, unchanged credential as needing sign-in. Preserve credentials that were refreshed after the failed run started. - Reuse the provider sign-in controls inside the task. Allow API key and subscription choices for Claude, Codex, and Grok. Keep names hidden and generate a default from the user, provider, and method. - Keep the existing account when reconnecting with the same method. Create and select another account when the method changes. Validate updates to explicit agent bindings through the normal agent save path. - Require explicit adoption for legacy agent authentication. Validate in the agent environment, then commit the binding, connection install, audit, and card completion in one transaction. Keep failed setup and account selection visible and retryable. - Add regression tests and update the specification and Connections documentation. ## Verification - Fresh local verification: 199 tests passed across the inline form, provider method selector, default naming, authentication and recovery classifiers, run liveness, OpenAPI routes, database adoption/rollback, and Cursor execution suites. The adoption database suite also passed against disposable Docker PostgreSQL. - Full repository `pnpm build` and `pnpm -r typecheck` passed on the latest commit. Token gates are clean. - Embedded browser: opened real task cards from seeded authentication failures; switched Claude from API key to subscription and back; switched Codex from subscription to API key; confirmed the name field stays hidden. Provider sign-in was not completed with real credentials. - The broad local `pnpm test:run` started before review fixes and was interrupted after the working tree changed; it is not counted as a passing full run. Fresh focused tests passed. CI supplies the full test and browser suite results for the current commit. - CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54 checks passed and two Storybook checks were skipped by their path rules. The workspace preview job passed on one rerun after a local-server startup timeout; its rerun passed 835 tests. - Greptile is 5/5 on the same commit with no actionable findings and no unresolved review threads. ## Risks - Incorrect authentication classification could prompt for a connection unnecessarily. Tests exclude tool authorization, quota, and unrelated runtime failures. - A method change selects the new personal provider default, which also applies to other agents that use that user's default. Explicit account bindings use the existing permission and runtime validation path. - Credential invalidation must not race with refresh or reconnect. The code compares the saved credential generation and grant update time under locks. - No database migration or new credential storage format is required. ## Model Used OpenAI GPT-6 through Codex. The exact model ID and context window size were not exposed in this session. Capabilities used: reasoning, repository editing, shell commands, database tests, and embedded-browser interaction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused tests listed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
eb31b926a1 |
fix(runner): keep the OpenCode session event stream open across turns (#14582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip Runner keeps provider sessions and their event streams for agent runs > - The OpenCode driver closed its event queue after each terminal turn > - A second turn on the same session then lost its response and completion events > - This pull request keeps the queue open between turns and rejects late events for sealed turns > - The benefit is reliable multi-turn OpenCode sessions with visible diagnostics for late provider events ## Linked Issues or Issue Description **What happened?** The OpenCode driver closed its event queue when a turn completed, was cancelled, or failed. A second turn on the same session then lost its response and completion events. **Expected behavior** The session must keep its event stream open between turns. Each turn must deliver its response and one terminal event. The session must close the stream only during session shutdown or an unrecoverable pump error. **Steps to reproduce** 1. Start one OpenCode session. 2. Run one turn and wait for its terminal event. 3. Run a second turn on the same session. 4. Confirm that the second turn delivers its response and terminal event. **Paperclip version or commit** `c0e1d87ddc181329471fa80a2061b2c538bb6618` **Deployment mode** Built from source with the Paperclip Runner package test suite. ## What Changed - Keep the OpenCode event queue open across completed, cancelled, and failed turns. - Track sealed turn ids and reject later events for those turns with a diagnostic event. - Preserve queue shutdown on session close and unrecoverable pump errors. - Give each simulated fixture turn unique provider event ids. - Add regression tests for completed, cancelled, failed, and closed-session paths. ## Verification - Type check: `cd packages/paperclip-runner && node ./node_modules/typescript/bin/tsc -p tsconfig.json --noEmit` passed. - Driver tests: `cd packages/paperclip-runner && npx vitest run src/drivers/opencode/opencode-server-driver.test.ts` passed except for the known pre-existing flaky test described below. - Consumer tests: `cd packages/paperclip-runner && npx vitest run src/native-session-runtime.test.ts src/backends/harness-driver-backend.test.ts src/cli/opencode-app-server-proxy.test.ts src/conformance/harness-driver.test.ts` passed. - The known flaky test reproduced on unmodified `master` because fixture event order depends on a local MCP HTTP round-trip. - CI must run the full pull request suite. ## Risks - The queue now retains sealed turn ids for the session lifetime. OpenCode does not reuse turn ids, so this set grows with the session. - A late provider event cannot reach a later turn. The driver emits a diagnostic event so the rejection remains visible. - The change does not alter session shutdown or unrecoverable pump error handling. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. This pull request fixes an OpenCode Runner bug and does not add a roadmap feature. ## Model Used OpenAI Codex, GPT-5, tool use and code execution; Anthropic Claude Sonnet 5 also assisted with the implementation. The runtime did not provide a context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f38b5693f6 |
fix: always enable keyboard shortcuts (#14643)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The web UI has keyboard shortcuts for the inbox, task lists, cases, and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and `]` > - Shortcut enablement was an instance-wide General setting until #14141 moved it to a per-user preference that defaults to off > - The move did not carry the old instance value over, so every existing user lost shortcuts on upgrade and had to find a new toggle under Profile settings > - A toggle that only turns off a standard, input-safe feature costs a setting, a database column, two API routes, and a React context for little benefit > - This pull request removes both the instance setting and the personal preference and enables keyboard shortcuts for every signed-in user > - The benefit is one less thing to configure, no silent loss of shortcuts on upgrade, and less code to maintain ## Linked Issues or Issue Description Refs #14141 (the change that introduced the personal preference). **What existing behavior does this improve?** Keyboard shortcuts in the web UI stay off unless each user turns them on in Profile settings. **Subsystem affected** Web UI shortcuts, Profile settings, instance general settings, the `/api/auth/preferences` routes, and the `user` table. **Current behavior** Shortcuts default to off per user. #14141 moved the toggle from Instance settings → General to Profile settings and did not carry the old instance value over. Users who had shortcuts on lost them after the upgrade and had to find the new toggle. **Proposed behavior** Keyboard shortcuts are always enabled for every signed-in user. There is no instance setting and no personal preference. Shortcuts already ignore key presses inside text inputs and modal dialogs, so an opt-out is not needed. **Reason and benefit** Fewer settings, no silent loss of shortcuts on upgrade, and removal of a database column, two API routes, a query hook, and a React context that existed only to gate this feature. **Breaking changes** `GET` and `PATCH /api/auth/preferences` are removed. `PATCH /api/instance/settings/general` no longer accepts `keyboardShortcuts`; that schema is strict, so the key now returns 400. `instance.general.keyboardShortcuts` is no longer a valid `PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a warning. ## What Changed - Removed the Keyboard shortcuts section from Profile settings, the `useUserPreferences` hook, `queryKeys.auth.preferences`, and `authApi.getPreferences` / `authApi.updatePreferences`. - Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list, legacy task list, cases, and task detail pages no longer gate their key handlers. - Removed the `enabled` option from `useKeyboardShortcuts`. The app shell always registers the global shortcuts. - Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI entries, and the `currentUserPreferencesSchema` / `updateCurrentUserPreferencesSchema` validators. - Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the general settings zod schema, the settings service defaults, and `HIDEABLE_GENERAL_SECTIONS`. - Added migration `0289_drop_user_keyboard_shortcuts`, which drops `user.keyboard_shortcuts`. - Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and `docs/deploy/environment-variables.md`. - Parsed the stored general settings row with `instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a retired key left in the row cannot reset the sharing preference to `prompt` and overwrite the stored choice. - Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before. - Updated the affected tests and added a Profile settings test that asserts the toggle is gone, a hook test for the modal dialog guard, and a feedback service regression test for the retired-key case. ## Verification - Typecheck passes for `@paperclipai/shared`, `@paperclipai/db` (including the migration numbering and safety checks), `@paperclipai/server`, and `ui`. - `pnpm exec vitest run server/src/__tests__/instance-settings-routes.test.ts server/src/__tests__/openapi-routes.test.ts server/src/__tests__/auth-routes.test.ts server/src/__tests__/sentry.test.ts` → 119 passed. - `pnpm exec vitest run ui/src/components/Layout.test.tsx ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed. - `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts` → 16 passed. - `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7 passed. - `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts` (embedded Postgres) → the new retired-key test passes with the fix and fails without it. - Manual: sign in with no settings changed, open the inbox, press `j` and `k` to move the selection, press `?` to open the cheatsheet. Open Settings → Profile and confirm there is no Keyboard shortcuts section. ## Risks - The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the column has no readers after this change. If you roll back to a build from before this PR after the migration has run, re-add the column first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean DEFAULT false NOT NULL;`. The older build's ORM selects that column when it loads users. - Any external client that still sends `keyboardShortcuts` to `PATCH /api/instance/settings/general` receives a 400. No in-repo client does. - Stored `instance_settings.general.keyboardShortcuts` values are stripped on read and ignored. - Users who never turned the toggle on now get shortcuts. The handlers skip text inputs, contenteditable regions, and modal dialogs, so typing is unaffected. ## Model Used Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended thinking and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
35a4448c02 |
fix(cursor): select failure diagnostics after trace notices (#14636)
## Thinking Path > - Paperclip manages work performed by AI agents. > - The Cursor CLI adapter turns process output into run results. > - Cursor can print a trace-file notice before a real error. > - The adapter used the first stderr line as the failure summary. > - This could hide the error behind an informational file path. > - This change selects the first diagnostic after that known notice and preserves the full logs. ## Linked Issues or Issue Description **What happened?** When Cursor exits with a nonzero code, a leading `cursor-retrieval: tracing to ...` notice can become the error summary. A later error remains in stderr but is absent from the summary. If the notice is the only output, the summary does not explain that the process exited unsuccessfully. **Expected behavior** Prefer the structured error, then a stderr diagnostic, then the exit code. Keep the run failed and preserve the original logs. **Steps to reproduce** 1. Use a fixture Cursor executable that writes the trace-file notice to stderr. 2. Write `Authentication failed` on the next line, then exit with code 7. 3. The old adapter reports the trace-file notice. This change reports the authentication error. 4. Repeat with only the notice. This change reports `Cursor exited with code 7`. **Paperclip version or commit** Reproduced against master commit `17780751551b3bc1c2521f7694026c34534c46c9` with local process fixtures. **Deployment mode** Local CLI adapter. The diagnostic helper is also used by environment probes. Searched open Cursor PRs and issues. PRs #14631 and #14435 concern native ACP support; #11106 concerns MCP configuration. None changes this legacy CLI diagnostic selection. ## What Changed - Skip only the exact trace-location notice when choosing a diagnostic line. - Remove terminal control codes from summary candidates. - Use the same selection for execution and environment probes. - Preserve structured-error priority, exit status, retry decisions, and raw stdout/stderr. - Add child-process regression tests and narrow parsing cases. Document the behavior. ## Verification - Two execution regression cases failed on the previous implementation; structured-error priority already passed. - `pnpm exec vitest run packages/adapters/cursor-local`: all 16 tests passed across five files. - `pnpm --filter @paperclipai/adapter-cursor-local typecheck` passed. - `pnpm -r typecheck` passed. - All GitHub CI checks passed on the PR head. One preview-runtime readiness test failed on the first attempt; its full local suite passed (28 tests, three skips) and the failed CI shard passed on retry. No unrelated source change was needed. - The broad local `pnpm test:run` command did not complete in the available verification window and was stopped; no full local-suite pass is claimed. The full sharded GitHub CI suite passed. `pnpm build` passed. - Tests use local fixture processes. They make no Cursor provider requests. ## Risks A future Cursor notice format may no longer match and will remain visible. Retrieval error lines and unknown diagnostics remain visible. This improves diagnosis; it does not claim to fix an unknown provider or machine failure. There are no schema, authentication, cancellation, or retry-policy changes. ## Model Used OpenAI Codex (GPT-6), with reasoning, repository inspection, and command execution. The session does not expose a more specific model revision or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run the targeted tests locally and they pass; full checks are in progress - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2de43fc909 |
fix(issues): keep agent mentions as context and defer personal app authorization (#14577)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Each task has one assignee. Explicit assignment and review requests select who should act. > - An agent mention started another agent on a task it did not own. Native attachment staging then rejected that run. > - Allowing that run through startup could also let two agents work on the same task. > - Mentions should identify relevant context. They should not start work or forward comments to other tasks. > - A personal app installed on a shared agent must also wait until tool use to resolve the current user's grant. > - This pull request removes mention dispatch and keeps missing personal app credentials from blocking startup. ## Linked Issues or Issue Description **What happened?** A native agent mentioned on another agent's task failed with `paperclip_runner_attachment_staging_not_authorized`. The source task could already be complete. A nearby optional-app warning was a separate problem: personal app tools were excluded when their shared health state required attention. **Expected behavior** An agent mention is context only. It does not wake the agent, take ownership, or copy a comment onto another task. Normal feedback still reaches the assignee. Assignment and explicit review requests still dispatch work. An unavailable personal app does not block startup or produce a startup warning. Tool use requests the current user's authorization and never uses another user's grant. **Steps to reproduce** 1. Assign a task to agent A. Post a comment that mentions agent B, including a comment that closes A's task or references B's child task. 2. Confirm the comment retains its agent link and B receives no run or deferred wake. A can still receive normal feedback. 3. Install an active personal MCP connection on B. Give only Alice a grant and leave shared health at `error`. 4. Explicitly assign work to B for another user. Confirm it can finish without using the app. 5. Ask B to use the app. Confirm its tool call shows an inline connection request for the current user. Related work: Refs #11144. This change uses the existing execution-time personal grant resolution. ## What Changed - Remove mention dispatch from standalone comments and issue updates. Remove implicit forwarding of parent comments to a mentioned worker's child task. - Ignore new requests with the legacy mention wake reason before creating a run or deferred request. Preserve already accepted queue entries, which can combine assignments and feedback with a later mention. - Remove the native mention admission, staging, and finalization exceptions from this PR. Native task ownership checks remain intact. - Keep active, installed personal app tools available despite shared health errors. Remove optional-app startup warnings. Tool execution retains the current user's grant and policy checks. - Update agent instructions and product/API docs. Refresh generated capability source anchors. ## Verification - Red: comment-route regressions reproduced extra agent wakes and child comment forwarding. A separate regression proved that cancelling by the last coalesced reason could drop an accepted assignment. - Green: the targeted route, wake queue, heartbeat, workspace, responsible-user, MCP discovery, and HTTP gateway suites passed. The final queue and heartbeat rerun passed 104 tests, the restored queue adapter passed 56, and both comment-route suites passed 135. These include accepted assignment preservation, rejection of new mention requests, and normal assignee feedback. - `pnpm -r typecheck` and `pnpm build` passed locally. The full local `pnpm test:run` attempt was interrupted for review/CI fixes, so it is not claimed as a completed local pass. It exposed a cleanup timing race in the concurrent-mention assertion, now fixed and verified across 10 repetitions. CI also exposed an obsolete test waiting for the removed mention lookup; it was reproduced and fixed, then both comment suites passed. Final full-suite verification is through CI. - Final head `bd9ea4cb05a8f081c54e017760a8999f9ea6ef44`: 54 checks passed, 2 Storybook checks intentionally skipped; no pending or failing checks. Full CI includes general and serialized suites, all 8 browser shards, runner verification, typecheck, build, and canary dry run. Greptile is 5/5 on this exact commit, with no unresolved findings. - One unchanged Cursor adapter test hit its 10-second CI timeout. All 5 tests in that file passed locally; one retry of its CI shard passed all 674 tests (3 skipped). The aggregate verification gate then passed. No code or timeout was changed for that retry. - Live browser check: inserted a structured mention with the picker on a human-owned task. The saved link remained visible. Database checks found zero new runs and zero wake requests. - Live Codex runner check: explicitly assigned that task with the unavailable personal app attached. The run succeeded and committed completion without using the app or creating a connection card. - Live browser follow-up: asked the assignee to call PostHog and mentioned another enabled agent as context. Only the assignee ran. It succeeded and displayed the existing inline connection card. Only Alice's grant existed; the run belonged to a different user. - The HTTP regression covers tool discovery with no provider calls or connection cards, first use returning the current user's authorization request, and successful retry after that user's grant exists. - App checks use an isolated local fixture and a fake MCP provider. They do not use production app credentials. ## Risks - Intentional behavior change: workflows that used mentions to wake agents must use assignment, a bounded child task, or an explicit review request. - Already accepted queue entries retain their prior rules. An old entry can combine assignment or feedback with a later mention; its last reason cannot safely identify mention-only work. New mention requests create no run or deferred wake. - Personal apps with a shared health error remain discoverable. Actual tool use still requires the responsible user's grant and existing policy gates. - No database migration or public API schema change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact serving model ID and context-window size are not exposed in this session. - Live native-run verification used `gpt-6-astra` through the Codex provider. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3b4b270650 |
fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The shared ACP adapter engine records agent failures for operators. > - ACP providers can report a failure category, title, and detailed cause. > - Our patch kept only the category in the saved error, so an operator could not diagnose a failure when tracing was off. > - This pull request preserves redacted provider diagnostics in the run error, transcript, and structured run result. > - Operators can now inspect the provider message and any supplied request ID or stack trace after the run ends. ## Linked Issues or Issue Description Refs #13889 (the diagnostic gap; this PR does not update the bundled Claude version). Refs #14484 (related model-refusal classification; this PR retains diagnostics for all terminal failure categories). **What happened?** An ACP turn failed with only `ACP agent reported a terminal service failure.` The provider's title and details were available in memory but absent from the saved error and transcript. **Expected behavior** The run retains useful provider diagnostics even when raw tracing is disabled. Credentials remain redacted. A size limit must report truncation instead of silently removing the cause. **Steps to reproduce** 1. Run an ACP agent that returns an error-severity typed session failure. 2. Include an HTTP error, request ID, and stack text in its title and details. 3. Inspect the failed run with tracing disabled. Before this change, only the category survives. ## What Changed - Both pinned ACPX patches pass complete error text to the in-memory callback, so redaction happens before truncation. - The shared engine retains the sanitized category, title, and details in `resultJson.terminalSessionFailure` and includes the text in the run error and error transcript. - Diagnostics redact configured environment values even under arbitrary names, unknown launch-environment values, connection URL passwords, run credentials, and common credential syntax. Known boolean settings remain readable, while credential values are redacted even when embedded in other text. Diagnostics remove control characters and invalid Unicode. - Title and detail limits keep escaped transcript JSON below the server's chunk limit. Truncated fields include an omission count. The safe run-result projection preserves a byte-bounded diagnostic preview when the result exceeds its byte budget, with an explicit pointer to the full adapter-bounded run error and transcript. - The existing UI and CLI display the error. Diagnostics do not become assistant output. Issue continuation summaries and session-compaction prompts receive only the generic category, preventing provider text from becoming handoff instructions. Existing quota classification, warnings, timeout precedence, and control-channel failure precedence remain in place. - Regression tests cover real ACP child processes with both pinned versions in one-shot and persistent modes, credential redaction, request IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and database retrieval of oversized multibyte diagnostics. ## Verification - Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2 intentionally skipped, no pending or failing checks**. Includes typechecking, build, all Vitest shards, Runner checks, browser E2E, and the canary packaging/public-install dry run. - Greptile: **5/5** on this commit. Superagent security scan passes. All review threads are resolved. - Local verification passed: shared ACP engine suite (395 tests); real Claude ACP child-process and diagnostic regressions across both pinned runtimes and both execution modes; run retrieval and model-handoff regressions (59 tests); ACPX patch packaging (16 tests); full typecheck and build. Affected package typechecks and focused tests were rerun after review fixes. - The broad local `pnpm test:run` was stopped after review edits made its cached imports stale. Fresh targeted runs pass, including both affected server suites. Cold-build import failures were also rerun after dependency builds: chat integration (1,063 tests) and tool access (351 tests) pass. The final commit's complete CI matrix is green. ## Risks - Provider diagnostic text is untrusted. This change retains more of it in company-scoped run records. Redaction and size bounds apply before persistence. - Diagnostics are limited to fields the provider supplies. Old runs cannot recover discarded error text. - No schema migration, recovery-policy change, or new Telemetry or OpenTelemetry export. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da887ea3e9 |
fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path > - Paperclip manages AI agents that work on assigned tasks. > - The task composer lets a person choose an assignee, model, and effort for the next run. > - A Paperclip Runner agent can use Codex as its provider. > - The composer hid Codex effort for that agent because it checked only the older Codex adapter. > - The native Runner input also did not carry an effort choice to Codex. > - This pull request carries the chosen effort from the composer to each Codex turn. > - People can now select a supported effort and get the effort they selected. ## Linked Issues or Issue Description Refs #14322 **What happened?** The composer showed a model but no effort slider when the assignee used Paperclip Runner with the Codex provider. A task-level model override also did not reach the native Runner input. **Expected behavior** The composer shows effort choices for a known Codex model. The next native Codex turn uses the selected model and effort. **Steps to reproduce** 1. Open a task composer. 2. Select an agent that uses Paperclip Runner with the Codex provider. 3. Select a known Codex model such as `gpt-6-astra`. 4. Open the assignee and model picker. The effort slider is missing before this change. ## What Changed - Show known Codex effort levels for Paperclip Runner Codex assignees. - Save the task effort override in the native run input and send it to Codex on each turn. - Apply the task's merged model and effort overrides when the native run starts. - Apply a task model override for OpenCode Runner without changing the agent's provider. - Add Runner effort tests and desktop and mobile Storybook cases. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm build-storybook` passed. - `pnpm check:token-gates` passed. - Focused UI, server, Runner contract, and Codex driver tests passed. - The full CI test matrix, build, typecheck, and canary dry run passed on the latest head. ## Risks - Native Runner inputs add an optional Codex effort field to the current v5 input. Older inputs keep their previous behavior. - A known model rejects an effort that its catalog does not support. Unknown models do not show a slider. > This fixes an existing composer bug. I checked `ROADMAP.md`; it does not describe this bug as planned work. ## Model Used OpenAI Codex, GPT-6. The exact deployment ID and context window are not exposed in this session. The model used reasoning, code execution, and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com> |
||
|
|
7636966452 |
fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Mine inbox shows work that needs the current user. > - Failed-run rows used the latest run for every agent in the company. > - A failure from another user therefore appeared in Mine and its badge. > - Run list responses also omitted the responsible user needed to filter these rows. > - This pull request uses run ownership for personal failure routing. > - Users see their own failures and can still inspect company failures in All. ## Linked Issues or Issue Description **What happened?** An agent run started for one user failed. Its row and failure badge appeared in another user's Mine inbox. **Expected behavior** Mine and its badge include failed runs for the current responsible user. Other users' failures remain available in All and run details. **Steps to reproduce** 1. Use a company with two human users. 2. Create a failed or timed-out run attributed to the first user. 3. Open Mine as the second user. Before this fix, the failed run appears there and increases the badge. **Paperclip version or commit** Reproduced in regression tests on master at `24beb0057`. **Deployment mode** Authenticated deployment with multiple users. Tests also cover the local single-user board. Related prior work: #933 addressed inbox dismissal and badge consistency. No duplicate ownership fix was found. ## What Changed - Return `responsibleUserId` in normal and summary run lists. - Share one ownership rule across both inbox versions and client/server badges. - Select the latest run per agent before applying the ownership filter. This prevents old failures from resurfacing on shared agents. - Keep unattributed historical failures in the local board's Mine view. Hide them from authenticated users with no matching owner. - Keep company health alerts outside the personal badge, consistent with the client. - Document the routing contract and add page, badge, and database regression coverage. ## Verification - Red: the new badge cases failed with three company failures instead of one personal failure; eight Mine page cases failed across both inbox versions. - Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`, `ui/src/pages/Inbox.test.tsx`, `server/src/__tests__/heartbeat-list.test.ts`, and `server/src/__tests__/inbox-dismissals.test.ts`. - `pnpm check:token-gates` passes. - Agent calls on behalf of a user have two additional red-to-green API regressions. - Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also passes after the agent-call fix. - All CI test shards and browser tests pass on `243bfa681`. The duplicate local `pnpm test:run` was stopped after the CI test lanes completed; it did not finish locally. ## Risks - Authenticated users no longer receive unattributed legacy failures in Mine. Those failures remain visible in All. - The server badge no longer counts company health alerts, matching the existing client badge. - No migration, run state, retry behavior, or company access rules change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, terminal execution, and browser tools. The exact deployment variant and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
d172197117 |
feat(storage): add plain directory sync with conflict preflight (#14416)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs use workspace transport to restore and collect files. > - Some directories belong to the agent across tasks. > - Those directories need plain file transport without task Git state. > - A concurrent file edit must be detected before a merge changes any file. > - This pull request adds optional plain-directory sync and conflict preflight. > - Existing task workspace sync keeps its defaults. ## Linked Issues or Issue Description Refs #14325. This is the transport prerequisite for a replacement of its instruction revision design with current agent files. ## What Changed - Add an opt-in plain-directory transport mode to command and sandbox runtimes. - Add strict merge preflight for file edits, deletions, and directory changes. - Accept identical replay after an interrupted merge. Preserve competing changes. - Set the compiled OpenCode test executable to 0755, independent of the CI host’s file-creation mask. Preserve the original startup error if cleanup also fails. ## Verification - At `ced53ae532ce6966cad5a83575d83db61af98126`, all 212 targeted transport tests passed across workspace restore, remote managed runtime, SSH fixture, and execution-target sandbox suites. Adapter-utils typecheck passed. - Real isolated SSH retry fixture previously passed with `PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1`; stale deleted files remain absent while gitignored binary bytes survive. - Dependent PR #14420 passed real native local, legacy local, and native Daytona persistence E2E at `169fab46d5af21caa2269b4c1b29b69c933a6951`, which includes all production transport changes through `ced53ae53`; the subsequent two commits only fix the OpenCode test fixture. Nine tasks, three server restarts, and all cleanup checks passed. - A hosted OpenCode fixture failed twice at `ced53ae53`. Reproduced the failure locally and in Linux with `umask 0002`: the compiler created a group-writable executable, correctly rejected by the qualified launch boundary. Explicit 0755 permissions fix the test without weakening the production guard. The focused test and non-root Linux reproduction now pass under that same mask. - Before rebase, head `69e97de0475d34aac5d532e559a405eaf015fd2b` includes the deterministic fixture permission fix and preserves original bootstrap diagnostics. All production transport code is unchanged since the 212-test validation. Fresh Greptile review is 5/5 on this exact head with no unresolved findings. All 54 current-head checks passed, with two conditional skips. The full CI run completed successfully, including the previously failing OpenCode runner shard. - Merge validation on rebased head `c509d79dd190c5cb00dc65edfde209097ff21465`: all five commits are patch-identical to the reviewed branch. All 54 checks passed with two conditional skips, and fresh Greptile review is 5/5 with no findings. One retry cleared an npm archive 404 and a Cursor fixture timeout. ## Risks - New behavior is opt-in. Existing task snapshot behavior retains its defaults. - Generic strict merge preflight remains opt-in. The dependent agent-folder feature rebases changed paths before applying them to provide per-file last-sync-wins; it does not create a conflict-review queue. - This change adds no database migration, dependency, or UI. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and tool use assisted this change. Live provider validation used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
53aad90b9e |
fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path > - Paperclip coordinates agent work through execution adapters. > - Sandbox ACP sessions send ordered input through a remote file queue. > - A temporary provider 502 currently closes the session during an input upload. > - A lost response can occur after the sandbox has consumed the message, so a blind retry can duplicate input. > - This pull request retries gateway failures with the same sequence and drops consumed sequences at the receiver. > - The session can continue through a brief provider failure without repeating a tool call. ## Linked Issues or Issue Description **What happened?** A sandbox ACP run can fail with `ACP agent disconnected during request (connection_close, exit=null, signal=null)` when a provider input upload returns HTTP 502. The bridge destroys its local socket on the first failure and can discard the diagnostic before the proxy reads it. **Expected behavior** A temporary gateway failure should get a bounded retry. A lost response after successful delivery must not duplicate input or reorder later messages. Permanent failures must still close the session. **Steps to reproduce** 1. Run the real sandbox process bridge with an echo child and a local test runner. 2. Inject a provider 502 before preparation, after a chunk upload, or after final publication and consumption. 3. Send the next input message. Before this change, the connection closes instead of delivering it. Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and `502 sandbox`. Related work: #13287 covers shutdown after bridge loss; #13793 covers large launch envelopes. This change covers ordered input delivery within a running legacy ACP session. ## What Changed - Retry input uploads up to three times for recognized Daytona and Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms delays. - Give each upload separate temporary paths and discard already-consumed input sequences, including late publication from an earlier attempt. Clean failed attempts in the background without removing a published message or another attempt’s files. Cleanup cannot delay retries or shutdown. - Keep later input behind the retry. Stop queued input on permanent failure and flush a fixed diagnostic before closing the socket. Neither failure-diagnostic persistence nor shutdown-warning persistence can block teardown. - Add real-process regression tests for lost responses, late publication, retry exhaustion, immediate permanent failure, and diagnostic redaction. - Give accepted run-log file appends up to three seconds to drain before finalization computes the size, hash, and durable copy. Close the run handle to later appends. This waits only for file writes, independently of later DB progress or live-event persistence. If writes remain stalled, return null size/hash metadata and skip the final durable copy so the run can settle. Late writes cannot restart mirroring. - Preserve legacy comment attribution when final log size is unknown by reading existing entries within the unchanged 2 MB scan limit. Storage errors or a three-second read deadline return the evidence already read instead of failing the comment listing; pagination stops at the deadline. The deadline requests cancellation of the underlying local stream or S3 HEAD, GET, and response stream. A separate response timeout returns partial evidence even when filesystem I/O delays cancellation; late reads cannot append evidence or start another page. Each listing retains its existing batches of eight reads, without a shared admission cap that skips readable logs under contention. - Document the retry and log-finalization boundaries in the development guide. ## Verification - Final commit `347daa564b`: [Linux CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2) passed. Greptile Apex review 13 scored this commit 5/5 with no new findings; all 12 review threads are resolved. - The final CI run initially hit a Cursor test timeout and four Discord credential-lock contention failures. All five cases passed in isolation. The two failed shards and their aggregate gate passed on retry without a code change. Those intermittent failures are not claimed fixed by this PR. - `pnpm --filter @paperclipai/adapter-utils typecheck` passed. - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests passed on the final implementation, including 21 new regressions. The original three fault-injection cases failed before the fix. - The regressions cover failed and indefinitely stalled cleanup, Cloudflare gateway responses and retry exhaustion, permanent errors that must not retry, and teardown while failure logging remains indefinitely stalled. Seven Apex regression cases failed before the review fixes. Adapter-utils typecheck and build passed again after the final review change. - `pnpm exec vitest run server/src/services/run-log-store.test.ts server/src/services/run-log-store-cancellation.test.ts`: all 25 tests passed, including four new regressions that failed before the finalization fix. They cover delayed and failed appends, late-write admission, agreement between the local bytes/summary/durable copy, and a stalled append that exhausts the three-second budget. The timeout case verifies unknown metadata, no final upload, and no mirror restart after late completion. New cancellation tests use the real AWS SDK against a local HTTP server. They verify that stalled HEAD, GET, and response-body connections close on abort and that a subsequent read succeeds. Local range and already-aborted read cases also pass. - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t 'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14 targeted tests passed. The null-size reader case, both storage-error cases, the stalled-read case, and the cancellation/concurrent-listing cases failed before their fixes. The new regressions verify that timed-out reads are cancelled, subsequent listings recover, and two concurrent listings both retain their attribution markers. A read that ignores cancellation still returns partial evidence at three seconds and cannot resume pagination when it finishes; this regression failed before the response-timeout fix. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/server build` passed after the response-timeout change. - Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR; the affected packages were rechecked after review fixes. - Full local `pnpm test:run` failed in the general-server group: 511 files passed, 40 failed, and 158 were skipped. Failures include embedded PostgreSQL initialization, read-only cache directory renames, a macOS long-path fixture, and a workspace exposure assertion. The PostgreSQL, cache-permission, and long-path failures also reproduce with both changed implementation files restored to baseline commit `24c58e479a`. The exposure suite passes in isolation both on baseline and the fixed branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later local test groups were not reached. - An earlier CI run hit the Telegram retry-timing failure fixed upstream in #14501. The branch includes that master fix. The selected recovery test passed against a fresh, migrated PostgreSQL 16 database. The embedded PostgreSQL runner is unavailable on this Mac; the isolated database was stopped and removed afterward. - No live agent turn was replayed. The tests use local child processes and injected provider failures. ## Risks Retries are restricted to recognized Daytona SDK and Cloudflare bridge gateway-error messages, which survive plugin RPC serialization. Other errors fail immediately. Temporary upload paths are now unique for all command-managed queue writes. Receiver sequence checks prevent duplicate input; retries do not restart an agent turn. Cleanup and failure logging are nonblocking and best effort; session teardown remains the final cleanup boundary. Log finalization now drains accepted local file writes for at most three seconds and ignores later appends on the closed run handle. A timeout leaves final size/hash unknown and skips the final durable upload; an existing partial mirror may remain available, but it is not claimed as a verified final snapshot. It does not wait for later DB progress or live-event persistence. Optional attribution keeps partial evidence when a read fails or times out. Cancellation closes S3 requests and response streams. Local filesystem I/O may finish after the caller deadline, but a late read cannot change the returned evidence or continue pagination. Later listings can retry after storage recovers. There is no schema, authentication, or permission change. Revert this commit to restore the previous behavior. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; targeted tests pass and full-suite limitations are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
8751e2de46 |
fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task board shows which recovery actions are active. > - A native run can stop while a person must repair its workspace. > - The board previously called that state “Recovery in progress.” > - The label implied that work would continue without operator action. > - This pull request derives the label from the recovery owner and live continuation. > - Operators can distinguish scheduled recovery from a repair that needs attention. ## Linked Issues or Issue Description **What happened?** A blocked task showed “Recovery in progress” after native finalization stopped and no automatic continuation remained. **Expected behavior** Show “Recovery needed” for an idle board repair or when no live recovery path exists. Show “Recovery in progress” while the recorded continuation can run, including an explicitly admitted export whose exact callback is executing even if the old recovery action remains board-owned. **Steps to reproduce** 1. Complete a native run whose workspace export cannot be recovered automatically. 2. Inspect the task recovery action and badge. 3. Compare the board-owned action with the old “Recovery in progress” label. Related work: #14314 serializes native workspace finalization and fences stale recovery outcomes. This change reports the recovery action that currently owns the task. ## What Changed - Show “Recovery needed” for board-owned active-run recovery unless the exact native export callback is positively verified as executing. - Require the native continuation run and a live or future continuation before showing progress. - Project native activity from the exact company, source issue, and run. Include active workspace export while the original heartbeat remains failed. - Require an executing callback before using a running export row as evidence. Preserve activity for long exports and clear it when the callback joins. - Give native resume its own card explanation. Preserve ordinary watchdog observation behavior. - Remove the redundant ownership sentence from all six recovery-card explanations that used it. - Document the labels and add regression cases for stopped, scheduled, and active recovery. ## Verification - Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests and token gates pass. No UI occurrence of the removed sentence remains. `pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54 successful checks, two optional Storybook checks skipped). Greptile is 5/5 with no open review threads. The duplicate local `pnpm test:run` was stopped after the full CI suite passed; it did not complete locally. - Original regressions: seven failures before the change, then 52 focused cases pass. - Review regressions: seven UI failures and nine database failures before the follow-up. All 75 database/API recovery tests, 167 UI tests, and 18 workspace lifecycle/finalizer tests pass. A further three RED cases cover explicit board retry activity; one RED case rejects orphaned running export rows after controller loss. Wrong company, issue, run, service, phase, and completed-operation cases remain inactive. - Final recursive typecheck, production build, token gates, and complete local suite coverage pass. Embedded PostgreSQL startup/socket failures passed in isolated retries with the canonical test environment; no expected behavior was weakened. The recovery regression added during the earlier full run passed in its final complete 75-case file. - Before the copy-only follow-up, all 56 CI checks passed on `c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no unresolved review threads. The final native-activity staging repeat passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`. - Verified on a separate staging instance: a real failed Daytona workspace export retains its board-owned repair action and displays “Recovery needed” in the task list. The repair card remains actionable without starting another provider turn. - Real staged export-only repair: the actual task list showed “Recovery in progress” while the original run had a positively identified running export operation, then Done after exact copyback of all 20,000 nonce-bound files. The accepted result and full provider session/turn/terminal envelopes remained unchanged. All 17 independent final checks passed; the browser downloaded the exact 19-byte result. The separate fixture was cleaned up with independent provider-absence verification. A control transport process restarted during repair; it did not submit another provider turn. - Final deployed-source repeat on `d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral Daytona allocation, actual browser export repair, and a saved full activity projection referencing the exact executing export operation with no scheduled retry. The task list showed recovery in progress, then Done; all 21 final checks passed, including 20,000 exact host files, unchanged provider provenance, and provider deletion only after committed copyback. The downloaded 19-byte result matched independently. This integrates #14334; no source change was required here. ## Risks - The label depends on the persisted recovery action. A separate runtime defect can still stop work; this change makes that condition visible. - Future recovery kinds must supply a valid continuation path before they can display progress. ## Model Used OpenAI Codex, based on GPT-6, with code execution, browser testing, and subagent tool use. The runtime does not expose an exact serving model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbc5132e6c |
fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters own local or remote provider processes. > - Run cleanup must know when the final provider process has stopped. > - A remote timeout or a lost transport does not prove that the provider stopped. > - Retried provider invocations also make an earlier stop signal stale. > - This pull request adds a verified final-invocation stop callback before workspace restoration. > - The dependent instruction revision change uses that boundary to preserve private instruction edits safely. ## Linked Issues or Issue Description **What happened?** Adapter completion did not expose a reliable point between provider shutdown and workspace restoration. Cleanup could lose provider-written files, or treat a remote timeout as proof that a process stopped. **Expected behavior** Cleanup runs once after the final provider invocation has a verified stop receipt and before workspace restoration. An incomplete remote command keeps collection pending. **Steps to reproduce** 1. Run a remote provider that returns a timeout without a numeric exit code. 2. Let the adapter return or retry the provider. 3. Attempt to collect provider-written files during cleanup. The adapter has no verified final-process boundary to use. This is the prerequisite for the stacked canonical instruction revision pull request. It has no database or UI dependency. Related #13291 verifies remote termination for later recovery; this change exposes the earlier adapter-owned stop boundary before workspace restoration. The scopes do not duplicate each other. ## What Changed - Add a stop callback to adapter execution context and a final-invocation fence. - Confirm local child closure and complete remote exit receipts. Reject remote timeouts, missing exits, transport failures, and SSH exit 255 as stop proof. - Invoke collection once before workspace restoration in eight CLI adapters and at the confirmed ACP stop boundary. - Preserve the stop observation when a later log flush fails. - Run bridge and workspace cleanup in `finally` even when collection rejects. ACP records a safe error without exposing a raw filesystem path. Grok keeps collection errors separate from workspace restore failures and preserves completed provider results when both cleanup steps fail. - Correct the existing Cursor test shell fixture so bounded remote file reads run against real fixture files. ## Verification - Three stop-boundary regressions failed before the callback implementation and passed after it. - Independent prerequisite branch: 377 tests passed across 24 adapter, process-target, and ACP suites. Two added collector-rejection tests failed before the cleanup fix and passed after it. - A third regression reproduced Grok misclassifying a collection failure as failed workspace restoration. Two additional cases covered completed and failed provider turns when collection and restore both fail. The Grok and restore-classifier suites passed 48 tests. - All nine affected package typechecks and affected package builds passed; Grok checks passed again after its classification fix. - Integrated instruction branch: native local, legacy local, and native Daytona each passed three browser tasks with exact persisted bytes, fresh-task readback, history/restore, and explicit conflict resolution. Unchanged warm Daytona passed three turns. Legacy Codex passed all five checks again after the exception-safe cleanup fix. - Alternate staging passed the same three-task native Daytona flow: exact stopped-run save, independent downloaded readback, browser history/restore, and explicit resolution of a real concurrent edit. The deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which covers the initial adapter callback. Later Grok cleanup failures are qualified by the adapter tests above. - The final combined native Daytona flow passed again on deployed `2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved ordinary instruction edits, independent readback, History/Restore, concurrent board conflict, and explicit candidate resolution. This native staging flow does not claim to exercise the Grok adapter. ## Risks - An unverified remote stop intentionally does not trigger collection. A later controller with verified stop evidence must recover it or report the copy unavailable. - The callback is optional. Callers that do not register it retain their existing behavior. - The callback runs before workspace restoration and can delay cleanup if its caller does not bound its own work. The dependent instruction collector uses bounded reads and retries. A rejected callback still permits bridge and workspace cleanup; it cannot claim an instruction save. ## Model Used OpenAI Codex, GPT-6, with tool use and code execution. The runtime does not expose the exact deployment variant or context-window size. Multiple Codex agents implemented and verified the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b38db9622 |
fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed runs copy a selected workspace to an execution environment and restore its changes. > - Git snapshots select the files for that copy and for later recovery. > - A fixed output limit stops large generated trees before the run can start. > - Increasing the limit still keeps the complete filename lists in memory. > - This pull request stores those lists and merge baselines in disk manifests. > - Large snapshots can now complete with bounded filename buffers and explicit failure handling. ## Linked Issues or Issue Description Refs #14194. This is the streaming follow-up to the merged 32 MiB limit fix. Related: #13619 and #11621 cover workspace scan admission and demand. This change keeps the shared scheduler and changes the snapshot data path. ## What Changed - Stream changed, untracked, deleted, and ignored paths through the shared scheduler and the standalone adapter path. - Use SQLite manifests for file selection, duplicate removal, ignored-path lookup, baseline capture, and merge lookup. - Set a configurable 30-minute snapshot deadline. Keep the existing interactive scan deadlines. - Wait for each child process and pending sink write before removing temporary storage after failure or cancellation. - Use NUL archive lists and bounded deletion batches. Preserve unusual names, explicit selection, nested repositories, and source root checks. - Store manifest references in native recovery format v2. Check their location and digest before recovery reads. Keep v1 descriptors readable. - Remove temporary manifests at lifecycle completion. Use fixed-size temporary copy names for long basenames. - Admit each manifest with a SQLite page allowance based on current disk capacity. Keep a configurable free-space reserve and fail explicitly when either limit is reached. - Preserve a host file that replaces a directory deleted by the sandbox, and continue the rest of the restore. ## Verification - Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54 successful checks/statuses and two skipped Storybook jobs. No checks failed or remain pending. - [CI passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368): typecheck, build, all test shards, E2E, Rust checks, and the aggregate verify job. - [Greptile is 5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338) on the current head. All four review threads are resolved. Security checks passed. - 229 focused tests passed across Git sync, runtime staging, merge, manifest integrity, native recovery, and the scheduler (214 adapter/runtime tests and 15 scheduler tests). - A real 40,000-file fixture produces 43,428,890 filename bytes. The original standalone and scheduled scans fail. The new test passes all four filename paths, complete staging, exclusion of late files, unusual names, and deletion replay. - Recovery tests reject changed bytes, symlinks, and paths outside the controller state directory. Adapter-utils typecheck passed. - A test executor returned buffered output and caused two retry integration failures. The fixture now uses the shared streaming scheduler. All 13 tests passed with `corepack pnpm exec vitest run server/src/__tests__/heartbeat-project-repositories.test.ts`. The same CI shard now passes. - Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally. Each full local command hit SIGKILL/exit 137 in the 4 GiB container. These local commands did not pass. The current-head CI gates above provide the full verification. - A real-Git disk-capacity regression confirms a typed failure and removal of the incomplete manifest. Repeated writer attempts cannot exceed the permitted page count. - Follow-up real Daytona and separate staging qualification passed with the related archive validator (#14315) and exact-owner finalization fix (#14314). Three successive turns copied back all 60,000 files with 39,828,890 filename bytes and five unusual names. Independent host inventories verified every file and the pinned Git HEAD. Native, provider, session, and process identities stayed fixed; no retry remained. The task reached Done, and its browser-downloaded final proof matched exactly. The reusable regression is #14316, including an assertion of the effective environment idle policy. ## Risks - SQLite manifests use disk space. Each receives one quarter of the available capacity above the host reserve at creation. The reserve defaults to 256 MiB and has a 64 MiB configuration minimum. Disk capacity, filesystem quotas, per-path limits, Git resource use, and execution deadlines remain limits. - Each path and sink chunk has a 64 KiB limit. SQLite connections use a 1 MiB page cache. Invalid or incomplete records fail explicitly. - Restore transport keeps fixed and configured archive exclusions. A remotely created Git-ignored file can be transferred, but the host merge excludes it through the manifest. - Provider archive buffers, Git and tar memory, repository metadata, legacy v1 arrays, and the separate referenced-source resolver retain their own limits. Existing provider safety validators still buffer textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64 MiB. These separate transport limits can stop a sufficiently large restore before merge. This change does not claim bounded total process memory or unlimited transport size. - New descriptors use v2. Existing v1 recovery remains supported; a downgrade cannot read v2 descriptors. ## Model Used OpenAI GPT-6 through Codex. The exact deployment ID and context limit are not exposed in this run. The agent used code editing, terminal execution, tests, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bacc0e6a98 |
docs: remove automatic agent escalation from coordination skill (#14188)
## Thinking Path > - Paperclip manages work for AI agents. > - The coordination skill tells agents how to handle blocked work. > - The skill directs blocked work to managers and other agents. > - Those agents may lack the same access or authority. > - These extra assignments delay the required human action. > - This pull request removes automatic escalation advice. > - Agents must identify the missing capability and use the correct approval or human-input path. ## Linked Issues or Issue Description **Issue type** Incorrect information. **Where is the issue?** The critical rules in `skills/paperclip/SKILL.md` and blocker guidance in `skills/paperclip/references/api-reference.md`. **What's wrong?** The skill tells agents to escalate through `chainOfCommand`, ask another agent for help, and avoid human help. A manager title does not grant permission to fix a connection or complete an administrator action. **Suggested fix** Remove blanket escalation and agent-first rules. Keep normal delegation when the recipient has a concrete capability for a bounded task. Use saved human-input interactions or existing connection and approval flows for human-only actions. Searches for open PRs with “escalation”, “chainOfCommand”, and “ask another agent” found no duplicate skill change. Recovery-routing PRs change server behavior, which is outside this change. ## What Changed - Remove the chain-of-command escalation rule and both repeated agent-first directives. - Replace manager handoffs in the API reference with direct blocker handling. - Keep reporting fields, normal delegation, approval gates, and the ban on bypassing permission denials. - Keep the ban on cancelling cross-team tasks. Request a decision instead of automatically assigning the task to a manager. ## Verification - Final head `1483e82cc3fe10a7c910230b3779635d6d54c162`: Greptile 5/5, no outstanding findings, all CI checks pass (optional Storybook jobs skipped). - Created a fresh worktree at `.worktrees/skill-blocker-guidance` from the current master commit. - Ran `git diff --check`: passed. - Ran Node assertions against both documents: passed. Removed directives are absent. Human-input and capability-based delegation guidance is present. - Checked that checkout conflict, approval, dependency, and normal delegation instructions remain. - Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. None could run because this environment has no `pnpm` executable. - Refreshed both generated capability inventories and their derived contract after skill heading positions changed. Both generator and live-inventory checks pass. The eval and MCP baselines remain unchanged. - Ran `node --test packages/paperclip-runner/scripts/check-capability-inventory.test.mjs`: all four tests pass. - Ran `node packages/paperclip-runner/scripts/check-capability-inventory.mjs` and `node packages/paperclip-runner/scripts/generate-capability-contract.mjs --check`: both pass. - Added explicit requester routing for agent and human scope questions. Focused Node assertions pass. ## Risks - Agents can request human input earlier for actions that require human authority. - Existing runs or installed copies can keep old skill text until refreshed. - This change does not alter server recovery routing or permission checks. ## Model Used OpenAI Codex, with reasoning, tool use, and shell execution. The runtime does not expose a verifiable exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f14d26123 |
fix(runtime): allow bounded large untracked workspace snapshots (#14194)
## Thinking Path > - Paperclip manages AI agents and their work. > - Remote runs need a snapshot of the task workspace before the agent starts. > - The snapshot lists untracked filenames through the shared Git scan scheduler. > - A generated directory with a few thousand long filenames can exceed the 1 MiB output limit. > - This stops setup and prevents the agent from continuing its task. > - This change gives that listing a 32 MiB bound and keeps the explicit file snapshot. > - Normal generated trees can now pass setup, while larger snapshots still fail at a finite limit. ## Linked Issues or Issue Description Refs #11572 and #12214 for the existing bounded scan and ignore-scan protections. **What happened?** An agent continuation failed during workspace setup with `Workspace Git scan exceeded its output limit`. A real Git fixture reproduces the untracked-file path: 5,000 long filenames in one generated directory exceed its 1 MiB output limit. **Expected behavior** The workspace snapshot must support ordinary generated trees with thousands of files. It must keep a finite output bound and select explicit files before staging. **Steps to reproduce** 1. Commit a base file in a Git repository. 2. Add 5,000 untracked files with long names in one new directory. 3. Call `readGitWorkspaceSnapshot` with the normal scan limits. 4. Observe the output-limit error before this change. **Paperclip version or commit** Base commit: `640dee1`. **Deployment mode** Remote sandbox execution from source. ## What Changed - Increase the untracked-file snapshot output bound from 1 MiB to 32 MiB. - Keep explicit file selection, the shared scheduler, the timeout, and the other scan bounds. - Test that 40,000 long filenames above the old 8 MiB bound reach the snapshot. Reuse those files with deeper paths to prove that output above 32 MiB still fails. - Test that files created after the snapshot, including an ignored secret, stay out of the overlay archive. - Check the workspace root identity and reject selected paths with symlinked parent directories before upload. - Test root replacement after path resolution, including root-level selected files. - Preserve existing workspace root aliases by capturing the resolved root before snapshot selection. Test that later alias retargeting cannot change the archive contents. - Stop staging on permission and I/O errors; continue to allow missing files. - Test these failures and preserve selected symlink entries. - Accept valid case-renamed directories by checking ancestor file types. A modeled case-insensitive regression failed before this correction and now passes. - Give the 40,000-file fixture enough time to remove its files. - Document the larger bound and staging behavior. ## Verification - Red: the 5,000-file regression failed with `stdout maxBuffer length exceeded` before the fix. - A separate check through the real server scheduler reproduced `workspace_git_scan_output_limit` on the original code. The revised code selected all 5,000 files. - The late-file regression failed against the first PR revision because the archive contained `drafts/late.secret`. It passes with the final explicit-file approach. - The staging regressions failed before the review fix: a substituted parent directory and permission/I/O errors were accepted. All three cases now stop before upload. - Green: 134 tests passed across `git-workspace-sync.test.ts` and `sandbox-managed-runtime.test.ts` with Vitest 4.1.11. This includes complete selection above 8 MiB and rejection above 32 MiB. - `pnpm --filter @paperclipai/adapter-utils... typecheck` passed after the revision. - The module-boundary check and `git diff --check` passed. - `pnpm -r typecheck` and `pnpm build` stopped in the Rust runner steps because this environment has no `cargo` executable. - The full `pnpm test:run` attempt ended with `SIGKILL` during the general server suite. It did not finish. That full-suite result belongs to the earlier revision. Fresh checks are required for this revision. - The new regression fails at the old 8 MiB bound with `stdout maxBuffer length exceeded`. All 134 focused tests pass with the 32 MiB change. - The affected typechecks and module-boundary check pass. The full local typecheck requires Cargo, which is absent in this environment. - Final verification for `b43e9bc95e55382c6a9bfe200487c164770c8be8`: 54 successful checks/statuses and two skipped Storybook checks. No checks remain pending or failed. - The [CI run](https://github.com/paperclipai/paperclip/actions/runs/36315569674) passes on attempt 2. The first attempt had one unrelated preview-fixture readiness timeout. That exact test passed locally; its CI shard passed on the single rerun. - [Greptile reports 5/5](https://github.com/paperclipai/paperclip/pull/14194#issuecomment-5852045501) on this revision. All review threads are resolved. - [Security review accepts the documented memory tradeoff](https://github.com/paperclipai/paperclip/pull/14194#discussion_r4115149465) for this finite mitigation. The separate streaming follow-up will remove full-list buffering. - This revision also passes affected local typechecks and the module-boundary check. Full local typecheck/build stop because Cargo is absent. The full local Vitest attempt was stopped after about 18 minutes once all remote gates passed; it did not complete locally. ## Risks - Each untracked-file scan can buffer up to 32 MiB instead of 1 MiB. The scheduler still limits concurrent scans and execution time. - Snapshots above 32 MiB still fail with the existing error. Tracked and ignored-file scan bounds stay unchanged. - A workspace that replaces a selected path’s parent with a symlink now fails staging. - Detailed logs from the reported host were unavailable. The exact command that exceeded its limit on that host is unconfirmed. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The runtime does not expose a more specific model build or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5ee9e751fb |
fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds its operation or byte limits. Reserve task and completion tools, then expose bounded discovery and call tools for large catalogs. Fetch oversized schemas in reauthorized chunks without blocking later search results. Retain task ownership, work-mode restrictions, pinned assignments, current gateway authorization, approvals, and audit. Small catalogs stay direct. Validation: 46 focused server tests, two Runner capacity tests, local typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b2e9e82f05 |
fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The execution service owns each run and saves messages sent while it runs. > - Interrupt must stop the current executor before it delivers those messages. > - Remote Grok commands did not register the host cancellation control. > - A cancelled task run could still write Done and prevent queue recovery. > - This pull request connects remote cancellation and revokes cancelled run writes. > - Saved input can use the existing queue admission rules after verified cleanup. ## Linked Issues or Issue Description **What happened?** Interrupting a queued message marked a remote Grok run cancelled before its sandbox stopped. The old run could still post a reply and mark the task Done. Its saved follow-up remained deferred behind execution recovery. **Expected behavior** Stop revokes run write authority and waits for verified termination. Saved messages remain durable and enter one successor through normal admission after cleanup. **Steps to reproduce** 1. Run a task with `grok_local` in a remote sandbox. 2. Send a follow-up and use Interrupt while the command runs. 3. Let the old command attempt a task status update after cancellation. 4. Observe the task disposition and the saved message queue. **Paperclip version or commit** The gap is present in master at `d3e0f0a238`. **Deployment mode** Authenticated server with a Daytona sandbox. Related work: #14028 and #14046 handle bounded continuation. #13291 covers infrastructure interruption and verified remote cleanup. #13332 addresses atomic recovery holds. This change handles direct Grok operator cancellation and stale task writes. ## What Changed - Register remote Grok cancellation before preparation. Keep command ownership until the host confirms sandbox termination. - Reuse the sandbox cancellation boundary for the direct CLI invocation. Reject fresh attempts after cancellation and preserve workspace restore failure evidence. - Reject writes from cancelled task JWTs and runs with a pending stop. Preserve diagnostic reads and existing conversation error codes. - Recheck run authority under a database lock before task updates and interaction responses commit. - Preserve authorized handoffs that stop their own run. Only the server-issued stop receipt for that request permits the final task update. - Add tests for hung commands, unverified stops, early cancellation, copy-back failures, late Done, late interaction responses, authorized handoffs, exact lease receipts, and one queue successor across concurrent restart sweeps. - Document the cancellation and write-authority contract. ## Verification - Targeted adapter, cancellation-boundary, authentication, queued-message, interaction-service, and activity-route tests passed. The expanded run passed 214 tests; one new test had an incomplete fixture. After correcting the fixture, all 8 selected follow-up cases passed. - `pnpm -r typecheck`: passed on `179c86caf1bf0d89914a503d46e24af7e4b8c557`. - `pnpm build`: passed on the same commit. - `pnpm test:run`: the general-server group completed with 13,521 passed, 99 skipped, and 18 failed tests. It then stopped, so the remaining local groups did not run. Five Slack, email, and wake-batching failures passed on focused reruns after correcting the local environment. The remaining 13 failures reproduce as `EACCES` on rename in unchanged skill-cache code on macOS. Two custom-image suite setup hooks also failed to start embedded PostgreSQL after the machine exhausted shared-memory slots; all 31 tests in that file passed on rerun after the local resource issue was resolved. CI covers all test groups. - CI: 53 checks passed and 2 were skipped on the latest commit, including the aggregate verification gate. The last server shard passed on its single rerun after a preview-server startup timeout. The affected file also passed locally with 28 passed and 3 skipped. - Greptile: 5/5 on the latest commit. Both review threads are resolved. - No live deployment or staging task mutation has been performed. ## Risks - Stopping the sandbox can prevent file copy-back. The result preserves workspace restore failure evidence; termination does not imply restored files. - If provider termination fails, the adapter keeps ownership of its outstanding command and does not acknowledge Stop. - The write restriction now applies to ordinary cancelled tasks. Reads remain allowed. Task and interaction checks add a shared run-row lock to agent mutations. An exact server-issued receipt permits the task request that stopped its own run to complete its handoff. - Existing terminal tasks are not reopened automatically. An operator must correct a historical late Done before its saved queue can continue. - No schema migration or UI change. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |