mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
dffc2b3ca1b9e88fa21cb17493083e682dffd1ca
1394
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
dffc2b3ca1 |
fix(claude-local): skip expired credentials file when reading the Claude token (#13505)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents on the Claude Local adapter run on a local Claude Code
subscription. A user connects that subscription from Agent → Harness /
Runtime → "Connect account".
> - The server verifies the login with `readClaudeToken` in
`packages/adapters/claude-local/src/server/quota.ts`. It reads
`~/.claude/.credentials.json` first and consults the macOS Keychain only
when no file is present.
> - On macOS the Claude CLI refreshes the Keychain item, not the file. A
leftover credentials file keeps an expired token forever, and the reader
ignores `claudeAiOauth.expiresAt`.
> - The stale file shadows the live Keychain login. The usage check
fails and the user sees "Could not verify the local subscription"
although `claude auth status` reports a valid login. Running `claude
auth login` again does not help.
> - This pull request skips a credentials file whose token has expired,
so the reader falls through to the Keychain or returns null.
> - The benefit is that a valid local Claude login connects on the first
try, and a dead token is never sent upstream.
## Linked Issues or Issue Description
No public issue exists for this bug. Description follows
`bug_report.yml`.
### What happened?
Agent → Harness / Runtime → "Connect account" → Claude (Subscription) →
Connect failed with:
> Could not verify the local subscription. Run claude auth login in a
terminal on the machine running Paperclip, then try Connect again.
`claude auth status` on the same machine reported `loggedIn: true`,
`authMethod: claude.ai`, `subscriptionType: max`. The Keychain item
`Claude Code-credentials` held a fresh token. `POST
/api/companies/:id/ai-connections/local/check` returned
`{"status":"sign_in_required"}`.
A leftover `~/.claude/.credentials.json` (written weeks earlier) held an
access token that expired the same day it was written. `readClaudeToken`
returned that token. `fetchClaudeQuota` got a non-OK response from
`/api/oauth/usage`, and the route threw the generic 422.
### Expected behavior
An expired credentials file must not block a valid login. The reader
skips the dead token and falls through to the Keychain. Connect
succeeds.
### Steps to reproduce
1. On macOS, sign in with `claude auth login` (credentials land in the
Keychain).
2. Place a `~/.claude/.credentials.json` with
`claudeAiOauth.accessToken` set and `claudeAiOauth.expiresAt` in the
past.
3. Open an agent → Harness / Runtime → Connect account → Claude
(Subscription) → Connect.
4. Before this change: the "Could not verify the local subscription"
error appears. After: the connection is created.
### Agent adapter(s) involved
Claude Code (`@paperclipai/adapter-claude-local`)
### Operating system
macOS (Keychain-backed credentials). On Linux the file is the live
store; an expired file token now returns null instead of a failing
request, so the user-facing message is unchanged.
## What Changed
- `packages/adapters/claude-local/src/server/quota.ts`:
`parseClaudeCredential` now returns the token plus `expiresAt` (epoch
ms) when the file records one. `readClaudeTokenFromFile` returns `null`
for a token whose `expiresAt` is in the past, so `readClaudeToken` moves
on to the next candidate (second file name, then Keychain when
`allowKeychain` is set). Files without an `expiresAt` keep the old
behavior. `parseClaudeCredentialToken` (used for the Keychain payload)
is unchanged in behavior.
- `packages/adapters/claude-local/src/server/quota-keychain.test.ts`:
three new cases — expired file falls through to Keychain; expired file
with no Keychain access returns `null`; a file with no expiry is still
accepted.
## Verification
- `pnpm --filter @paperclipai/adapter-claude-local exec vitest run
src/server/quota-keychain.test.ts` → 7 passed (4 existing + 3 new).
- `pnpm --filter @paperclipai/adapter-claude-local exec tsc --noEmit` →
clean.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/local-ai-credentials.test.ts` → 9 passed.
- Manual, on the affected machine: with the stale file in place, `POST
/api/companies/:id/ai-connections/local/check` (provider `anthropic`,
method `subscription`) returned `sign_in_required`; with the stale file
removed it returned `ready`. This change makes the first case behave
like the second without touching the file.
## Risks
- Low risk. The only behavior change is for a credentials file that
carries a numeric `expiresAt` in the past. Such a token is already
rejected upstream, so the change removes a guaranteed failure rather
than a working path.
- Clock skew: a machine clock that runs ahead of real time could treat a
token as expired slightly early. The fall-through then reads the
Keychain (macOS) or returns null, which triggers the same "sign in"
message the user already sees for an expired token.
- Keychain payloads are not expiry-checked in this PR. The CLI refreshes
that item itself, and `getQuotaWindows` already falls back to the CLI
`/usage` probe when the OAuth call fails.
## Model Used
- Claude — `claude-fable-5-1` (Claude Fable 5.1) via Claude Code, with
extended thinking and tool use (shell, file edit). Root cause found by
reproducing the server's read path against the local credential file and
Keychain.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
d49f168381 |
fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must publish generated files so users can inspect their results after a sandbox stops. > - Legacy sandbox bridges blocked attachment listing and could not carry multipart binary uploads through the queue transport. > - The native runner has a separate verified file registration path that needs the same durable result. > - This pull request repairs legacy binary transport, makes native download receipts explicit, and reveals new outputs in the task Artifacts tab. > - Users can open generated files from either runner without a transport flag change. ## Linked Issues or Issue Description Refs #13355 for the existing native file publication path. Related filename fixes: #2615 and #4788. Related sandbox persistence work: #13376. This change repairs attachment delivery through the existing API; it does not add workspace persistence. **What happened?** The upload helper first lists task attachments to avoid duplicates. Both legacy bridge allowlists rejected that GET request with 403. A direct multipart upload also failed: the queue bridge accepted only JSON, excluded attachment uploads, and converted bytes to UTF-8 text. Enabling HTTP/2 alone did not fix the missing listing route. These failures occurred before attachment storage. **Expected behavior** Both runners can publish a workspace file, register its work product, bind it to a response, and return a working download. The file stays accessible after sandbox deletion. A new output opens the task Artifacts tab. The agent receives accurate errors and decides how to retry or report a failure. **Steps to reproduce** 1. Run a legacy agent in Daytona with the duplex bridge disabled. 2. Invoke the bundled upload helper with Bash on a PNG or PDF. 3. Repeat with the duplex bridge enabled. 4. Register the same file through the native runner with generic API tools disabled. 5. Retry registration, delete the sandbox, and compare the downloaded bytes with the original file. **Paperclip version or commit** The failing baseline was `f2c5e54dc`. This branch is rebased onto `6cfe4acff`. **Deployment mode** Source checkout with a local API and real isolated Daytona sandboxes. ## What Changed - Allow authenticated attachment listing, upload, and content download through both legacy bridge transports. - Add optional base64 body encoding to queue envelopes. Preserve the existing UTF-8 contract when the encoding field is absent. Decode binary bodies before forwarding them. - Preserve multipart headers. Bound raw bytes, encoded envelopes, and in-flight reservations. Retain timeout and uncertain-write behavior. - Preserve helper deduplication and return structured uncertain-write failures. Document explicit Bash invocation in live skills. - Add attachment IDs and content/download paths to native registration receipts. Reuse verified local and remote file reads, attachment storage, work-product registration, and response binding. - Preserve Unicode upload filenames and provide a valid Content-Disposition header. - Open the task Artifacts tab when new stored outputs arrive, including a closed desktop panel or mobile drawer. Deduplicate upload and registration events by object ID. Preserve manual selection on refetches, edits, and panel remounts. - Remove task artifact filters, the company Artifacts footer link, and the unassigned group heading and timestamp. ### Screenshot  This is the local display fixture. The image was generated separately and published through the attachment and work-product APIs. ## Verification - Post-rebase `pnpm -r typecheck` and `pnpm build` pass. - The post-rebase local `pnpm test:run` passed 12,369 tests before one existing conversation reset test timed out; all 33 tests in that suite pass when rerun with isolated test configuration. The aggregate command stopped before its remaining groups. GitHub runs the complete suite in separate shards. - All [GitHub verification checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350) pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks and two configured skips. The native-session recovery assertion initially raced its fire-and-forget Sentry report; all 13 tests pass locally, and the same-commit CI rerun passes all 170 suites (3,079 tests). - Live post-rebase Daytona: all three file-delivery tests pass. They cover the real Bash helper with the queue bridge, the helper with HTTP/2, and native `register_deliverable` with generic API tools disabled. - Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate registration, response binding, authorization controls, and byte-for-byte downloads after sandbox deletion. - Local focused coverage includes transfer bounds, malformed encoding, interrupted transfers, remote path containment, and native file verification. The attachment route suite passes all 32 tests, including an eight-case filename-header matrix for Unicode and special characters, inline and forced downloads, and full and partial responses. - Browser verification confirms image previews, persisted downloads, automatic Artifacts selection, and preserved manual selection after edits and reloads. Desktop/mobile component coverage passes. The latest UI cleanup passes its 10 affected tests and token gates. - Coverage limit: the Daytona tests call the real helper and native registration path directly. They do not replay a complete model-led image-generation task through the browser. Live command (requires a configured Daytona credential): ```sh PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts ``` ## Risks - Binary queue bodies use more memory because base64 adds encoding overhead. Transfer and process limits must remain aligned. - An interrupted write can have an unknown result. The bridge reports this state and preserves stable retry identities. - New artifacts intentionally change the active task tab. Existing history and repeated updates must not take focus again. - Transport flag defaults, server authorization, frozen skill snapshots, and completion policies remain unchanged. No schema migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cceeb0aa66 |
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9ed55f6931 |
fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts local command-line sessions and stores provider credentials through managed connections > - One OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run > - A second run then waited for the first run, and concurrent runs could overwrite a newer credential > - The write-back must compare fresh credentials while it holds the row lock > - This pull request removes the run lease, keeps the revocation guard, and bounds Codex timestamps against the host clock > - The benefit is safe concurrent use of one subscription connection with newest-credential selection ## Linked Issues or Issue Description **What happened?** A managed OpenAI Codex or xAI Grok subscription connection held a credential lease for the full agent run. A second run waited for the first run to finish. The write-back gate also rejected any row change before it compared credential freshness. **Expected behavior** Concurrent runs should start on one subscription connection. The server should keep the newest valid credential and reject a credential write after a person revokes the connection. **Steps to reproduce** 1. Start two runs that use one OpenAI or xAI subscription connection. 2. Let both provider tools refresh the credential. 3. Finish the runs in either order. 4. Confirm that the newest valid credential remains in the connection. **Paperclip version or commit** `ec25bf1e4a81d1729a6d7276e6a486587cff4a0b` **Deployment mode** Local dev (`pnpm dev`), built from source. **Agent adapter(s) involved** Codex. The server path also covers xAI Grok subscription connections. **Database mode** Embedded PGlite for local development, and external Postgres for deployments. **Access context** Both board and agent runs can use managed connections. **Additional context** The provider command-line tool refreshes credentials inside the sandbox. The server copies the result back after the run. Two long runs can still refresh one token hours apart, so the provider can reject the second refresh. The server cannot observe that provider call. ## What Changed - Remove the full-run credential lease for OpenAI Codex and xAI Grok subscription connections. - Lock and re-read the connection row before credential write-back. - Accept only a strictly newer credential, while keeping the connection revocation guard. - Reject Codex freshness timestamps more than five minutes ahead of the host clock. - Add tests for both completion orders, xAI cleanup, revocation, the Codex time bound, and agent hiring. - Update the connection and run-log documentation. ## Verification - `server/src/__tests__/ai-connections.test.ts` passes with 42 tests. - `packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts` covers the five-minute boundary and the one-millisecond overflow. - `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts` passes. - `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI and Anthropic. - The project type check reports no new error in changed files. - GitHub Actions must pass on this pull request. ## Risks The write-back now permits concurrent runs, so the provider may reject a later refresh when both long runs use one token. The server keeps the revocation guard and rejects future-dated Codex timestamps. No schema change occurs. ## Model Used OpenAI Codex, GPT-5. Context window and exact deployment build are not exposed in this run. The model used tool calls, code inspection, and test verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b64469e403 |
feat(workspaces): add an operator default for isolated execution workspaces (#13444)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The execution workspace subsystem decides if a task run uses the
shared project checkout or an isolated per-task git worktree
> - The mode comes from the project policy, then the task settings. A
project that stores no policy always falls back to the shared checkout
> - An operator who wants every project to use isolated workspaces must
therefore edit each project one at a time, and must repeat this for each
new project
> - There is no instance-level control, so a fleet operator cannot set
this default at all
> - This pull request adds a managed experimental flag that moves the
default for projects that store no policy of their own
> - The benefit is that an operator sets the workspace default one time,
and every current and future project follows it
## Linked Issues or Issue Description
No public issue exists. The description follows the feature request
template.
**Subsystem affected**
Execution workspaces. The files are
`server/src/services/execution-workspace-policy.ts` and the run dispatch
path in `server/src/services/heartbeat.ts`.
**Problem or motivation**
`resolveExecutionWorkspaceMode` reads only the project policy, the task
settings, and a legacy field. Its last statement returns
`shared_workspace`. A project that stores no policy always gets the
shared project checkout.
An operator has no way to change this default for many projects at the
same time. The operator must edit each project, and must edit each new
project again later. Tasks in one project therefore share one checkout,
and they run one at a time when the environment driver makes the
scheduler serialize them.
**Proposed solution**
Add the managed experimental flag `enableIsolatedWorkspacesByDefault`.
When the flag is on, a project that stores no policy of its own resolves
as if it selected isolated workspaces. A project that stores a policy
keeps that policy.
The new helper substitutes a project policy. It does not move the last
statement of `resolveExecutionWorkspaceMode`. Two behaviors make this
necessary:
- A task that has no project must keep its current behavior. An isolated
workspace needs a repository to cut a worktree from.
`isUnrunnableWorktreeCombo` blocks an isolated task that has no
`projectId` and no `projectWorkspaceId`. A moved fallback would resolve
isolated for project-less tasks, such as agent chat, and stop them
before dispatch.
- The mode and the strategy must agree.
`buildExecutionWorkspaceAdapterConfig` supplies the default
`git_worktree` strategy only when one layer asserts workspace control. A
moved fallback would leave isolated mode with a `project_primary`
strategy.
**Alternatives considered**
- Change the last statement of `resolveExecutionWorkspaceMode` to
`isolated_workspace`. This is one line, but it changes the default for
every deployment. It is also not gated, so it would apply where isolated
workspaces are off.
- Write the policy to each project row with a script. This does not
cover new projects, and it does not cover new instances.
- Add an instance defaults section to the managed-config document. This
needs a new document key, new validation, and new delivery code. A
boolean flag reuses the delivery machinery that exists today.
**Roadmap alignment**
`ROADMAP.md` does not list execution workspace defaults. This change
adds an operator control to an existing capability. It does not add a
new capability.
**Additional context**
The flag is `tier: "managed"`. A cloud operator can therefore deliver it
with the managed-config machinery that exists today. No new delivery
code is needed.
## What Changed
- Add `enableIsolatedWorkspacesByDefault` to the feature catalog with
`tier: "managed"`. Both defaults are off.
- Add the flag to the experimental settings schema, the type, and both
branches of `normalizeExperimentalSettings`.
- Add `applyDefaultIsolatedExecutionWorkspacePolicy` to
`execution-workspace-policy.ts`. It substitutes `{ enabled: true,
defaultMode: "isolated_workspace" }` only when the flag is on, the task
has a project, and the project stores no policy.
- Apply the helper in the run dispatch path in `heartbeat.ts`, after the
existing `gateProjectExecutionWorkspacePolicy` call. The `hasProject`
argument reads the resolved project row, not the raw `projectId` of the
task.
- Gate the new flag behind `enableIsolatedWorkspaces` at the call site.
The new flag does nothing on its own.
- Add a toggle card to the instance experimental settings page. The card
shows only when isolated workspaces are on.
- Add eight tests for the new helper.
## Verification
Commands:
```
pnpm --filter @paperclipai/shared typecheck
pnpm --filter @paperclipai/ui typecheck
cd server && ../node_modules/.bin/tsc --noEmit
```
The server typecheck script also builds a Rust binary. I ran `tsc`
directly because this machine has no `cargo`. The server package reports
no type errors.
Tests:
```
./node_modules/.bin/vitest run \
server/src/__tests__/execution-workspace-policy.test.ts \
server/src/__tests__/instance-settings-service.test.ts \
server/src/__tests__/instance-settings-cloud-defaults.test.ts \
server/src/__tests__/instance-settings-managed-overlay.test.ts \
server/src/__tests__/instance-settings-routes.test.ts \
server/src/__tests__/managed-config.test.ts \
server/src/__tests__/heartbeat-workspace-busy.test.ts \
server/src/__tests__/heartbeat-workspace-session.test.ts \
server/src/__tests__/heartbeat-workspace-ready-comment.test.ts \
server/src/__tests__/execution-workspaces-service.test.ts \
server/src/__tests__/issue-runtime-workspace-binding.test.ts \
server/src/__tests__/run-trust-preset.test.ts \
packages/shared/src/feature-catalog.test.ts \
packages/shared/src/settings-visibility.test.ts \
packages/shared/src/validators/instance.test.ts \
ui/src/pages/InstanceExperimentalSettings.test.tsx \
ui/src/components/Sidebar.test.tsx
```
All of these files pass. The new tests cover each of these cases:
- The helper substitutes an isolated policy for a project that stores
none.
- The helper changes nothing while the flag is off.
- The helper changes nothing for a task that has no project.
- The helper keeps a stored policy, including a policy with `enabled:
false`.
- The resolver returns `isolated_workspace` for an unpolicied project.
- An explicit task setting still wins over the operator default.
- The substituted policy produces the `git_worktree` strategy.
- A project-less task does not become an unrunnable worktree.
To confirm the behavior by hand:
1. Turn on Isolated Workspaces, then turn on Use Isolated Workspaces By
Default.
2. Open a project that has no execution workspace policy.
3. Start a task in that project.
4. The run gets its own worktree. Tasks in that project no longer wait
for each other.
## Risks
Low to medium. The details:
- The flag defaults to off, and it is inert unless
`enableIsolatedWorkspaces` is also on. An instance that does not turn on
both flags sees no change.
- A project that stores a policy keeps it. This includes a policy with
`enabled: false`, which the helper reads as a decision to stay on the
shared checkout.
- When an operator turns the flag on, the workspace configuration
fingerprint changes for projects that store no policy. Their next run
creates a new workspace. This is correct, because the mode did change,
but the first run after the change does more setup work.
- A task that is in flight when the flag changes resumes with a
different workspace path than the path its session remembers. An
operator should let current runs finish before turning the flag on.
- Isolated workspaces use more disk, because each task gets its own
worktree.
## Model Used
Claude Opus 5 (`claude-opus-5`) in Claude Code, with extended thinking
and tool use. The model read the repository, made the change, and ran
the typechecks and tests above.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
a2e7ffdc34 |
fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in remote sandboxes. > - Connection checks must use the selected execution target. > - Recovery must respect the controller that owns an active run. > - A host path or PID does not describe a remote sandbox. > - This pull request checks sandbox paths on the target and preserves live controller leases. ## Linked Issues or Issue Description **What happened?** Selecting an AI account for a sandbox agent could fail because the Claude ACP environment check tried to create the sandbox directory on the Paperclip host. The recovery sweep could also interrupt a sandbox run while its controller lease was still valid. It treated a PID absent from the local host as proof that the run had stopped. **Expected behavior** ACP checks directories on the selected execution target. Recovery leaves a run with a live controller lease alone. Its final database write rejects a stale snapshot after renewal, a claim, a controller change, or a runtime change. **Steps to reproduce** 1. Test a Claude ACP sandbox agent with a directory that cannot be created on the host. The check fails before this fix. 2. Give a running sandbox task a valid controller lease and a PID absent from the host. Run the stale-lock sweep without an in-memory handle. The sweep interrupts the run before this fix. 3. Renew or replace the controller between the sweep's read and write. The old snapshot must not end that controller's run. **Paperclip version or commit** Rebased onto `origin/master` at `0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases still fail against this base. Refs #13438, which supplies the managed hiring and task-connection behavior, and #13433, which preserves non-assignee subscription comment wakes. This PR preserves both upstream changes and addresses the two remaining sandbox failures. ## What Changed - Resolve and create Claude ACP test directories through the execution-target helpers. - Preserve active legacy controller leases during stale-lock recovery, including finalization after a task becomes terminal. - Recheck the controller, lease, runtime mode, and native ownership in the terminal database write. - Add two sandbox-directory cases and five database-backed controller-lease cases. - Document the target used for ACP directory checks. ## Verification - Red: all seven new cases fail against `0e9b24c82` without these two implementation changes. - Before the final upstream sync, 192 focused tests passed. Full `pnpm test:run` coverage completed using the repository's group/shard runner: all general server and workspace groups passed, and all 147 serialized server suites passed across the initial run and isolated continuations. Five cold-import timeout suites passed with `--experimental.fsModuleCache`; their assertions and deadlines were unchanged. - The hiring routes, default-selection service, and upstream hiring tests match `origin/master` exactly. The two remaining fixes are unchanged by the final rebase. - On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372 focused tests pass across 16 suites covering both upstream changes and these fixes. Two timeouts in the combined run (database setup and an existing ACP case) pass in isolated reruns with fresh test homes and temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on this head. - All 32 active checks pass on final head `bd28d5cef`, including the full test matrix and browser shards, in [CI run 34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849). Two Storybook checks are skipped by path filters. - Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or unresolved review threads. ## Risks - A failed remote directory check still blocks connection adoption. - A live controller retains finalization authority after its task becomes terminal. Cleanup waits for ownership to expire and must pass the final ownership check. - No schema or credential-storage changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code execution, and API tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f1905d34d |
fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects can now attach several source repositories. > - Task preparation still treated these sources as alternative workspaces. > - Sandbox sync preserved Git history only for the selected repository. > - A task needs every attached repository to complete work across the project. > - This pull request prepares all distinct project repositories and preserves their separate Git histories through sandbox restore. ## Linked Issues or Issue Description **What happened?** A user reported that a project with two repositories received only the first repository in Daytona. Repository-only project rows also reached the agent with null local paths. Managed checkouts with matching repository names could resolve to the same directory. **Expected behavior** Local and sandbox tasks receive every distinct repository attached to their project. Repository-only sources work without preconfigured local folders. Each repository keeps its own Git history and working files. **Steps to reproduce** 1. Create a project with two repository sources and no local folder paths. 2. Assign a task to the project and run it in Daytona. 3. Inspect the task workspace and the repository paths exposed to the agent. 4. Observe that the original implementation supplies only the selected checkout. Related change: #13010 added multiple repository selection. The open repository-catalog proposals #11234 and #11228 cover a different data model. This fix uses the existing project workspaces. ## What Changed - Materialize each additional distinct repository as an editable checkout inside the task root. Seed configured local sources with their current working files and retain task edits across runs. - Pass materialized repository paths to local agents and native sandbox task prompts. Apply existing run-scoped Git credentials to each remote clone. - Preserve each repository's Git history, dirty files, and restore baseline during sandbox staging and durable recovery. Apply each repository's ignore rules and the operator's workspace exclusions. - Keep same-name managed repositories in separate directories. Report additional clone failures before the task starts. - Add task-level, checkout, sandbox round-trip, environment-hint, and recovery-descriptor regression coverage. Document checkout and restore behavior. ## Verification - Red: the original implementation fails the sandbox test because the second repository has no Git directory. It also fails the same-name checkout test and both real-database task tests because repository hints have no local path. - Green: focused tests pass for one and two repository-only sources, local source edits, clone failures, per-repository credentials, separate Git histories, ignored files, and recovery from remote or durable seed state. - Live Daytona smoke passed with two disposable repositories through the production provider sync functions. Both repositories arrived with Git history. Commits from both restored locally. Ignored files stayed excluded. The disposable sandbox was deleted. - Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm build`. - Final focused coverage: 254 assertions across the six changed test areas passed across the serial run and an isolated rerun of the existing process-kill timing test. The live Daytona smoke also passed. - The local `pnpm test:run` overlapped source edits and retained stale transformed code. Its first phase reported 12,240 passed assertions, nine failed assertions, three hook failures, and one worker error; later phases did not run locally. This run is not claimed as green. Fresh focused tests verify the changes, and every general/workspace and serialized-server CI shard passes on the final commit. - Final CI is green on `93ab76763`: all test shards, all three browser shards, typecheck, build, runner verification, canary dry run, and security checks. The initial unrelated chat-delivery browser timing failure passed in the final CI run. Optional Storybook visual checks were skipped. - Greptile is 5/5 on the final commit with no unresolved review threads. Its checkout-race finding was reproduced with a failing test, fixed, and rechecked. ## Risks - Additional repositories need disk space and clone time. Access failure for an attached repository stops preparation. - Additional checkouts live under `.paperclip-repositories/` and keep independent histories. Changes stay in those task copies; they do not overwrite configured source folders. - Detached or reconfigured repository copies are retained under `.paperclip-runtime/detached-repositories/`. Sandbox recovery retains per-repository merge baselines. - No database migration, UI contract change, or new credential delegation is required. Referenced projects retain their separate read-only behavior. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and tool use. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5054c9ef9b |
fix(grok): stage the environment test from a host directory, and survive an absent workspace (#13416)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent runs through an adapter. Before you use an environment, the
adapter's environment test probes it and reports structured checks
> - The grok adapter's environment test stages the managed account into
a remote environment first. It gave the runtime its resolved `cwd` as
the host workspace directory
> - For a remote target that `cwd` is a path inside the sandbox. It does
not exist on the host that runs the test
> - A new workspace ignore scan reads that host directory. The scan
failed on the missing directory, and the error escaped the environment
test
> - The user got a server error instead of a result with checks
> - This pull request stages from a host temporary directory, sends the
remote path separately, and reports a staging failure as a check
> - The benefit is that the grok environment test always answers with
checks, and a workspace directory that does not exist no longer stops a
runtime preparation
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
The grok environment test failed with `Error: Workspace ignore scan
failed: git-ignore-scan-failed`. The route returned a server error, so
the user saw no checks at all.
**Expected behavior**
The environment test always returns `{status, checks}`. Every other
problem it finds (an invalid working directory, a command it cannot
resolve, a probe that times out) becomes a check with a level. A
credential staging problem must do the same.
**Steps to reproduce**
1. Configure a grok agent with a managed AI connection.
2. Point the agent at a remote (sandbox) environment.
3. Run the environment test for that environment.
**Paperclip version or commit**
Present on master. Both halves landed on 2026-09-12: the call site in
#13247, and the scan that it trips in #13353.
## What Changed
- `packages/adapters/grok-local/src/server/test.ts` stages the managed
account through a fresh host temporary directory and passes the remote
path as `workspaceRemoteDir`. This is the split `codex-local` and
`opencode-local` already use.
- A failure while staging becomes a `grok_environment_unprepared` check
with level `error`, instead of an exception that escapes the function.
The probe does not run after it, because there is no prepared
environment to probe.
- The temporary directory is removed in the existing `finally` block.
- `packages/adapter-utils/src/sandbox-managed-runtime.ts` treats a
`workspaceLocalDir` that does not exist as "nothing to sync". A
directory that is not there has no files for ignore rules to govern,
nothing to stage, and nothing to restore.
- Tests: the grok environment test now covers the managed-connection
remote branch, which had no coverage. `sandbox-file-sync.test.ts` covers
a preparation whose workspace directory does not exist.
## Verification
```
npx vitest run packages/adapters/grok-local/src/server/test.test.ts # 9 passed
npx vitest run packages/adapter-utils/src/sandbox-file-sync.test.ts \
packages/adapter-utils/src/sandbox-managed-runtime.test.ts # 90 passed
cd packages/adapter-utils && npx tsc --noEmit # clean
cd packages/adapters/grok-local && npx tsc --noEmit # clean
```
The two new grok tests fail against the old call site: the first asserts
the runtime never receives the remote path as its host workspace
directory, and the second asserts a staging failure becomes a check
instead of an exception.
## Risks
Low risk, and limited to environment preparation.
- The adapter change only affects the managed-connection remote branch
of one adapter's environment test.
- The `adapter-utils` change makes a preparation that used to throw now
continue with no workspace sync. A real workspace is unaffected, because
the directory exists in that case and the scan runs exactly as before.
- No migration. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
46bcc0863c |
test(runner): stop two CI-contention flakes that starve the canary lane (#13418)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Each master commit is published as an npm canary, and the nightly release gate installs the newest canary and runs the onboarding smoke against it > - A canary only publishes after the Cloud readiness workflow passes for that exact commit > - Cloud readiness fails often, and almost every failure stops that commit from ever becoming a canary > - The nightly gate then keeps testing an old canary, so a fix that landed after it cannot reach the gate, and the gate stays red for reasons no longer in the code > - Most of these failures are short timeouts that only expire because the runner is loaded > - This pull request gives two such timeouts the budget the surrounding tests already use > - The benefit is that more master commits publish a canary, so the nightly gate tests current code ## Linked Issues or Issue Description No existing issue. The problem, in the bug report format: **What happened** Cloud readiness failed on 25 of 85 runs over three days (29%). 24 of those failures stopped the commit from publishing a canary. Two timeout clusters account for 10 of the 25: - `packages/paperclip-runner/src/control-plane/durable-prp-control-plane.test.ts` — `Error: Test timed out in 5000ms.` in "promptly fails the real transport request and notification paths on authenticated bad semantic input". Seen in runs 34652752663, 34654215468, 34656885170, 34657111560, 34661910120. - `packages/paperclip-runner/src/live/live-session.test.ts` — `Turn settled before the intentional SIGKILL: Error: Capability Codex turn timed out after 2000ms`. Seen in runs 34632264036, 34658026290, 34658106386, 34727991195, 34760283177. **Expected behavior** A commit that is good publishes its canary. A test fails because the behavior is wrong, not because the runner was busy. **Steps to reproduce** 1. Run the Cloud readiness workflow on master repeatedly. 2. Watch the Runner protocol and chaos jobs. 3. The two tests above fail intermittently, and the commit then publishes no canary. **Paperclip version or commit** Present on master. Measured over runs from 2026-09-11 to 2026-09-14, which is the workflow's whole history. ## What Changed - `durable-prp-control-plane.test.ts`: the "promptly fails the real transport request..." `it.each` now takes the same 15s budget as the test directly above it. It ran on the 5s default, although it builds a real control plane over a socket. The neighbouring test already asks for 15s, so this reads as an omission. - `live-session.test.ts`: the real-runner SIGKILL test gives the turn a 60s timeout instead of 2s. The test kills the runner process and then asserts the turn was still running at that moment, so the timeout is only a backstop. At 2s a loaded machine settles the turn first and the assertion fails. Neither change weakens an assertion. No test in this PR asserts that a timeout happens, and the tests that do assert one (`turnTimeoutMs: 500` and `turnTimeoutMs: 20`) are untouched. ## Verification ``` cd packages/paperclip-runner npx vitest run src/control-plane/durable-prp-control-plane.test.ts # 87 passed npx vitest run src/control-plane/durable-prp-control-plane.test.ts \ -t "promptly fails the real transport" # 2 passed (both it.each cases) ``` The live-session change is verified by the readiness lane itself, because that suite needs a real runner and a provider. ## Risks Low risk, and limited to test budgets. - A genuine hang in either test now takes longer to report: 15s instead of 5s, and 60s instead of 2s. - A real regression still fails, because the assertions are unchanged. - Timeouts are not a complete fix for this lane. Port races (`EADDRINUSE`), lease races, and lost runner VMs cause about another 40% of its failures. A retry on the `verify` matrix would cover all of those classes, and none of them reproduce on a second attempt. That is a larger policy change, so it is not in this PR. ## Model Used - Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run through Claude Code with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes — no documented behavior changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green — pending first run - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending first review - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
e1d2e279a3 |
fix(db): replay a query whose socket write failed on a recycled pooled connection (#13417)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - A hosted instance keeps its data in Postgres, and it reaches that
database through a connection pooler
> - A pooler recycles the server side of a connection that sat idle. The
client still holds the socket and believes it is open
> - The next query fails while the driver writes it to that dead socket,
with `write CONNECTION_CLOSED host:5432`
> - The request that happened to draw the recycled connection fails,
although nothing is wrong with the query or the database
> - Only one code path guards against this today, so the same error
keeps reaching users and error reporting from every other path
> - This pull request replays a query when the write itself failed,
because those bytes never reached the server
> - The benefit is that a recycled connection costs one retry instead of
one failed request
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
Requests fail with `write CONNECTION_CLOSED <host>:5432` (driver code
`CONNECTION_CLOSED`). It is most visible on the first queries after an
idle period, when the pool holds connections the pooler already
recycled.
**Expected behavior**
A connection that the pooler recycled while it was idle is not a
user-visible failure. The client notices the dead socket and uses a live
connection.
**Steps to reproduce**
1. Run the server against a pooled Postgres endpoint (for example a Neon
`-pooler` host).
2. Leave the instance idle until the pooler recycles the server side of
the pooled connections.
3. Issue any request that queries the database.
**Paperclip version or commit**
Present on master.
## What Changed
- `packages/db/src/transient-write-retry.ts` (new) wraps the root
`postgres.js` client. When a query fails while the driver writes it to
the socket, the wrapper runs it again on a fresh connection. Three
attempts, 50 ms then 100 ms backoff.
- `packages/db/src/client.ts` gives Drizzle the wrapped client. The
teardown registry keeps the real client, because shutdown must end the
actual pool.
- The wrapper is deliberately narrow:
- It matches only the write phase (`code === "CONNECTION_CLOSED"` and a
message that starts with `write CONNECTION_CLOSED`). The write failed,
so the server never saw the query, and a replay cannot run anything
twice. That makes it safe for reads and writes alike.
- An error after the write propagates untouched, because the server may
have acted on the query.
- `CONNECTION_ENDED` and `CONNECTION_DESTROYED` propagate untouched,
because they mean a deliberate shutdown.
- Queries inside `db.transaction()` run on the scoped client that
`sql.begin()` returns, which the wrapper does not touch. A transaction
that loses its connection must abort, not replay.
- A query runs once however many handlers attach to it, and it still
executes lazily, like `postgres.js` itself.
## Verification
```
npx vitest run packages/db/src/transient-write-retry.test.ts # 7 passed
npx vitest run packages/db/src/client-teardown-registry.test.ts \
packages/db/src/client-options.test.ts # existing db suites pass
npx vitest run server/src/__tests__/cloud-tenant-transient-db-retry.test.ts # 6 passed
cd packages/db && npx tsc --noEmit # clean
```
New tests cover the replay, the `.values()` form Drizzle uses, the
attempt budget, an error that must not replay, and one execution per
pending query. A Drizzle round trip runs through a fake wire-protocol
server, which shows the wrapper is transparent to ordinary queries.
## Risks
Low risk.
- The retry only fires for a failure during the socket write, where the
server never received the query. A replay therefore cannot duplicate an
effect.
- A permanently unreachable database costs two extra attempts and 150 ms
before the same error surfaces.
- Transactions keep exactly their current behavior.
- Existing `retryOnTransientDbConnectionError` in the auth middleware
stays. It wraps a broader set of codes for one path, and it is
unaffected.
- No migration. No configuration change. No API change.
## Model Used
- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior or configuration changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — pending first run
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
pending first review
- [x] I will address all Greptile and reviewer comments before
requesting merge
|
||
|
|
02c7175e72 |
feat(agents): hired agents inherit provider credential references (#13268)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters need provider credentials to start work > - A hired agent can lack the credential reference that its hiring agent already uses > - The child agent then cannot authenticate, even when the company has a valid credential > - This pull request copies matching credential references from the hiring agent to the hired agent > - The benefit is that hired agents can start with the provider access that the hiring agent already uses ## Linked Issues or Issue Description This change relates to [PR #9920](https://github.com/paperclipai/paperclip/pull/9920), which covers credential inheritance for other agent creation paths. This pull request covers hiring-specific inheritance and fixed Claude OAuth binding checks. ### What existing behavior does this improve? The agent hire route builds the child adapter configuration from the hire request only. ### Current behavior A hired agent does not receive matching provider credential references from its hiring agent. The child agent cannot run when the request omits the credential. ### Proposed behavior The hire route inherits matching credential references from the hiring agent. The request keeps priority. A Claude hire that supplies any Claude credential inherits none. ### Reason and benefit The child agent can use the provider access that the hiring agent already uses. The change copies references only and never copies raw token values. ### Breaking changes None. The change affects only hires that need an inherited reference. ## What Changed - Copy matching credential references from the hiring agent into the hired agent adapter configuration. - Preserve pinned versions, `required`, and `allowMissingOverride` fields on each copied reference. - Keep hire-request credentials ahead of inherited credentials. - Reject inherited fixed Claude OAuth bindings unless the parent agent passes company, adapter, and exact-binding checks inside the same transaction. - Add route and service tests for inheritance, precedence, and binding validation. ## Verification - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-auth-inheritance-routes.test.ts src/__tests__/agents-claude-oauth-binding.test.ts` — 60 passed. - `pnpm exec vitest run --project @paperclipai/server src/__tests__/agent-hire-idempotency-routes.test.ts src/__tests__/agents-service-secret-bindings.test.ts src/__tests__/secrets-service-user-secret-owner-scoped.test.ts` — 28 passed. - `tsc --noEmit` in `server/` — the error count matches the merge base, with no error in either changed source file. - `git diff --check` — clean. ## Risks Low risk. The route copies references, not raw tokens. The request keeps precedence. Company, adapter, and exact-binding checks protect the inherited Claude OAuth path. ## Model Used Codex, OpenAI GPT-5, with code execution and review support. The implementation commit predates this pull request handoff. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
78ce96a48b |
fix(runner): restore native Claude context and read permissions (#13422)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner supplies each agent with instructions, assigned skills, and tools. > - Native Claude lost its skill snapshot before provider launch. It also changed exact model IDs to aliases. > - The read permission mode denied ordinary Paperclip reads because this runner has no interactive approval handler. > - This pull request restores the missing context and permits only assigned tools that Paperclip defines as reads. > - Other requests that need approval stop with a clear action for the operator. They do not retry automatically. > - Agents can complete read tasks. Users can see why a restricted task stopped. ## Linked Issues or Issue Description **What happened?** Native Claude could not find assigned skills. The ACP layer could replace an exact model ID with an alias and then fail identity verification. The default `approve-reads` mode denied Paperclip read tools. A denied write showed a generic transport failure. The tool descriptions also called live operations “mock” operations. The completion prompt said to call a completion tool once, although the protocol can reject a claim and require a corrected call. **Expected behavior** Load assigned skills before launch. Keep the selected model ID. Allow assigned Paperclip reads. Stop an operation that needs approval with clear instructions when no approval handler exists. **Steps to reproduce** 1. Configure a native Paperclip Runner agent with ACPX Claude and an exact model ID. 2. Assign a skill and ask the agent to use it. 3. Select the read permission mode and ask the agent to read task context and list documents. 4. Ask the agent to write a document. Check the task state and recovery message. **Paperclip version or commit** Reproduced from `d351e08deee1b49d3467a950d1a3f01131943441`. **Deployment mode** Local native runner with real Claude, Rust runnerd, the Paperclip server, and embedded PostgreSQL. Related: #13196 fixed remote skill staging in the legacy `claude_local` adapter. The native runner uses a separate path, which this pull request fixes. ## What Changed - Carry the runtime context through Rust and the ACPX sidecar. Load assigned Claude skills after the provider lifetime lease is held. Refresh the files on each open. - Write the exact requested model ID into the isolated Claude settings. - Grant exact MCP permissions for the intersection of assigned tools and Paperclip's read catalog. Provider hints cannot grant access. Existing task-control permissions stay in place. - Stop requests that need an unavailable approval handler. Preserve the typed error through the server. Show “Approval required” on the task and require operator action without automatic retry. - Label the setting “Allow Paperclip reads.” Remove “mock” from live tool descriptions and regenerate the contracts. - Change one completion-prompt sentence to require one accepted result. Add regression tests and update the runner documentation. ## Verification - Red/green regression tests cover read admission, unavailable approval handling, server recovery, and the task error message. - A real Claude read trial failed before the fix and completed after it: [red trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=95e6a326e24437f1), [green trace](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=85a1f347db5604d1). - Full browser tests used real Claude and the production tool authority. Reads completed. A write stopped with “Approval required.” No document was created and no automatic retry was scheduled. All nine checks passed: [events and screenshots](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=1c4ff10be0d779cc). - A separate full browser test assigned a skill, invoked it, and completed with a marker absent from the task prompt: [skill evidence](https://www.braintrust.dev/app/Paperclip/object?object_type=project_logs&object_id=fda74078-b00d-495a-90d3-ea1be019d71b&id=ba5fbcb29a60be52). - Post-rebase checks passed: 128 targeted runner tests, 46 Rust tests, 27 sidecar/protocol contract tests, `pnpm -r typecheck`, and `pnpm build`. The server/UI red-green checks and `pnpm check:token-gates` also passed. The full suite passed in CI, including all general and serialized test shards, all three browser shards, and runner `check:all`. The duplicate unsharded local full-suite run was stopped after CI passed. Braintrust links require project access. The traces contain provider/runner events and app outcomes. They do not contain raw model HTTP requests. ## Risks - `approve-reads` now allows assigned Paperclip reads. Other operations that need approval stop the turn. Users must review the operation and change permissions before retrying. - The Claude settings depend on the pinned ACP and SDK behavior. Automated tests and real Claude trials cover this boundary. - The native Codex path is unchanged. The ACPX Codex fallback shares the clearer approval failure handling. - Local execution was tested end to end. Remote execution was not run. This change has no database migration. ## Model Used OpenAI GPT-6 through Codex. The agent used reasoning, code execution, and browser tools. The exact deployment ID and context-window size were not exposed in the session. Live acceptance tests used `claude-sonnet-5`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
728f7185f6 |
feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Self-hosted boards need a way to show occasional product announcements. > - An app release should not be required to publish or withdraw a card. > - Native card controls keep publishing consistent; the hero can use a static image or isolated HTML/CSS animation. > - This pull request renders a validated JSON feed with native components. > - It stores dismissals per account on each instance, so a closed card stays closed across companies and browsers. > - Named staging feeds let authors test content before production publication. ## Linked Issues or Issue Description **Subsystem affected** Board application shell, announcement delivery, and user preferences. **Problem or motivation** Operators need a small, optional announcement card. Users need reliable dismissal state. Authors need to test remote content without changing the production feed. **Proposed solution** Add one non-modal AnnouncementWell. Fetch validated JSON and content-addressed media through the instance server. Keep card controls native, with optional sandboxed HTML/CSS animation in the hero. Use stable announcement IDs for dismissal, an explicit empty manifest and quiet 404 handling. Provide a staged publishing helper and isolated test-drive guide. **Alternatives considered** Hosting the entire card as a page would move navigation and dismissal into remote content. This change limits HTML to a scriptless, isolated visual hero and keeps controls native. Browser-only storage would lose dismissals across browsers, so the instance stores account preferences. **Roadmap alignment** ROADMAP.md has no overlapping announcement feature. A GitHub title search found no related announcement pull requests. This work implements a maintainer-requested feature. ## What Changed - Add shared feed types, strict validation of every object, supported routes, expiration and version checks. - Add a board-only current-feed API, constrained media proxy, and idempotent dismissal API. Store the first dismissal and its company audit entry in one transaction. - Cache upstream data for one hour. Use conditional requests, request deduplication, response limits, public destination checks, and a three-second deadline. Treat a remote 404 as an empty feed with a fifteen-minute retry cooldown. - Keep announcement visibility stable when focus moves to browser chrome or another app pane; only tab visibility starts a return check. - Add a responsive native announcement card. Respect onboarding, dialogs and toast placement. Sync pending dismissals across tabs and retry after reconnect or return. - Add idempotent migrations for dismissals and validated publication IDs, design-guide examples, static and animated Storybook examples, and focused tests. The publication registry supports offline retries without accepting caller-invented IDs. - Add HTML/CSS animated heroes with static posters, automatic playback, reduced-motion handling, strict DOMPurify validation, an empty iframe sandbox and CSP that blocks scripts/network resources. - Add validated staging publication, content-addressed assets, an empty production manifest, preview fixtures, and authoring/operator documentation. ## Verification - The preceding implementation passed 98 targeted shared/server/publisher/route/OpenAPI/UI tests and 127 tests including the master rebase. The playback-control removal passes all 21 announcement UI tests, covering the rendered sandbox, fallback, reduced motion, dismissal and slow/stale state lookups. The preceding shared/server tests cover HTML validation and response sandbox headers. - The playback-control removal passes UI typecheck, production UI build, Storybook build and token gates locally. Browser verification confirms the animated card has only its dismiss button and two links, with no page errors. The full canonical CI matrix passed on current head `00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two optional Storybook deployment checks skipped. This run needed no retries. Greptile reviewed this same head at 5/5 with no outstanding findings. - The local canonical general-server run passed 12,063 tests before reporting embedded-PostgreSQL startup failures in an unrelated fixture. All 31 tests in that fixture passed across isolated retries. The UI group passed 6,219 tests and other workspace groups passed 3,201; two CLI database-startup failures also passed individually. Serialized server suites were verified by the full CI matrix rather than repeating them locally. No source changes were needed for these environment failures. - The real S3/CloudFront staging manifest and both media asset headers were verified. Production remains empty/unpublished. The guide distinguishes the preview host's disabled edge cache from production cache requirements. - In the isolated test-drive, the animation visibly moves without playback controls. A 390×844 browser viewport keeps the card above navigation. Reduced motion makes no animation request. Both themes render correctly and browser page errors are empty. Browser fault injection verified that scripts cannot execute and CSS cannot make network requests; a missing animation leaves its poster and controls. - Refresh leaves the animated card visible. Closing it persists after reload and the API returns null. Earlier live checks verified dismissal across browsers, company-relative CTA navigation, modal deferral/restoration, and new-ID eligibility after restarting the same database. - The deployed empty feed and a real remote 404 return HTTP 200 with null from the board API, with a usable dashboard and no announcement popup or browser warnings. - Authoring documentation covers staging, animated HTML constraints, test-drive, withdrawal, ID reuse and cache-refresh steps. ## Risks - Animation supports self-contained visual HTML/CSS and inline SVG, without JavaScript or external resources. A static image is required. Older builds that do not recognize the optional animation field quietly hide that unsupported feed. - The default feed makes an outbound request from an instance when a board is used. Operators can disable it. Requests contain no account IDs, company data, cookies or interaction events. - Feed publication and withdrawal can take about 65 minutes to reach returning users because of CDN and instance caches. Expiration also removes visible cards locally. - Dismissals follow an account within one instance. No-login instances share the existing local-board identity. Separate installations do not share state. - Both tables are additive. A unique key prevents duplicate dismissals; the transaction prevents duplicate first-dismissal audit entries. The publication registry retains only validated IDs. AGENTS.md and the implementation spec document the required exception to company scope for these instance-level records. - Publication was limited to separate public staging prefixes on the existing preview host. Production remains empty/unpublished. No AWS policies or infrastructure were changed. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context-window size are not exposed in this session. Capabilities used: reasoning, code editing, shell execution, tests, browser interaction, and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1d57863d2 |
fix: make connection checks and task handoffs reliable (#13404)
Preserve connection-probe outcomes through cleanup, reduce unrelated startup work, and report selected Claude authentication accurately. Make artifact download actions match their labels. Route delegated feedback through its active child, retain accepted messages across completion, and avoid redundant worker runs for proven closing notes. Preserve explicit follow-ups, human input, company boundaries, source provenance, and mixed issue references. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2d47a8058d |
fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly approves. ## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connector screens need recognizable app artwork. > - Several bundled marks have inconsistent artwork or dark-theme behavior. > - The shared logo component should retain the existing frame. > - This change replaces selected artwork and fixes local theme fallback. > - Users see consistent connector icons across shared component callers. ## Linked Issues or Issue Description **Current behavior** Some connector marks use outdated artwork or unsuitable theme variants. A remote dark logo can override a canonical local mark that works in both themes. **Proposed behavior** Use the selected bundled artwork with the existing gray rounded frame. Use local light artwork in both themes unless a distinct local dark variant exists. **Subsystem affected** Connector artwork, the shared AppLogo resolver, and its validation. **Breaking changes** No connector capability, permission, credential, or catalog activation changes. No duplicate PR was found in the earlier search. ## What Changed - Update 46 artwork files and only the app definitions whose logo paths need to change. - Keep a compact public manifest of identities, paths, visibility, and aliases. - Prefer local artwork in both themes and preserve the existing frame and padding. - Add locally runnable artwork safety checks and a light/dark Storybook gallery; leave PR workflows unchanged. - Document artwork conventions. Source research is kept outside the public manifest. ## Verification - Artwork check: 69 identities pass. - Node artwork validation tests: 13 pass (rechecked September 14). - Focused resolver, component, and catalog tests: 40 pass (rechecked September 14). - Maintainer decision: dedicated icon-validation CI is not required; the workflow remains unchanged. Greptile acknowledged 5/5 with no remaining code concerns on September 14. - UI typecheck, token gates, and Storybook build pass. - Earlier full build and repository typecheck passed. The broad local test run was stopped after workspace-runtime dependency fixture failures outside this change. Current GitHub CI remains the full-suite gate. - Visual approval remains outstanding. Review the canonical icon registry in both themes at 24–48px. ## Risks The SVG checker rejects common active features; it is not a general sanitizer for arbitrary uploads. Optical balance still requires human review. Remote fallback for unknown brands retains existing behavior. This change does not add an asset importer or new connector capabilities. ## Model Used OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and browser tooling. Exact hosted model IDs and context window sizes were not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Artwork comparison Before and after for every affected identity. Images follow your GitHub light/dark theme and are pinned to the base and PR commits. This compares artwork; the existing gray rounded product frame and padding are unchanged. | Connector | Before | After | |---|:---:|:---:| | AgentMail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | | Airtable | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | | Asana | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | | ClickHouse | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | | Cloudflare | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | | Cloudinary | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | | Discord | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | | GitHub | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | | Gmail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | | Google Calendar | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | | Google Chat | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | | Google Docs | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | | Google Drive | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | | Google People | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | | Google Sheets | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | | Google Slides | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | | Google Workspace Search | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | | Hugging Face | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | | Jam.dev (library only) | New library entry | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg" alt="Jam.dev" width="48" height="48"></picture> | | Jira | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | | Linear | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | | Manufact | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | | Microsoft Teams | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | | Miro | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | | Mixpanel | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | | Netlify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | | Notion | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | | PagerDuty | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | | PostHog | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | | Postman | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | | Shopify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | | Slack | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png" alt="Slack" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg" alt="Slack" width="48" height="48"></picture> | | Stripe | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | | Supabase | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | | Telegram | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | | Todoist | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | | Wix | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | | Zapier | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | |
||
|
|
f2c5e54dca |
fix(runner): preserve handoff work and publish requested files (#13355)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - A task keeps its instructions, progress, and files when its assigned agent changes. > - The replacement runner lost the interrupted run's context and could overwrite an existing draft. > - A saved message also stayed attached to the former agent and could reopen the task after the replacement finished. > - File tasks could report Done with only a local path that the user could not open. > - This pull request transfers handoff context and saved messages, and makes requested files accessible through the existing attachment contract. > - Users can change agents and collect completed work without repeating instructions or confirming bookkeeping. ## Linked Issues or Issue Description Refs #13338. Builds on merged #13354 for queue admission and #13353 for remote workspace retry. #10123 concerns restricted recovery-model escalation; this change instead covers ordinary native handoff and file completion. **What happened?** Codex wrote a draft before a user assigned the task to Claude. The replacement lacked continuation context and replaced the draft. A queued user message could later restart the former agent and reopen the completed task. Separately, a runner could finish a requested file but return only a machine-local path. Remote native runs had no bound file publication tool. **Expected behavior** The replacement reads and preserves existing work, receives saved messages once, and keeps each message's author. The former agent stays stopped. A requested file has a working attachment or accessible work product before Done. Text-only tasks do not require attachments. **Steps to reproduce** 1. Ask Codex to save three newsletter names and then wait. 2. Queue an instruction to keep those names and expand the draft. 3. Use Interrupt and assign to select Claude. 4. Verify the original names survive, the result has a working download, and only the source and replacement runs exist. 5. Ask either provider for a Markdown checklist and open the file from its completed response. ## What Changed - Carry the exact same-task interrupted run's summary, semantic receipts, and history into handoff context. Tell the replacement to inspect existing files before editing. - Adopt saved ordinary task comments into the successor's receipt under the task lock. Preserve authors and separate mention, chat, and interaction contracts. - Prevent a former-assignee comment wake from reopening a completed task or starting a stale execution. - Reject workspace-only, fabricated, and cross-task file completion references with actionable runner feedback. New file output also needs a matching current-run publication receipt and asset filename/size/hash, or an accessible work product registered by the current run. Prior output can remain context alongside a current file, or be verified and re-registered internally. Authorized chat attachment reuse retains its verified current-run clone receipt; older receipt shapes require an intact matching source. - Bind remote file reads to the active environment runner and reuse the existing attachment and work-product publication path. - Enforce workspace confinement, regular single-link files, stable identity, a 10 MiB limit, and exact size and SHA-256 checks. Rotate the native session fingerprint for the updated tool contract. - Contain rejected remote signals and protocol-failure cleanup, including logging failures. Preserve the original cleanup rejection for its owner; a rejected operation never supplies stop acknowledgement or cleanup proof. - Allow exactly one maximum-size base64 file through the native SSH command adapter, preserving a finite output cap. - Document handoff and accessible file completion rules. ## Verification - Each observed bug has a failing regression before its fix. Final post-rebase integration passed 732 tests across 13 files before the final receipt and signal guards; final affected results are below. - Publication provenance and compatibility: 8 provenance regressions and 2 compatibility regressions failed before their fixes; the final four affected suites pass 53 tests, including mixed old/new references and real authorized chat reuse. Controls cover old attachments and work products, filename/size/hash/origin mismatch, missing/wrong receipts, current-run publication, same-run durable proof, internally re-registering preserved bytes, and no-new-file follow-ups. - Remote signal rejection: the real Node subprocess previously exited 1 when the production launcher signalled a deleted sandbox. It now stays alive for both a failed signal and failed logging; all 349 executor tests pass. The failed signal still provides no termination proof. - Remote file reader and SSH command boundary: 33 tests passed, including real Linux descriptor reads and the actual SSH adapter subprocess output cap (network executable replaced by a deterministic fixture). Exact 10 MiB bytes pass, one byte beyond the encoded cap fails. - Live local Claude and Codex Stop journeys preserve the saved file, deliver queued instructions once, and reach Done with two total runs. The handoff journey preserves the original names and download with exactly two runs. Both local providers deliver exact checklist files without a completion confirmation. - Live combined Daytona verification passed: the original failed task's Retry reused its sandbox; a selected Git subfolder produced an exact downloadable file; warm and deliberately resumed Claude runs took about 33 seconds. Codex produced a 240-byte download in 33.1 seconds after 134 seconds of contention/backoff. Both cloud downloads retained exact bytes after the two owned sandboxes were deleted. - The sandbox-deletion retest identified a separate ignored promise in protocol-failure cleanup. Two real Node subprocess regressions failed under fatal unhandled-rejection policy before the fix; all 36 protocol, lifecycle, and integrity tests now pass. The original close promise still rejects to its owning runtime. The final live retest passed: a normal Claude Daytona task completed in 132.352 seconds, then its sandbox was deleted. Thirteen samples over 361 seconds confirmed the same controller stayed healthy, the task stayed Done with unchanged run IDs, and its attachment retained exact bytes. The post-deletion browser download passed with zero page errors; all five owned sandboxes are confirmed absent. - Full repository typecheck and build passed on final commit `de64d16f1`. Final-head Greptile is 5/5 with no unresolved threads. [Final-head CI](https://github.com/paperclipai/paperclip/actions/runs/34736623758) passed: 32 successful checks and two conditional skips. The earlier mixed-source full local test invocation was deliberately stopped before rebase, so no pristine green full local aggregate is claimed. Its known failures passed in later affected suites. ## Risks - Handoff may adopt only ordinary comments from its validated former owner. Other delivery contracts must remain independent. - File verification fails closed if a remote file changes during reading. The runner must retry publication or explain a blocker. - The updated session fingerprint starts a fresh provider process where needed to install the new tool contract. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI `gpt-6-astra` through Codex, with reasoning, code execution, browser testing, and tool use. The context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
827ba8a434 |
fix: cancel stalled sandbox startup without waiting for setup (#13352)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox runs prepare credentials and files before an agent starts. > - Stop must work during that preparation. > - ACPX registered cancellation, but did not handle it while a setup command was waiting. > - Daytona cleanup waited for that same command before stopping the sandbox. > - This change stops the run's sandbox first and requires proof before abandoning setup. ## Linked Issues or Issue Description **What happened?** Stop left a sandbox run active when remote credential setup stalled. The run kept its connection lease until the sandbox was stopped separately. **Expected behavior** Stop terminates the selected run's sandbox, prevents later setup from launching the agent, and lets run cleanup finish. It must not report success without proof from the provider. **Steps to reproduce** 1. Start an ACPX agent in Daytona. 2. Hold a command during remote credential or file setup. 3. Select Stop before the agent starts. 4. Before this fix, cleanup waits for the held command and never reaches sandbox stop. Related: #13351 exposed this during connection acceptance testing. #12150 addresses scheduler load and session initialization limits, a separate startup problem. ## What Changed - Handle cancellation during ACPX sandbox preparation with a host-owned stop callback. - Pass an explicit active-work cancellation flag through environment cleanup. - Stop Daytona before draining setup commands. Keep normal graceful cleanup. - Require an exact run and lease termination receipt. Keep ownership of outstanding requests when stop cannot be verified. - Reject late setup work and defer sandbox resume until old requests settle. - Add regression tests and document the cancellation boundary. ## Verification - Red: both the stalled ACPX setup test and the Daytona cancellation test failed before the fix because Stop never reached the provider. - Green: adapter and Daytona suites passed, along with cancellation boundary and database-backed receipt/isolation tests. - Two real Daytona probes ran a five-minute setup command. Cancellation returned matching stopped receipts in 6.01 and 6.533 seconds. No later setup command ran. Both test sandboxes were deleted. - Repository typecheck and build passed locally. The full required CI suite passed, including all server/workspace tests, browser shards, Paperclip Runner verification, and canary packaging. The duplicate local full-suite run was interrupted after CI passed; it is not claimed as a completed local pass. - No UI changes. The live probe uses the actual adapter cancellation boundary and Daytona plugin; it is not a browser acceptance test. ## Risks - Provider stop failures remain unacknowledged. The adapter keeps ownership while its original requests remain active. - A retry can receive a settling-work error until old provider requests finish. - Cancelled sandboxes are stopped and retained under their existing provider expiry policy. - Local execution and cancellation after the agent turn starts keep their current behavior. - Other providers must return an exact termination receipt to permit early setup cancellation. No database migration. ## Model Used OpenAI GPT-6 through Codex, with code execution and tool use. The exact deployment identifier and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c9e3bb7ca4 |
fix: preserve queued work after native Stop and honor steering support (#13354)
## Thinking Path > - Paperclip lets people manage AI agents and their tasks. > - The runner owns execution, while the task keeps user instructions and status. > - Stop must stop the current response without losing instructions that the user already sent. > - The queue stored its original reason inside saved context. Recovery checked the outer deferred reason and left the message waiting. > - Claude also exposed Steer through a shared method even though its driver did not support it. A rejected request could remove its own error row. > - This pull request keeps queued work until execution has stopped, uses the driver's real capability, and preserves completion event order. > - Users can continue work without repairing task state or repeating messages. ## Linked Issues or Issue Description Refs #13338. Related recovery work: #13353. **What happened?** A message sent during a native run stayed queued after Stop. Claude exposed an unsupported Steer action. A steering failure could hide the queued row and its error. A terminal event could also precede the final provider result, and subtree Stop omitted the board actor. **Expected behavior** Stop ends the current execution. Once Paperclip proves that execution has stopped, it delivers the saved instruction once through normal admission. Pause and recovery holds still prevent dispatch. Unsupported controls stay disabled, and a rejected action leaves an actionable error visible. Final results precede terminal events. **Steps to reproduce** 1. Start a Claude or Codex task that writes a file and then waits. 2. Send a follow-up instruction while it runs. 3. Press Stop. Check that the queued instruction runs once and preserves the file. 4. Check Claude's Steer control and simulate a server rejection on the only queued message. ## What Changed - Recover saved native comments after acknowledged Stop using their original wake reason. - Require durable remote termination receipts or verified local process termination before dispatch. - Preserve actor identity, queued-message deduplication, Pause, and recovery gates. - Derive steering support from the driver descriptor and reject unsupported calls. - Keep the queue mounted until a steering request succeeds so its error remains visible. - Emit provider results before terminal events and pass the board actor into subtree Stop. - Document Stop and steering behavior. ## Verification - Final-head continuation suite: 104 passed, after failing regressions for saved wake reasons, cleanup proof, and deduplication of every queued message. Steering UI: 121 passed. Driver capability: 26 passed. Runner backend/transport: 205 passed; Rust library: 285 passed. - Real Claude and Codex browser journeys both preserved the saved file, delivered the queued instruction once after Stop, and reached Done with exactly two total runs. The process Stop browser fixture also passed. - Local full repository typecheck and build passed on `afaa35139`; server typecheck and the affected 104-test suite passed after the final queue changes. Token gates passed. Final-head CI verifies the complete integrated source. - Local aggregate evidence has explicit limits: the general-server invocation overlapped the queue fixes and finished with 11,989 passed, 2 failed, and 80 skipped; both failures are covered by the final 104-test pass. The UI and CLI then passed all 6,184 and 485 tests; the complete 145-file serialized rerun passed all 2470 tests. The shared-package lock fixture passed unchanged on rerun, but the package phase subsequently stopped at an embedded-Postgres bootstrap resource failure. No single pristine green local full aggregate is claimed. - Greptile reviewed `fa66e2bd5` at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13354#issuecomment-5650334242). [Final-head CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733781888/attempts/2): 33 successful checks, 2 conditional skips, including all server, package, UI, browser, runner, typecheck, and build gates. The first attempt hit a preview-readiness/port-collision fixture; its unchanged local control passed 25 tests with 3 skips, and one supported unchanged CI retry passed the affected shard and aggregate gates. ## Risks - Queue recovery must never overlap an old execution. Unknown cleanup state remains blocked. - Driver descriptors are now authoritative for steering; a wrong descriptor disables the action instead of attempting it. - No schema migration or historical status reconciliation is included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, browser testing, and tool use. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6809314a3f |
fix: recover sandbox workspace setup and retries (#13353)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - Sandbox tasks need a workspace and a provider session before work can start. > - A selected Git subfolder is a valid workspace, but it is not a repository fetch source. > - A resumed sandbox must keep one resource identity when the provider fills in its region. > - Workspace reuse does not prove that a provider session has started. > - A recorded continuation must not leave an old failure blocking Retry for a newer attempt. > - This pull request fixes those setup and recovery boundaries while preserving ownership checks. ## Linked Issues or Issue Description **What happened?** A Daytona task failed before provider startup when its project folder was inside a parent Git repository. After the folder was repaired, Retry resumed the same sandbox and verified its workspace sentinel, but workspace preparation reported that the lease was no longer active. A later attempt could fail because the resumed workspace had no provider checkpoint. The old run also kept a recovery projection after the server recorded its explicit successor, hiding Retry for the new failure. **Expected behavior** Sync the selected folder without importing parent files or history. Keep the existing sandbox identity stable. Allow a new provider session only with proof that its exact session has never started. Let the current failed attempt retain Retry once the old recovery has a recorded successor. **Steps to reproduce** 1. Select a subfolder of a Git repository as a project workspace and start a Daytona task. 2. Leave the target region unset. Release its reusable sandbox after setup fails, then resume and realize the workspace. 3. Retry a provider setup that failed before any session directory or checkpoint was created. 4. Record an explicit successor for a native failure, fail that successor during setup, and inspect Retry in the task thread. **Paperclip version or commit** Reproduced against `8d1f0c20a` with new failing regressions before each production fix. **Deployment mode** Local controller with a Daytona environment. Automated tests use deterministic provider fixtures and real filesystem operations. Related: #13338, #13163, #13264, #13349. The open work-folder stack in #13264 includes a broader fresh-session authority change. This patch addresses the independently reproduced startup failure with existing durable bootstrap proof and atomic directory creation. It does not include the work-folder migration or credential changes from that stack. ## What Changed - Classify only the selected Git repository root as a fetch source. Sync subfolders as directories and preserve their enclosing Git ignore rules on upload and restore. Recognize the shared scheduler’s completed non-repository result so ordinary folders still sync; timeout, cancellation, and output-limit failures remain closed. - Remove the placement target from Daytona account cache identity. Keep API endpoint, credential digest, company, environment, driver, and sandbox ID boundaries. - Preserve closed-lease admission until a sentinel-verified resume reopens that same resource. - Show the already-recorded explicit successor of a resolved recovery. Keep the old failure evidence and unresolved holds. No historical status writes occur. - Permit a resumed workspace to create a new session directory only with matching durable identity, zero connections and events, untouched bootstrap commands, no backup, and an absent remote session. Claim the directory atomically. Existing, partial, or ambiguous state still blocks startup. ## Verification - Final head `c43c7403f`: [CI completed successfully](https://github.com/paperclipai/paperclip/actions/runs/34733846820/attempts/2), with 32 successful checks and 2 conditional skips. This includes every server, UI, package, serialized-route, browser, native-runner, typecheck, and build gate. Greptile reviewed the same head at [5/5 with no unresolved findings](https://github.com/paperclipai/paperclip/pull/13353#issuecomment-5650270168). - New regressions failed before each of the four production fixes. Git/archive/restore suites: 148 passed. The scheduler-wrapped non-Git regression also failed before its fix; 135 affected Git/sync/Codex tests then passed. - Daytona plugin: 230 passed, 6 opt-in live tests skipped. Native executor, projection, and TaskChatThread: 494 passed, including existing/partial state, wrong identity, prior connections or turns, backups, unavailable proof, and unresolved recovery controls. - Local full-repository typecheck, build, and token gates passed on the final head. Local CLI: 485 passed. Complete single-worker package rerun: 3,224 passed, 19 skipped. - Local verification is an aggregate with recorded retries, not one pristine green invocation: the general-server run began on `e721a920a` and finished with 12,002 passed, 3 failed, 70 skipped. Its real Codex scheduler failure is fixed above; the socket and workspace-runtime timeout failures passed unchanged in focused reruns. Both Inbox failures passed unchanged in the full 27-test Inbox file; database/shared-package failures passed in the single-worker package rerun. The supplemental local serialized-route rerun remains in progress; all five corresponding final-head CI lanes passed. - The first final-head CI attempt hit a Daytona fixture-readiness race and a signoff-browser heartbeat receipt timeout. One supported unchanged failed-job rerun passed both and the aggregate gates. Live combined user-journey verification is tracked in the related follow-up; this PR's provider tests use deterministic fixtures and real filesystem checks. ## Risks - A selected subfolder uses directory sync and does not carry parent Git history. Its ignored files stay local. - The target region remains a creation setting and part of workspace reuse policy; it does not split the account identity of an existing sandbox. - Incomplete or conflicting provider state still fails closed. This change does not erase a session, infer completed work, bypass a user decision, or replay uncertain actions. - A later remote setup failure can leave a claimed partial session directory. It remains blocked rather than being overwritten. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact hosted model ID and context-window size are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d1f0c20af |
fix: let responsible users choose either AI subscription or API key (#13351)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - AI Connections select the account used for each run. > - A responsible-user binding must follow the person whose work the agent performs. > - The saved sign-in method currently blocks users with another method for the same provider. > - This pull request resolves a personal default by company, user, and provider. > - Each user can use a subscription or API key with the same bot and model. ## Linked Issues or Issue Description Refs #13247, #13248, #13346, #13347. **What happened?** A bot configured with a Claude subscription rejects another responsible user’s Claude API key. Inline repair also limits that person to the original sign-in method. **Expected behavior** The same bot uses each responsible user’s default Claude account, whether it is a subscription or API key. Explicit shared account selections remain fixed. **Steps to reproduce** 1. User A connects a Claude subscription and creates a bot using the responsible user’s connection. 2. User C connects a personal Claude API key. 3. User C runs the same bot. Before this fix, credential resolution fails. ## What Changed - Add a personal provider-default table. Preserve legacy per-method preferences and backfill the most recently updated preference, including unavailable defaults. Repeated migration does not replace a selection. A database trigger propagates old-server default updates without treating new accounts as replacement defaults. - Resolve responsible-user bindings by provider. Retain the method as a wire compatibility hint for old servers. Explicit selections still require the exact method and grant. - Use the selected account’s method for credential isolation, refresh locking, and run attribution. - Update onboarding, agent setup, the picker, and inline task repair. Keep existing authentication components and harness/model settings. - Add mixed-method runtime, migration, repair, and Storybook coverage. Include upstream’s duplicate Anthropic option fix through the base branch. ## Verification - Focused resolver, migration, connection-intent, onboarding, agent setup, model, and connector UI suites: 331 tests passed. - Onboarding and new-agent regression suites passed during the initial focused run. - UI typecheck, token gates, and Storybook build passed. - Live browser checks passed for Claude and Codex API-default execution, switching both back to subscriptions, and both existing shared-account bots. Bot configuration remained unchanged. - One Daytona startup command stalled before Claude launched. The test run was cancelled, its sandbox stopped, and the same account/task passed on retry. Startup cancellation remains a separate environment finding; this PR does not change that command transport. - Browser review: all eight assertions passed in the new mixed-method story, including shared selection, return to responsible-user selection, and unchanged harness/model. - Repository build and typecheck passed after refreshing upstream dependencies. Final resolver and historical rollback verification: 38 tests passed. All latest-head CI gates passed, including server/workspace/serialized suites, browser E2E, build, typecheck, and runner verification. The extra serial local full-suite run was stopped after equivalent CI passed; focused local checks completed. ## Risks - Users with both historical method defaults get their most recently updated preference as the initial provider default. They can change it explicitly in Connections. - A revoked or unavailable default blocks. Connecting an additional account does not silently replace it. - Existing legacy authentication is unchanged. Managed responsible-user bindings intentionally stop pinning a method. - Live staging: the same Claude and Codex bots completed real API-key runs after changing only the personal default, then completed subscription runs after restoring the original defaults. Read-only database verification confirms unchanged bot configuration and actual method attribution. Distinct-user concurrency is covered by automated real-database tests with synthetic credentials, not two live human logins. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, code execution, and browser interaction. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
422287eecd |
fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task messages, provider execution, and task outcomes. > - First-time user tests exposed gaps in recovery, completion permissions, message delivery, and Stop behavior. > - These gaps left usable output hidden, completed work waiting for bookkeeping, or safe work unable to continue. > - This pull request fixes the shared lifecycle and receipt paths while preserving process ownership and action checks. > - Users can continue work with accurate task state and durable messages. ## Linked Issues or Issue Description **What happened?** A stopped local Codex execution could remain blocked even after its processes had stopped and its complete transcript proved that no external action needed replay. Claude under Conservative permissions could fail to call task completion tools. Recovery could reuse an assistant item ID and overwrite prior output. A delivered comment could remain marked uncertain after navigation. Stop could look like Pause or a new recovery incident. Workspace contention could look like cancellation. A direct reply reopening Done could enter a clarification loop. **Expected behavior** Recover automatically only with verified termination and complete action receipts. Preserve answers and messages. Keep task completion available under Conservative permissions without broad tool access. Show crashes as Blocked, actual human decisions as In Review, and ordinary workspace contention as waiting. Stop the current response and allow a new direction. **Steps to reproduce** 1. Create ordinary response tasks with local Codex and Claude Code, then send follow-up messages through the task composer. 2. Interrupt a disposable local Codex runner during text-only work. Verify automatic continuation and retained output. 3. Stop a response, send a new request, answer a clarification, and reopen completed work with another message. 4. Navigate or reload while a comment submission is pending. Confirm the exact persisted request receipt settles it without removing newer draft text. 5. Run two tasks in a shared Daytona workspace. Confirm waiting does not appear as failure. **Paperclip version or commit** Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`. Current integration base: `6cef9743c`. Both operator-interruption and workspace-waiting guards are preserved; native restart and legacy permission rules remain documented. **Deployment mode** An isolated source-built test-drive instance, with real local Codex and Claude Code providers and disposable Daytona environments. Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163. This PR addresses additional failures from ordinary task journeys, including controller restart handoff and repeated warm sandbox setup. Historical task status reconciliation is excluded. ## What Changed - Persist runner ownership immediately at spawn and resume an explicitly adopted runner even when the controller crashed before the first driver checkpoint. Detach the controller safely across graceful restarts, including session startup. Prevent an old finalizer from suspending or signaling an adopted runner. Checkpoint idle warm sessions before shutdown. Preserve the same run and queued follow-up messages. - Scope saved legacy queue successor checks to the queue owner while preserving ordinary task locks, operator identity, assignment gates, and exactly-once delivery. - Preserve managed Codex credential files when an old session is detached for restart; normal owned cleanup still copies refreshed auth back and removes the scoped copy. - Reuse the bound warm shared sandbox and fully verify an existing staged provider pack before using it. This avoids repeated uploads when the pack is already valid. - Add a narrow local Codex replacement path with stopped-process proof, a closed transcript inventory, exact completion receipts, and fresh-session lineage. Preserve no-replay holds when evidence is incomplete. Recovery may clear only the same run's recorded Blocked status version; manual re-blocking and dependency changes invalidate that receipt, while queued comments do not. Later blocks stop scheduled, queued, and final dispatch; queued/final checks re-read dependencies even when the task status stays In Progress. - Permit only task delivery and human-input tools through the isolated Claude runner's exact task bridge. - Scope assistant item identity to the provider turn and ignore only authority-free Codex skill-change notifications during startup. - Reconcile composer submissions by client request ID across response loss, navigation, and reload. Retain text typed during delivery. - Keep acknowledged run-only Stop neutral and show workspace contention as waiting. Project exhausted native failures as Blocked. - Restore the guarded task-page retry action for failed legacy runs, including the server-supported explicit new-attempt path for stopped conversation adapters. Preserve native/process recovery holds and avoid promising Retry while a decision or execution gate hides it. - Refresh delivered artifacts and handle direct user replies that reopen completed work without a clarification loop. - Check the embedded PostgreSQL PID, data directory, and actual port before connecting or migrating. - Document accepted behavior and add focused regressions at lifecycle, route, transcript, and UI boundaries. ## Verification - Final head `fece606ac2` passes the complete GitHub CI matrix: **34 green checks, two expected Storybook skips, no failures or pending checks**, including `ci / verify`, `ci / e2e`, full runner verification, typecheck, build, every server/workspace shard, and all browser shards. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34727183287). Greptile is **5/5 with no open findings**. The final two commits only refine test fixtures; both affected suites pass 24/24 locally and in CI, with server typecheck green. - Complete local Vitest coverage uses the canonical groups/shards: all 635 general server suites, all 145 serialized suites, and all workspace packages. The aggregate began on `0a8001c18` while the final queue fix arrived: 23,903 passed, five failed, 87 skipped. The five port/socket/timing failures passed unchanged in follow-ups (60 tests in the exposure/file suites and 412 tests covering the serialized failures and unrun tails). The final queue/operator-identity suites separately passed 52/52. This is aggregate coverage plus explicit reruns, not a pristine single-command final-head run. - After integration with current master, queue/operator-identity/continuation suites passed 162/162 and affected UI suites passed 140/140. ACP Stop/continuation and legacy task/Inbox/message browser suites passed 9/9, including both task recovery Retry and thread Try again, automatic saved-message delivery, exactly one new run, Done, and retained output after reload. The default process Stop/Pause/Resume browser case passed (the native-provider case is opt-in and skipped by default). The complete Board attachment/receipt browser suite passed 11/11 on a disposable instance, covering both composers, exact receipts after lost responses, no replay, bound attachments, and newer drafts after reload. - Blocking-intent regressions cover pre-existing Blocked, a mismatched run/cause, an explicit manual re-block, changed dependencies, a queued comment after failure, and a block arriving between scheduling and provider dispatch. The negative cases reproduced before the fix. All 478 affected executor/recovery/dispatch tests passed; both database suites ran separately after availability-probe skips in the first combined command. The final late-dependency check passed all 143 affected recovery/dispatch tests (zero skips) after two new negative cases reproduced the bug. - Focused runtime regressions cover awaited runner ownership publication, authenticated adoption before the first checkpoint, old-finalizer detachment, idle and busy warm-session shutdown, rejected checkpoint propagation, provider-pack verification, and managed-Codex credential preservation. Four managed credential detachment cases reproduced the bug before the fix; normal owned cleanup still succeeds exactly once. - Live local Claude: SIGKILL 2.6 seconds into startup recovered the same run automatically in 53 seconds, then a normal follow-up completed in 24 seconds. SIGTERM 2.5 seconds into startup preserved the same run (54 seconds) and its queued follow-up (21 seconds). Answers remained visible and the task reached Done. - Live Claude Daytona: a warm follow-up retained its sandbox and fell from 121 seconds to 44 seconds. A separate cold turn took 127 seconds; after controller shutdown and checkpointing, its follow-up completed in 33 seconds with the same sandbox, workspace, native session, and runner. Both answers remained visible and the task was Done. - Other live journeys covered task completion and follow-up with local and Daytona Codex, local Codex crash recovery, Stop then new direction, clarification response, live artifact refresh, and shared-workspace waiting. - Validation limits: the opt-in native composer Stop/Pause→subtree Resume fixture exposes terminal/result ordering and subtree-cancellation attribution bugs that can leave a child task blocked; that new finding is assigned to a separate follow-up and is not claimed fixed here. Default CI skips this optional native-provider fixture. Managed-Codex credential handoff and the queue-agent integration use automated regression evidence. Cold custom provider-pack uploads still add startup latency. ## Risks - Automatic replacement remains deliberately narrow: local Codex, verified stopped identities, unchanged retained state, and a complete text/completion-only turn. Unknown actions, partial history, or changed ownership remain blocked. - Claude completion permission handling changes an upstream package patch. The exact isolated task bridge must remain pinned; unrelated tools keep their existing permissions. - New task failure projection changes user-visible status. No historical status backfill or database migration is included. - This is a broad lifecycle fix across server and UI. Live proof covers graceful local Claude restart during startup and idle Claude Daytona session recovery across controller shutdown. Live abrupt SIGKILL during local Claude startup also recovered the same run. Unknown ownership or missing action evidence still blocks reuse. Cold custom provider-pack uploads still add startup latency; this change avoids unnecessary repeat uploads. ## Model Used OpenAI GPT-6 (Codex), with reasoning, code execution, browser automation, and tool use. The exact hosted model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
13bae6fa21 |
fix: reject unsupported REST tool connections without stdio validation (#13346)
## Thinking Path > - Paperclip manages AI agents and their connections. > - Connection checks must use the configured transport. > - The tool service treated every remaining transport as local stdio. > - Anthropic's old REST method therefore failed with a templateId error. A REST connection with a valid stdio template could incorrectly pass. > - Anthropic now has a supported AI-account flow. This pull request removes its obsolete REST setup option and limits stdio checks to stdio connections. > - Users can connect an AI account, and existing unsupported connections receive an accurate error. ## Linked Issues or Issue Description Related: #13248 added the supported AI-account flow. Searches for related REST health and templateId bugs found no duplicate fix. **What happened?** The Anthropic REST API-key connection showed `Local stdio MCP connections must use an approved templateId`. Health checks and catalog discovery both fell through to the local stdio path. A REST connection with an approved template could report success and expose the template's catalog without a REST integration. **Expected behavior** Only local stdio connections use command templates. Unsupported transports return an accurate HTTP 422 error. New Anthropic accounts use the supported runtime authentication flow. **Steps to reproduce** 1. Check out the test-only commit `924e6e85a` in a separate worktree and install dependencies. 2. Run `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts -t 'unsupported REST|obsolete Anthropic'`. 3. The tests exercise saved Anthropic REST configuration and an unsupported REST connection containing an approved stdio template. They cover health checks and catalog discovery separately. 4. Run the same tests on the fix commit. They pass. The full affected files also pass. **Paperclip version or commit** Reproduced against master `6cef9743c`. **Deployment mode** Server transport handling. Reproduced with an isolated embedded PostgreSQL test database. No provider account or live credentials are required. ## What Changed - Restrict stdio health checks and tool discovery to `local_stdio`. - Return and audit `tool_connection_transport_unsupported` with HTTP 422 for unsupported tool transports. - Remove Anthropic's obsolete REST method from the generated catalog and its durable ingestion source. Keep its subscription and API-key AI methods. - Cover the reported error, false-success case, rejected obsolete setup, connection removal, and the UI's AI-account submission path. - Replace impossible reconnect forms for removed methods with supported setup, while preserving connection removal. - Preserve AI-versus-tool intent isolation for legacy requests and reject new unsupported Anthropic tool requests. - Document recovery for existing unsupported connections. ## Verification - Clean-worktree red/green: the same command failed all six regression cases at `924e6e85a` and passed all six at `4d3de9de0`. The failing run includes the reported templateId error. - Green: all 555 tests across the six affected test files passed. - Recovery UI red/green: three added cases failed before the recovery fix and passed afterward; all 200 tests across setup, detail, and advanced controls passed. - After the recovery UI update, UI typecheck/build and token gates passed again. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm check:token-gates` — passed. - Catalog regeneration — passed with the documented `PAPERCLIP_CONTENT_TEMPLATES` override for the local capture corpus. - Full CI on `69fb31fd4` — passed all general and serialized test shards, browser shards, typecheck, build, runner verification, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/34726975425. - The local serial `pnpm test:run` was stopped after the fixture correction superseded that run; full-suite verification above comes from CI. All 555 affected tests passed locally, including all 17 connection-intent tests after the correction. - Greptile — 5/5, successful check on final commit `69fb31fd4`, no unresolved findings. - No live Anthropic validation was performed. The UI regression uses a fake key and a mocked AI-account response. ## Risks Existing obsolete REST connections remain in needs-attention state. Users must add an account through the supported flow and remove the old connection. Credentials and grants are not transferred automatically. Removal remains covered. The specialized AgentMail and Composio paths keep their existing behavior. There are no schema or permission changes. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window capacity are not exposed in this session. Used reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
df984cbc2c |
fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task conversations save messages that arrive during an active turn. > - Legacy adapters deliver these messages in a later turn. > - A run can stop before the saved queue is delivered. > - The Interrupt button previously required an active run, so it could not release this queue. > - This pull request lets a board operator send the saved queue after the run stops and retries queues missed during finalization. > - Manual dispatch must use the clicking operator and must not require permission to create agents. > - The task can continue without a duplicate message or a second execution owner. ## Linked Issues or Issue Description **What happened?** A legacy task retained a queued message after its run stopped. Interrupt was disabled because the queue had no active target. Finalization and deferred message admission can also leave a queue without a successor. **Expected behavior** Interrupt sends the saved messages when no runner is active. Messages that arrive during normal completion are delivered automatically. An uncertain previous execution still requires proof that its process or sandbox stopped. **Steps to reproduce** 1. Queue a user message during a legacy conversation turn. 2. Let the turn stop or simulate a server restart before queue promotion. 3. Open the task with a deferred queue and no active run. 4. Try Interrupt. Before this change, the button is disabled. **Paperclip version or commit** Reproduced against `8f40b4ad4`. **Deployment mode** Legacy conversation adapter. The same persisted queue state is covered with an isolated PostgreSQL fixture. Related public work: #13275 adds active legacy interruption. #13291 addresses automatic sandbox conversation recovery. This change handles explicit saved-queue delivery and late queue promotion. ## What Changed - Accept a null Interrupt target while retaining queue identity, revision, company, and assignee checks. - Save the operator's request on the existing queue. Reuse normal admission after verified stop, including older messages, different authors, and queues whose original wake came from the system. - Strip interruption authority from caller-supplied wake payloads. Only the board queue route can persist that authority. - Retry durable interruption requests after restart and deferred queues after legacy cleanup. - Let an explicit Interrupt retry cleanup for its stopped run, including old ephemeral leases that recorded success without a provider stop receipt. Preserve retained resources, other lease owners, and the automatic retry limit. - Preserve the server's waiting explanation when normalizing and combining queue entries. - Revalidate the consumed board queue receipt at dispatch so a different message author does not cause setup failure. - Use the Interrupt user's execution identity for the new run. Preserve original message authors. Validate the receipt independently at startup and inherit the resulting identity on retry. - Persist authenticated board authority for ordinary manual wakes too. Adopting someone else's queued messages cannot switch a manual run to that author's permissions. Strip caller-supplied authority markers and retain private conversation ownership checks. - Keep the clicking user when a manual wake is merged into an older deferred receipt. Update its requester and payload in the same transaction. - Use the same current-queue/revision API on task details and pipeline conversations; show Interrupt after a legacy target stops. - Keep manual wakes out of active runs, including unscoped agent wakes. They receive their own execution identity; a matching receipt requester is not sufficient because an exact retry can retain a different originating identity. - Authorize both existing-agent wake endpoints with `agent:wake`, available to active non-viewer company members. Keep `agents:create` for hiring. Validate the stored task and current assignee before an exact task retry. - Reject viewer Interrupt requests before saving intent or stopping execution. Keep external chat retry authorization and per-action agent/user permission checks. - Preserve edits and discards until dispatch. Prevent another queue promotion when the same agent already has a successor. Keep independent reviewer recovery available. - Suppress cancelled/failed run toasts for intentional operator interruption. Keep ordinary runtime error notices. - Add UI, route, admission, restart, successor ownership, and toast regression tests. Document the behavior. - Reuse the existing socket reservation helper for both credential-quorum test cases after CI exposed an ambient-port collision. This changes test preparation only; production credential staging is still called exactly once. ## Verification - Failing regression tests reproduced the message-author identity bug and an operator's `agents:create` rejection before the fixes. - All 316 focused tests pass across eight route, queue, identity, authorization, continuation, and responsible-user suites, including the 44-test rerun of queue admission and actual startup after the final manual-wake restriction. Regressions reproduce cross-user merging both with and without a task, and same-requester receipt ambiguity. The cross-company existence guard also passes both tests. - Startup integration tests reach adapter execution under the clicking operator and retain that identity through follow-up. Coverage includes mixed authors, adopted queues, system-origin queues, restarts, forged or stale receipts, viewers, suspended memberships, changed assignees, private conversations, and caller-supplied authority markers. - The earlier queue/cleanup/UI regression suite passed 402 tests. The final review corrections pass another 180 tests across queue admission/persistence, real heartbeat startup, UI API, conversation rendering, and pipeline suites. Regression tests reproduced both review findings before correction. The final head has a 5/5 review with no unresolved threads. Full CI passes on `c2002979c`, including every general and serialized server shard, all browser shards, Paperclip Runner verification, typecheck, build, canary dry run, and the aggregate gates. - Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after the final application changes. CI identified an outdated task-page API mock after the shared helper extraction; the fixture now exercises the real helper, and all 131 task-page/API tests pass. The final application build passes with the additional manual-wake restriction. - CI exposed a pre-existing port collision in the Codex credential-quorum fixture. It reproduced locally; both listener cases now use the existing bounded reservation helper. All 41 credential tests pass on rerun. One intervening local run hit a separate ambient bind collision in the two-occupied-port case. - The full local `pnpm test:run` attempt was stopped after host contention caused focused-suite timeouts. The affected focused tests passed on rerun. An expiring trace fixture and a missing private-conversation state were corrected. The successful full CI run is the complete-suite verification. - Hosted Interrupt previously cleared the original queue and produced exactly one successor with neutral interruption feedback. It exposed the dispatch authorization defect. Retry on that earlier build was rejected for missing `agents:create` before creating another run. - Deployed the final application build (`38257f391`) to the scoped hosted instance and verified readiness. The latest PR commit changes only the credential test fixture; application code matches that deployment. A live Retry by the same operator without `agents:create` created one successor attributed to that operator, passing the former dispatch permission gate. Startup then stopped at `configuration_incomplete` because that operator has not configured their required personal Claude Code OAuth secret; the post-deployment run page confirms the operator identity and no provider work started, and the My secrets UI still shows the token as not set. Provider execution remains unverified pending that credential. No permission grants or credentials were changed. ## Risks Queue admission and finalization can race. The task lock, durable queue receipt, current comment IDs, and successor guard prevent duplicate dispatch. Process and lease stop checks, task pauses, approvals, ownership, and budgets remain in force. The API change only allows null on legacy Interrupt; native steering still requires an active run. No schema migration is required. Active non-viewer board members can now invoke existing agents without agent-creation permission. Agent self-invocation rules, raw provider-trace admin access, task retry scope, external chat authorization, and action-specific user/agent permissions remain enforced. ## Model Used OpenAI GPT-6 through Codex. The session does not expose an exact backend model ID or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8d6232e7b0 |
feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect provider accounts during onboarding and agent setup. > - They should reuse and manage those accounts through the existing Connectors interface. > - A second login wizard would diverge from the established provider workflows. > - This pull request composes the existing sign-in components into Connections and agent configuration. > - Users can select accounts without changing their agent's harness or model. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add compact AI-account management to the existing Connectors pages. - Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome, and authentication controllers. - Add the shared connection picker to agent setup/settings and task requests. - Preserve onboarding's sequence and reuse existing accounts. - Add local-login recovery, retry, cancellation, and React StrictMode handling. - Add interactive Storybook scenarios, design-guide examples, and app acceptance checks. This is part 2 of the AI Connections change. The runtime foundation in #13247 is merged. This PR now targets master. ## Verification - Updated against master `47ded8bf9`, including the landed runtime foundation and upstream task-search changes. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All CI test, browser, build, packaging, and runner jobs passed on final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh Greptile review is 5/5, the security scan passed, and there are no unresolved review threads. The final CI aggregate gates passed. - Local general-server coverage passed 11,804 tests; three port-collision failures passed in an isolated 25-test rerun. All 6,111 UI tests passed. CLI coverage passed 484 tests; its remaining doctor test requires port 3199, which is occupied by an unrelated report server on this Mac. The complete CLI suite passed in CI. - Live browser testing verified Codex API-key reconnect inside a task card on desktop and phone. Real provider runs resumed and completed with unchanged connection/grant identity and agent routing. - Tested opening, cancelling, reopening, and completing connection creation. A regression confirms Connect another account cannot submit the new-agent form or copy provider keys into agent settings. - Added shared inline repair, automatic local sign-in checks, and responsive connection dialogs. Standalone Daytona installation ignores workspace configuration and suppresses dependency scripts. Its standalone build also passed with CI's exact pnpm 9.15.4. - Destructive live tests are excluded by default. Explicit opt-in, local deployment checks, and matching disposable fixture identities are required before any mutation. ## Risks - Local Codex/Grok creation requires the connection-specific terminal login command. - Browser sign-in uses the existing supported-environment controllers. - This update verifies live local Claude detection and Codex API-key task repair. New subscription authorization/refresh and independent-human/native-runner isolation were not reverified in this update. - No agent automatically adopts managed Connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ed50a39c3f |
fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c9021c6721 |
fix: require explicit native completion reviews (#13314)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs report their outcome through paperclip_finish. > - The server previously turned incomplete reports into human approval requests. > - Those requests could block a later successful run, even when no person had requested review. > - This pull request creates review cards only for explicit attention requests and withdraws proven old fallback cards. > - Agents receive useful completion feedback, while explicit approval gates and task state protections remain in force. ## Linked Issues or Issue Description Related: #13266 removed reviews caused by policy upgrades. This change removes the separate completion fallback. **What happened?** An agent reported needs_review while waiting for checks without requesting a human decision. Paperclip created a generic Native completion review. A later successful report could not complete the task because that old card remained pending. **Expected behavior** Ordinary low-risk work completes after a valid done report, a successful run, and workspace finalization. Incomplete work stays with the agent. Explicit approval requests remain visible and must be resolved. **Steps to reproduce** 1. Complete a native run with needs_review and no attention requests. 2. Continue the task and submit a successful done report. 3. Observe that the old implementation leaves the task in review behind a generic confirmation card. ## What Changed - Require explicit attention requests to create native review cards. Route each request independently and preserve pending or declined decisions. - Withdraw only pending system cards with matching old decision, assessment, effect, contract, and prompt provenance. Preserve history and explicit or answered requests. - Reassess an affected current result without overwriting later task edits, runs, contracts, or workspace failures. - Return pending approval links and required actions through the completion tool. Reject contradictory done reports and empty review requests before accepting a result. - Allow one corrective continuation for incomplete results, then expose a recovery action. - Update status fixtures, database regressions, runner tests, and the completion contract documentation. ## Verification - `pnpm -r typecheck` passed after merging current master. - `pnpm build` passed after merging current master. - The combined branch passed 89 completion and Agent Chat tests. Other targeted tests passed: 170 external-chat and reconciliation tests; 50 runner-resume and control-plane tests; 13 arbiter tests; 7 chat delivery tests; 21 runner completion and runtime-context tests. - The full local test attempt exposed old review fixtures and a missing fake-provider binary. The fixtures are fixed and the helper is built. All affected suites pass in fresh reruns. The timing-sensitive Discord test also passed on rerun. - All latest-head CI checks passed, including build, typecheck, general and serialized tests, runner verification, browser tests, and canary dry run. Greptile is 5/5 with zero unresolved comments. ## Risks - Cleanup changes existing pending cards. It requires exact system provenance and only applies to low-risk agent-claim contracts. It does not delete history or dismiss explicit requests. - Status still commits after the turn and workspace finalization. Completion feedback reports current constraints and does not claim an early status commit. - Incomplete reports now request a bounded corrective run instead of an automatic approval. Repeated failures expose recovery. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and test tools. The exact deployment identifier and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5cc7784986 |
test: remove cold executable reads from runner integrity deadlines (#13301)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runner protocol tests verify that invalid authenticated input fails closed. > - Cloud readiness requires those tests to pass before a source commit can deploy. > - Two test fixtures hash the host Node executable twice even though its process launcher is synthetic. > - Cold reads of a large Linux executable consume time unrelated to protocol failure handling. > - This PR gives both synthetic runner fixtures tiny real artifacts and keeps their existing deadlines. > - Real authentication, encrypted frames, and request and notification failures remain covered. ## Linked Issues or Issue Description **What happened?** Cloud readiness repeatedly failed in `DurablePrpControlPlane > promptly fails the real transport request and notification paths on authenticated bad semantic input (throwing observer: false)` with `Test timed out in 5000ms`. The image built, but the failed source check prevented deployment. See runs [34656885170](https://github.com/paperclipai/paperclip/actions/runs/34656885170) and [34657111560](https://github.com/paperclipai/paperclip/actions/runs/34657111560). The first case took 5.6–11.4 seconds; the following case took about 0.4 seconds. The original test reads and hashes `process.execPath` once itself and once through the real transport. The Node executable is 122,678,944 bytes in the local Linux Node 24 container, versus 68,672 bytes on the development Mac. Instrumented Mac runs pass; cold executable reads are a likely cause of the CI-only timeout. **Expected behavior** The deadline should measure real protocol rejection and consumer failure, without reading a large unrelated executable as test fixture data. **Steps to reproduce** 1. Run the named test with the real transport and synthetic process launcher. 2. For a deterministic probe, inject a six-second delay into the first `readFileSync(process.execPath)` call. 3. The original test exceeds its existing five-second deadline. Both cases pass after this change because neither reads the host executable. The probe is temporary instrumentation, not part of this commit. **Paperclip version or commit** f12b647ae; the same failure also occurred on |
||
|
|
dbf5ea432d |
fix: protect starting runs during overlapping deployments (#13285)
## Thinking Path > - Paperclip controls agent work across service deployments. > - A run can provision a remote sandbox before a process or invocation event exists. > - Each container previously treated its own missing process handle as proof that the run was orphaned. > - Overlapping deployments could therefore fail a run owned by another container. > - This pull request records and renews a controller lease before provisioning. > - A recovery worker must revoke an expired owner before it finalizes the run. ## Linked Issues or Issue Description Merged PR #13272 records startup adapter identity and restores explicit user continuation. This PR adds controller ownership on top of current master. Refs #7997 and #10442 for related replica and ownership problems. Related #13138 addresses silence and detached local processes; this change does not infer death from silence. **What happened?** During an overlapping hosted service deployment, a new container reaped a legacy conversation run that another container was provisioning. The run had no PID or adapter invocation yet. **Expected behavior** A live controller keeps its run. After controller loss, one recovery worker takes cleanup authority and the old controller cannot dispatch further work. **Steps to reproduce** Claim a legacy run in controller A. Start controller B against the same database before A finishes provisioning. Run the startup reaper in B. **Paperclip version or commit** Observed on `663c44cb2b9c28336d38d0b4a6971f4f1964bce6` in a hosted Railway deployment with a Daytona environment. ## What Changed - Add nullable controller boot ID, lease deadline, and execution stage columns. Claim ownership in the queued-to-running update. - Renew ownership independently of run output. Abort and reject dispatch if renewal fails. - Serialize reaper revocation against renewal. Let unfinished recovery claims expire after a restart. - Restrict graceful shutdown to legacy runs owned by the current controller. - Hand ownership back to the existing native coordinator when runtime selection becomes native. - Add twelve database regressions and document the lease contract. Update the task-drain regression to require controller expiry before reaping. ## Verification - `pnpm exec vitest run server/src/services/legacy-controller-lease.test.ts server/src/__tests__/heartbeat-task-drain-admission-release.test.ts`: 14 passed after rebasing onto master (`f12b647ae`). - Queue-interruption regressions in `heartbeat-process-recovery.test.ts`: 2 passed after preserving the new cleanup promotion from #13275. - `pnpm --filter @paperclipai/server exec tsc --noEmit`: passed after rebuilding runner TypeScript outputs for the updated master. Broad local tests are omitted at the maintainer’s request; CI owns broad coverage. - Latest-head CI passed on `f255e8e4d5ab2b24b12638a02434e6aa8a2285c5`: [run 34658248569](https://github.com/paperclipai/paperclip/actions/runs/34658248569). All test shards, browser suites, typecheck, build, canary, and security checks passed. Greptile is 5/5 with no unresolved review threads. ## Risks - Additive, idempotent migration; historical rows retain the previous recovery behavior. - Database unavailability aborts new dispatch rather than permitting an unfenced controller to continue. - Lease expiry is permission to clean up, not evidence that remote inference stopped. Follow-up PRs add persistent cleanup and automatic continuation. - Mixed-version deployment still includes old binaries whose reapers do not understand controller leases. ## Model Used OpenAI GPT-6 through Codex, using reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f12b647ae8 |
fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A task can collect more messages while its agent works. > - Legacy runners must stop the active process before they can receive those messages. > - The old Interrupt action cancelled the run but could leave the queue idle and hidden. > - Codex could also classify a cancelled run as successful or start a fresh process after cancellation. > - This pull request joins cancellation, preserves the provider session, and dispatches the current queue after cleanup. > - The benefit is reliable interruption with the saved message order, edits, and deletions. ## Linked Issues or Issue Description **What happened?** Interrupt could strand a legacy message queue. The UI could hide pending messages after the run stopped. A Codex signal exit could race the cancellation write. A stale session warning could also trigger a fresh process after an interrupted resume. **Expected behavior** Interrupt stops the active turn and sends the remaining messages once, in their saved order. Deleted messages stay deleted. An interrupted Codex turn keeps its session and does not restart itself. **Steps to reproduce** 1. Assign a task to a legacy Codex agent that runs a long command. 2. Queue three messages. Edit one, discard another, and move the last message first. 3. Click Interrupt in the queue. 4. Repeat the interruption while the resumed session runs another command. Related work: Refs #13160, which moves native queue steering into the wake-queue module. This change fixes legacy interruption and keeps native steering unchanged. ## What Changed - Add a revision-checked, company-scoped endpoint for legacy queue interruption. - Promote only the requested queue after the provider stops and releases its lease. Retry its persisted interrupt intent from the scheduler after a promotion error or server restart. - Keep pending legacy queues visible after a run stops. Use server state for the interrupt result. - Serialize owned process cancellation before classifying the adapter result. Preserve late session and log metadata. Acknowledge cancellation only when an actual process or process group was owned; scheduler placeholders retain their normal release policy. - Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the session has started. - Add cancellation race, multi-actor queue order, durable retry, resume fallback, and stale request regression tests. Document the behavior. ## Verification - Real browser tests passed with legacy Codex CLI and ACP engines, using Codex 0.153.4 and gpt-5.6-sol. - All three automated ACP browser scenarios passed locally: immediate Interrupt delivery, no replay of an unfinished write, and pause requiring Resume. Updated the old test expectation that required a separate “go” after Interrupt. - Browser tests covered queued edits, deletion, reordering, deleting the final message, and repeated interruption. - Two consecutive CLI interrupts kept one provider session. Both stopped processes exited. The final message arrived once. - `pnpm -r typecheck` passed. - `pnpm check:token-gates` passed. - All 346 post-review scheduling, recovery, queue-route, archived-company, worktree-suppression, and stale-queue regression tests passed. - All 318 process-recovery and durable-chat tests passed after the final cancellation guard. - Codex adapter, queue UI, issue-page, and OpenAPI contract tests passed. - `pnpm build` passed. - Full local suite coverage completed with `PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI shards: 618 general server suites, all 145 serialized server suites, and all workspace groups. Every failing suite passed a targeted rerun after the fixes, rebuilding the native test fixture, correcting macOS temporary-path setup, or retrying setup/timing failures. Existing skips remain. - The original monolithic run reported failures before the final fixes; its failed suites were rerun rather than rerunning all 618 suites again. The final process-recovery/durable-chat regression run passed all 318 tests. - All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: [run 34654820774, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2), including typecheck, build, all test shards, E2E, and canary. The signoff and Cursor sandbox tests each hit a timeout in the initial attempt; both suites passed locally, and both failed shards passed their single CI rerun. All three corrected ACP browser scenarios passed in CI. - Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no open review threads. ## Risks Cancellation order affects local adapters. The tests cover signal exits, graceful exits, adapter exceptions, termination errors, and cancellation write errors. Embedded adapters keep their cancellation controls. Ordinary run cancellation and task pause keep their distinct queue policies. No database migration is required. ## Model Used OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code execution. The exact serving model ID and context-window size are not exposed in this session. The live test runner used OpenAI gpt-5.6-sol. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9031516a7e |
fix: recover legacy Daytona startup failures from task and inbox (#13272)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Legacy conversation adapters can run in Daytona sandboxes. > - A server restart during provisioning can occur before the invocation event exists. > - Recovery then lacks the old adapter identity and leaves a hold that ordinary user retries cannot clear. > - A remote launch can also fail when its host relay looks for Node in the sandbox PATH. > - This pull request records the adapter at claim time and restores explicit user continuation after verified cleanup. > - Users can recover from the task or inbox while the failed run and uncertain action history remain intact. ## Linked Issues or Issue Description Refs #13237, #13239, #13254. Those changes cover recorded conversation runs, native user continuation, and explicit remote Stop. This change covers legacy failure before `adapter.invoke` and exact task/inbox Retry. Refs #9771 for overlapping generated-command quoting. This change also supplies the absolute host Node executable. Refs #13163 and #13264 for the separate native restart and retained-workspace work. **What happened?** A legacy Daytona run interrupted during provisioning became `process_lost` without an invocation event. Recovery preserved an execution hold, and Retry or a new task reply could not resume it. Cleanup could also run before the Daytona plugin was ready. On a macOS host, a subsequent ACP relay launch failed with `env: node: No such file or directory` because the remote launch environment did not contain the host Node path. **Expected behavior** An interrupted conversation can continue after its previous execution stops. Explicit Retry and new user replies should start a fresh turn with the task history. Cleanup failures must remain visible and recoverable. The host relay must use the host Node executable. **Steps to reproduce** 1. Use a legacy Claude adapter with a Daytona environment. 2. Interrupt the server after it acquires the sandbox lease and before it records `adapter.invoke`. 3. Restart and inspect the task hold. 4. Retry from the task or inbox, or send a new task reply. 5. Confirm the old sandbox has stopped and one new response arrives. **Paperclip version or commit** Reproduced from master at `3bafac12f796fbea02e609e1074a9639f872e9c4`. The branch is rebased on `51b0e01ea`, including #13261 and #13270. **Deployment mode** Built from source on macOS with a real Daytona sandbox and the legacy Claude ACP adapter. ## What Changed - Count new browser specs with the scheduler's median duration in the shard-balance check. This fixes a false policy failure after new specs arrive from both branches. The balance threshold is unchanged. - Persist server-owned adapter identity in the queued-to-running claim before provisioning starts. - Wait for provider plugin startup before restart cleanup. Keep failed cleanup leases as active ownership blockers. - Admit exact board retries and new user comments after verified termination. Retain the old run, task history, approvals, and unknown action outcomes. - Adopt repeated Retry requests. Permit one scoped cleanup attempt per explicit user Retry after the automatic limit, with an activity record. A later user Retry can recover after a transient provider failure; automatic attempts remain capped. - Resume replies deferred during cleanup, including historical legacy startup failures. - Launch the host ACP relay through the absolute host Node executable. - Add a task-level Retry button and return actionable blockers when retry admission is refused. - Add database regressions and three browser recovery journeys. Exclude installed third-party dependency skills from the shipped-skill audit. ## Verification - Current head: `d23c84181`, rebased on `51b0e01ea`. Conflict resolution retains the saved-message recovery, local stop receipts, and wait reasons from #13270 alongside exact legacy Retry support. - Real Daytona: interrupted the server after lease acquisition and before adapter invocation. Restart cleanup confirmed provider termination. Task Retry cleared a seeded historical hold and a real Claude agent returned `Recovery verified.` in the task. Removed the disposable sandbox and environment after testing. - All three browser recovery journeys passed again after the final rebase. Task Retry, Inbox Retry, and a new reply each produced one fresh successor, completed the task, preserved the failed run, and retained the answer after reload. - All 29 e2e/server shard-partition tests passed. The balance check now uses the scheduler's median fallback for unmeasured specs, with the same balance threshold. - Server typecheck passed after rebuilding the generated runner dependencies. The combined recovery/route run passed 136 of 137 tests. Its remaining route test timed out during the first cold module import at its explicit 10-second limit; an isolated rerun reproduced that timeout and passed the other 51 route cases. The complete CI suite passed on this head. The same route file passed all 52 cases in CI, including the first cold import in 7.5 seconds. - Before the final rebase, recursive typecheck, full build, UI token gates, 132 targeted server tests, and the complete [CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34650004085) passed. The subsequent CI failure was the shard-balance accounting mismatch fixed here. - Greptile reviewed `d23c84181` at 5/5 with no outstanding actionable findings. The complete [current CI workflow](https://github.com/paperclipai/paperclip/actions/runs/34653327949) passed on attempt 2. All test, typecheck, build, and canary jobs passed on the first attempt. Docker setup timed out fetching BuildKit from Docker Hub; retrying that job and its dependent aggregate succeeded. ## Risks - Recovery admission changes executable authority. Company, task, agent, user, approvals, process ownership, and provider termination checks remain required. - Explicit continuation starts a fresh conversation with history. It does not certify unknown external action outcomes or rerun non-conversation adapters automatically. - Changing task status alone does not clear an execution hold. The task now offers an explicit Retry action. - Historical adapter claims and invocation events take precedence over current agent settings. Known process or webhook runs retain their hold. Pre-upgrade rows with no adapter evidence may receive only a new explicit user turn after termination proof; they do not become eligible for automatic replay. - No schema migration or sandbox-image change is required. This branch has not been deployed to production. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution, browser automation, and test execution. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
51b0e01ead |
fix: resume saved user messages after execution recovery (#13270)
Preserve verified native process-stop evidence and retry saved user messages through normal continuation admission after recovery cleanup. Show the current wait reason and serialize delivery so a saved message starts one fresh turn. Validated with 410 focused tests, typecheck, build, token gates, all PR CI checks, and Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
250deab910 |
fix(runner): keep healthy native sessions alive (#13261)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native sessions keep a provider process and its work alive across control-plane operations. > - A hidden 15-minute turn deadline stopped work even when the agent timeout was zero. > - A one-hour runner lifetime and fixed connection lease added two more limits. > - Recovery also rejected goal commands because it reconstructed their startup summary with the wrong protocol version. > - This pull request removes implicit duration limits and renews authenticated leases in the harness. > - Healthy sessions can continue without model action or a user interface change. ## Linked Issues or Issue Description Refs #13092 and #12845. Related: #13163 covers sandbox recovery after app restarts; this change covers session duration and lease renewal. **What happened?** A native Codex session stopped after 15 minutes while a tool was still running. The agent had `timeoutSec: 0`. Recovery then rejected a `session.goal.get` startup command with `invalid provider startup ownership fence`. **Expected behavior** An unlimited session keeps working while its provider and authenticated controller remain healthy. Lease maintenance is transparent. Explicit timeouts, cancellation and revoked authority still take effect. **Steps to reproduce** Start a native session with `timeoutSec: 0` and run a tool beyond 15 minutes. Before this fix, the runtime cancels the turn. A recovery startup that uses a goal command also exposes the protocol-version mismatch. ## What Changed - Honor the agent turn timeout. Zero disables the timer. Long explicit durations use timer chunks to avoid Node timer overflow. - Default native runner lifetime to unlimited. Keep bounded startup, reconnect and control-operation deadlines. - Renew leases over the authenticated connection. Persist renewal before the reply. Validate identity, epoch and expiry. Handle duplicate requests and a lost reply on reconnect. - Freeze renewal during warm ownership transitions and terminal handling. - Validate persisted goal startup commands with protocol v2. - Add duration, renewal, ownership, recovery and real-process regression tests. Update runner protocol and recovery docs. - Add no UI components or controls. Renewal requires no model output or user action. ## Verification Current head: `348e369c35c5da8bb8be378f4b35dcf6f40882e7`. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34649767113). - All 32 checks pass on this head. The two Storybook checks are skipped as expected. CI includes full build, typecheck, runner verification, browser suites, server tests and the canary package dry run. - Greptile reports 5/5 on this head. All review threads are resolved, and the security scan passes. - Passed `pnpm -r typecheck` and `pnpm build` locally. - Passed 219 native-runtime and controller tests, including fake-clock tests for three weeks of renewal and 30-day explicit timeouts. Six denial tests confirm that renewal cannot extend expired, revoked or mismatched authority. - Passed 337 executor, cancellation and restart-recovery tests, plus 278 Rust runner-core library tests. - Passed a real runner with a silent fake Codex provider across its original lease expiry. Runner PID, provider PID, thread and active turn stayed unchanged. Warm-attach recovery tests also pass. - Passed all 83 plugin-worker tests and 159 of 161 workspace-runtime tests locally. The two remaining assertions passed with a canonical macOS temporary directory, as did the changed runtime fixture. The full affected server shard passes in CI. - An unchanged GitHub callback-ordering test failed once in CI, passed locally in isolation, and passed its one test-shard retry. The final CI summary is successful. - The full local `pnpm test:run` sweep was interrupted after dependency setup failures and load-related timeouts. Identified failing suites passed in isolated reruns after the dependency repair. The complete test matrix passed remotely in CI. ## Risks - Deploy the controller and runner together to enable renewal. Older peers keep their existing bounded lease behavior. - Unlimited runtime permits long resource use until completion, explicit cancellation, configured timeout or loss of valid authority. - Lease renewal changes authenticated protocol handling. Regression tests cover stale, revoked and mismatched authority, lost replies and warm handoff behavior. - Simulated multi-week tests and a real lease-boundary test do not constitute a weeks-long production soak. ## Model Used OpenAI GPT-6 through Codex, with repository inspection, code execution and TypeScript/Rust test tools. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2083bf6f9a |
feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents controlled access to external services. > - Experimental channels already map conversations to tasks and durable work queues. > - Email needs inbox ownership, recipient envelopes, delivery records, and explicit sends. > - This pull request adds AgentMail to that infrastructure and keeps the provider key in the server vault. > - Agents can receive and send email from local or sandbox execution while the board follows each conversation in its task. ## Linked Issues or Issue Description **Problem or motivation** Agents need dedicated email addresses. Incoming email should become assigned work. Internal task comments and progress must never become outgoing email by accident. **Proposed solution** Add experimental AgentMail connections, an inbox assignment wizard, durable email intake and publication, task email cards, and authenticated API, CLI, and native runtime actions. Agents use Paperclip credentials to request sends. Paperclip owns the provider key and enforces access and task authority. **Alternatives considered** A general mailbox MCP connector does not provide durable task binding or publication boundaries. A separate mailbox application duplicates task collaboration. The board instead directs the agent through the normal task conversation. **Roadmap alignment** This extends the existing experimental connections and task infrastructure. Product scope and interaction design were reviewed with the maintainer. Related connection authority work: #11831 and #11818. The duplicate search found no competing task-based AgentMail integration. ## What Changed - Add AgentMail catalog data, shared contracts, company-scoped email records, and an additive migration. - Add vaulted setup, inbox assignment, access grants, trust guidance, and provider-side allowlist guidance. - Support WebSocket and signed-webhook intake through a shared durable pipeline, deduplication, catch-up, and task wakeups. - Queue explicit new conversations and replies with immutable send intents, idempotency, delivery state, and uncertain-send resolution. - Show inbound and outbound email cards in normal task conversations. Keep internal messages internal. - Add task-scoped CLI actions and the sandbox callback routes required for Daytona execution. - Provide a dedicated AgentMail skill automatically only to agents with active authorized inbox assignments. Keep email instructions out of the universal Paperclip skill. - Advertise connector-owned `agentmail_inboxes`, `agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery` tools only in eligible native sessions. Recheck live authority on execution. - Isolate Codex CLI connector skills by agent and skill revision. Deliver the assigned skill in the run prompt for adapters that use shared skill directories, including resumed turns. Keep automatic skills out of manual persistent sync. Show them as read-only and document the pattern in the connector playbook. - Fix AgentMail health checks that entered local-stdio validation and optional missing Codex credential cleanup in sandboxes. - Add API, pipeline, authorization, sandbox, browser, and Storybook coverage. ## Verification - Live AgentMail testing covered WebSocket intake, signed webhooks, restart catch-up, and a full receive → task → Daytona Codex CLI → explicit reply → Delivered round trip. The reply was verified in the other inbox. The normal task composer also initiated an outgoing email child task. - The connector-skill change was verified in the browser: AgentMail appears once as an automatic, read-only skill with its assigned address. Disabling experimental chat connections removes it; re-enabling restores it. A regression test covers assignment data arriving after library data. - Connector regression coverage passed 178 runtime utility, email integration, skill-route, and heartbeat tests. All 17 Codex execution tests passed, including per-agent skill isolation, model identity, revision changes, removal, and prompt delivery without shared skill files. - After rebasing onto master, all 44 focused email, heartbeat, and native-authority tests passed. All 313 native-session executor tests passed. The UI regression suite passed all 3 tests. These test sets overlap earlier focused runs. - Full workspace typecheck and build passed after the rebase. Token gates passed. Earlier focused Playwright task/setup coverage and the Storybook build also passed. - Native connector tool execution uses deterministic integration tests. Live Daytona qualification used the Codex CLI adapter; the new shared-home prompt fallback has deterministic coverage. - The full repository suite is run by CI. The earlier unsharded local full-suite attempt was stopped after the equivalent CI suites passed and is not reported as a completed local run. Greptile reviewed `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved threads. All server, workspace, serialized server, and browser suites passed in CI. The build job hit a five-second timeout in a runner transport test; both variants and the full 80-test file passed locally with unchanged timeouts. The build passed on retry on the same commit without code or timeout changes. All required CI gates, including the final `ci / verify` and `ci / e2e` summaries, are green on `7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`. ## Risks - Email from external senders can start normal agent work. Setup recommends a low-trust agent and AgentMail sender controls. Sender addresses never grant board membership. - Provider timeouts can leave uncertain sends. Retries retain their idempotency key; expired windows require reconciliation or operator resolution. - Connector skills and native tools are assignment-dependent and require current access. Revocation denies retained calls; assignment changes select a new runtime context. - Activation remains behind the experimental-channel setting. The native runner path has deterministic coverage; live Daytona qualification used the Codex CLI adapter. - Schema changes are additive. Inbox ownership is unique across companies. Disconnect preserves provider inboxes and task history. ## Model Used OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a12bbd1824 |
fix: stop completion reviews caused by policy upgrades (#13266)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runtime records completion assessments and task decisions. > - An application update can change the assessment policy version. > - The previous code treated that version change as a reason for human review. > - An upgrade alone does not give the user a new decision to make. > - This change keeps existing decisions and withdraws obsolete upgrade review cards. ## Linked Issues or Issue Description **What happened?** A rules version change moved unfinished tasks into review and created a card that said, "Review the superseding native policy assessment." The task did not need new work or a human decision. **Expected behavior** New runs use the current rules. An upgrade leaves existing task decisions alone. The saved policy version remains available in the audit history. **Steps to reproduce** 1. Save a native run assessment and a task decision. 2. Change the native status policy version. 3. Run finalization reconciliation without new task evidence. 4. The previous code created a review card. The corrected code keeps the saved assessment and decision. Related context: #13038 changed the native policy version. Searches found no duplicate fix for upgrade-only completion reviews. ## What Changed - Remove policy-version mismatch as a reconciliation trigger. - Withdraw pending cards only when their source decision, effect ledger, creator, key, and prompt match the old upgrade-only review. - Restore the previous status only while that decision and status version remain current and no other review gate is pending. - Preserve answered cards, real review requests, later task changes, and historical assessments and decisions. - Record cleanup activity and retire any corresponding chat review actions. - Isolate cleanup failures so one old card cannot block other cleanup or normal finalization. - Update the status conformance fixture, regression tests, and architecture documentation. ## Verification - Passed: `pnpm exec vitest run server/src/__tests__/native-status-arbiter-corpus.test.ts` (23 tests, including the 53-fixture status corpus). - Passed: `pnpm -r typecheck`. - Passed: `git diff --check`. - Passed: `pnpm build`. - Passed: fresh Greptile review at 5/5 on `d24e5ecaf`, with no open review threads. - Passed: the 23 focused tests in the isolated full-suite environment. - Local `pnpm test:run`: the server group finished with 10,638 passed and one missing-fixture failure. Built the required `fake-codex-app-server` fixture and reran the entire affected native session-resume suite: 37/37 passed. The initial full command exited on that server-group failure, so remaining groups are covered by CI. - Passed: the workspace-runtime-exposure suite (25 tests, 3 platform skips). - Passed: all CI gates, including every test shard, browser tests, typecheck, and build. The server shard passed on one retry after an unrelated host-port conflict in the unchanged workspace-exposure tests. - Cleanup tests cover repeat runs, real reviews, answered cards, later statuses, newer decision identity, status changes back to review, other pending requests, run scope, cleanup failures, retry, and live publication failures. ## Risks - Cleanup changes stored task state. It checks the exact obsolete decision and status version under database locks before restoring status. - Pending approvals, interactions, and execution stages prevent restoration out of review. - The cleanup handles at most 100 matching cards per reconciliation pass. It does not rewrite old decisions or accept an agent's completion claim. - New evidence and explicit task changes still use the existing reconciliation paths. This change does not reevaluate old work merely because Paperclip was updated. ## Model Used OpenAI Codex, GPT-6. The exact deployment identifier and context window are not exposed in this session. Used reasoning, repository inspection, code editing, shell execution, and automated tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad4f0b5867 |
Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent runtime settings can bind organization secrets to an adapter environment > - Paperclip redacts plain environment values when it returns a saved agent to the UI > - A saved-agent test sent the redacted `CODEX_HOME` value back to the server > - Codex ACP also received the API key without an ACP API-key authentication request > - This pull request restores saved environment values for tests and selects API-key authentication for Codex ACP runs > - The benefit is that Codex agents can test and run with an organization-scoped OpenAI API key ## Linked Issues or Issue Description **What happened?** Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`. Secret normalization rejected that placeholder. Remote Codex ACP runs received `OPENAI_API_KEY`, but session creation stopped with `Authentication required`. **Expected behavior** Paperclip must use the saved `CODEX_HOME` value when it tests an existing agent. Codex ACP must select API-key authentication when `OPENAI_API_KEY` is available. **Steps to reproduce** 1. Create an organization-scoped secret named `OPENAI_API_KEY`. 2. Give a Codex agent access to the secret. 3. Save the agent runtime settings. 4. Test the saved agent again. 5. Run the agent in a remote sandbox through ACP. **Paperclip version or commit** Reproduced on master before commit `68c17709d7c051a804a416263e2e08920f1dfcb1`. **Deployment mode** Self-hosted server with a remote sandbox environment. **Installation method** Built from source. **Agent adapter(s) involved** Codex. ## What Changed - Send the saved agent ID with adapter environment tests. - Restore redacted plain environment values from the saved agent before test-time secret resolution. - Select the Codex ACP `api-key` authentication method when `OPENAI_API_KEY` is present. - Add focused regression coverage for saved-agent tests and remote ACP launch configuration. ## Verification - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts` - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm --filter @paperclipai/ui exec vitest run src/lib/test-agent-setup.test.ts` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `git diff --check` ## Risks - Low risk. The test route reads saved configuration only when the request supplies a compatible agent ID and the caller can update that agent. - The Codex ACP change applies only when `OPENAI_API_KEY` exists and no explicit `DEFAULT_AUTH_REQUEST` exists. - There are no schema migrations or telemetry changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex with `gpt-5`. The context-window size is not exposed in this runtime. The model used reasoning, repository search, file editing, command execution, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3bafac12f7 |
refactor: remove automatic productivity reviews (#13263)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its recovery loop keeps assigned work moving after execution failures. > - Productivity review used run counts, comment counts, and elapsed time to create management tasks. > - Infrastructure failures could satisfy those rules and create more tasks without evidence that the source work needed management review. > - This pull request removes that detector and its continuation holds. > - Bounded recovery, budgets, explicit blockers, and normal review stages remain in place. > - Existing task records stay readable and unchanged. ## Linked Issues or Issue Description Refs #5897. That request describes unwanted automatic productivity reviews and asks to preserve existing tasks. This change retires the feature instead of adding another configuration switch. Related prior approaches: Refs #9191, Refs #12489. Those changes excluded infrastructure failures or bounded review creation. This removal replaces the detector rather than tuning its thresholds. ## What Changed - Delete the scheduled detector, automatic task creation, evidence refresh, and productivity continuation holds. - Remove computed productivity fields, special attention items, badges, and Storybook fixtures. - Retain historical origin values, decision compatibility, and recovery recursion exclusions. Add no migration and change no existing task data. - Update the execution contract. Replace feature tests with regressions for legacy task reads, ordinary attention, and bounded continuation in the presence of an old review. ## Verification - Targeted attention, issue-route, startup, and UI tests: 4 files and 101 tests passed. - Updated issue-route and UI tests: 2 files and 61 tests passed. - Bounded continuation regression: 2 cases passed, including a legacy review plus pre-dispatch cancellation churn. - `pnpm check:token-gates`: all four gates passed. - `git diff --check`: passed. - `pnpm build-storybook`: passed. - Greptile: 5/5 on `a5a612eea`, with no actionable findings. - Scheduler and historical recovery regressions: 2 files and 28 tests passed. - Repository `pnpm -r typecheck` and `pnpm build`: passed. - The complete `pnpm test:run` suite passed across the CI server, serialized-server, and workspace shards on `a5a612eea`. Stopped the duplicate local monolithic run after the full CI suite passed; no completed local full-suite result is claimed. The targeted local suites above passed. - CI serialized shard 5 initially hit a 10-second timeout in the first interaction-route test. The complete file passed locally (78 tests), then the single CI rerun passed. - All CI gates are green, including the build and end-to-end suites. - A local merge check against current `master` (`ce09ea40b`) completed without conflicts. ## Risks - API responses no longer include the computed `productivityReview` field. Consumers must stop using it. - The scheduler no longer creates management work from elapsed time, run counts, or missing comments. This is the intended behavior change. - Existing review tasks and explicit dependencies remain in place. Historical origins still prevent recursive recovery treatment. No task cleanup or data migration occurs. - The native review handoff repair is separate from this removal. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
663c44cb2b |
fix: continue conversations after confirmed remote runner stop (#13254)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can stop a run and send another message on the same task.
> - Remote runners need evidence from their sandbox provider that
execution stopped.
> - Local process checks cannot prove that a remote process exited.
> - This pull request records provider stop receipts and uses them for
conversation admission.
> - New user messages can proceed after confirmed cleanup without
repeating interrupted actions.
## Linked Issues or Issue Description
Refs #13237 and #13239. Related: #13163 covers app-restart recovery;
this change covers an explicit stop followed by a new user message.
**What happened?**
A stopped remote Claude ACP task kept its execution hold after Daytona
cleanup succeeded. Native runners also rejected remote process
identities and retained stale session cleanup gates. A message sent
during cleanup could stay deferred after the sandbox stopped.
**Expected behavior**
After the provider confirms that the old execution stopped, a new user
message starts a fresh turn. Pending user messages must not need another
message to trigger admission. Prior action outcomes remain recorded.
**Steps to reproduce**
1. Start a long-running task in Daytona with a legacy Claude ACP or
native ACP runner.
2. Cancel the run while its tool is active.
3. Send a new message immediately, or after cleanup completes.
4. Observe the execution hold despite the old sandbox having stopped.
**Paperclip version or commit**
Reproduced on master at
|
||
|
|
b1efd65edc |
fix: continue interrupted task conversations with bounded retries (#13237)
## Thinking Path > - Paperclip manages AI agents and their tasks. > - A task can outlive a provider process or a server restart. > - Legacy recovery treated unknown tool outcomes as a permanent execution hold. > - That hold could also reject a later user message. > - A conversation turn can use prior history without replaying prior tool calls. > - This pull request lets supported conversation adapters continue within the existing retry budget. > - Users can send a new message after automatic attempts stop. ## Linked Issues or Issue Description **What happened?** A server restart could interrupt a local ACP run and leave its task behind a permanent recovery hold. A later user message could be cancelled before the provider answered. The immediate recovery path could also create a successor outside the durable failure counter. **Expected behavior** Continue with a bounded new conversation turn. Preserve a compatible provider session or use full task context when it is unavailable. Do not replay recorded tools. When automatic attempts stop, allow a new user request through the normal execution gates. **Steps to reproduce** 1. Start a task with a local conversation adapter. 2. Restart the server while the provider is working. 3. Let the previous run become interrupted. 4. Send a follow-up message and observe the recovery hold on the old behavior. Related work: Refs #13075 for durable task recovery. Refs #12946 for retry-limit and checkout-lock handling. This change routes conversation recovery through the existing bounded scheduler. ## What Changed - Mark supported local conversation failures for continuation. Keep native-runner and non-conversation recovery rules. - Carry an interruption notice into the next turn. Retain stopped ACP session history even when a write outcome is unknown. - Clear unavailable ACP sessions so the next bounded attempt can use full task context. - Route immediate failure recovery through the same durable scheduler as process-loss recovery. Release only the predecessor checkout when its retry takes ownership. - Retire obsolete conversation holds using immutable run evidence, in bounded batches with an activity record. Preserve outcome evidence and do not wake historical tasks. - Block actual admission and Resume while a predecessor process or environment lease is still active. Keep the original interruption notice after a rejected wake. Preserve the upstream blocked-wake waiting contract: bounded retry planning can happen during cleanup, while deferred messages and execution remain gated. - Add subprocess and database regression tests. Update the execution contract. - Add the current thread-status field to the native recovery provider fixture so its damaged-journal test reaches the intended boundary. Tolerate an already-exited fixture process during test cleanup while still asserting both processes terminate. ## Verification - Workspace typecheck passed: `pnpm -r typecheck`. - Build passed: `pnpm build`. - Module boundaries passed: `pnpm check:module-boundaries`. - Focused tests passed: 293 recovery/session/dispatch tests, 66 retry and response-gate tests, and 37 native-session tests. Some suites overlap. - Tests cover interrupted writes, missing sessions, concurrent retries, restart persistence, pending questions and approvals, execution gates, and historical holds. - Built the Rust test executables with `pnpm --filter @paperclipai/paperclip-runner build:rust` for native-runner verification. - Full Vitest coverage verified locally using the repository’s general and serialized shards, with focused reruns for failures and files not reached after a shard stopped. The ownership-gate regression is fixed and the complete affected server shard passes (1,390 tests). Local parallel runs also hit temporary-directory, resource, and timing failures; those suites pass with canonical temporary paths and sequential reruns. No test timeouts were increased. - Final merged-branch regression run: 577 tests pass across process recovery, retry scheduling, liveness, durable chat, wake-queue application/adapter, dispatch, continuation, native sessions, and task chat. Earlier focused verification also passed 19 native control tests. Token gates and whitespace validation pass. - Browser verification passed all three ACP Stop/continue/pause scenarios, including a rerun after merging the upstream waiting behavior: `PAPERCLIP_E2E_PORT=3397 pnpm test:e2e tests/e2e/acp-stop-continuation.spec.ts`. The interrupted-write case verifies that follow-up completes without a repeated write. - Final-head [CI run 34625037394](https://github.com/paperclipai/paperclip/actions/runs/34625037394) passed on `06ac4bd9d150f8b209a96e5fd609c696958794a0`: all 31 reported checks are green, including server/workspace suites, all browser shards, native runner verification, build, typecheck, release dry run, and aggregate gates. The two conditional Storybook checks were skipped. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - A new model turn can choose to repeat an action. Paperclip does not replay recorded tool calls and does not certify unknown action outcomes. - Conversation adapters now stop after their retry budget instead of requiring action reconciliation. Explicit Stop, pause, dependency, approval, budget, and ownership gates remain in force. - No schema migration or dependency changes. Historical holds are folded without changing task status or waking work. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and test execution. The session does not expose a more specific model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1616046c24 |
fix(runner): prevent trace scans from delaying live events (#13228)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner sends events that keep task state and steering
controls current.
> - Debug trace correlation reread and parsed the full trace for each
pending event.
> - Repeated scans blocked event delivery and left task state minutes
behind the provider.
> - This pull request indexes appended trace records once and reuses
those correlations.
> - Operators get current run state while native trace evidence remains
available.
## Linked Issues or Issue Description
**What happened?**
Native runs appeared live after their provider turn ended. Steering was
unavailable while the board still showed an active run. A CPU profile
attributed 93% of sampled time to trace lookup and file reads.
**Expected behavior**
Debug trace correlation must not delay run events or task state by
minutes.
**Steps to reproduce**
1. Enable native provider trace capture for a Codex run.
2. Produce a long event stream with pending correlations.
3. Compare provider event times with persisted run event times.
**Paperclip version**
Reproduced on source build
|
||
|
|
87b3e5fc61 |
fix(adapter-utils): stage selected skills into the sandbox for a remote Claude ACP run (#13196)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - A user selects skills for an agent, and the host materializes those skills into a bundle the agent reads > - An agent can run in a remote sandbox, where the host must stage every file the agent needs > - On the Agent Client Protocol lane the host built that bundle and then named its host path in the prompt, but it never staged the bundle into the sandbox > - The agent therefore read a path that does not exist inside the sandbox, and the run failed on the missing skill file > - The command-line lane of the same adapter already stages a `skills` asset and reads the in-sandbox directory back from the staged runtime > - This pull request carries that proven pattern to the Agent Client Protocol lane, so a selected skill reaches the agent in a remote run ## Linked Issues or Issue Description No public issue exists for this change. The description follows. **What happened?** A remote run of the Claude adapter on the Agent Client Protocol lane could not read any selected skill. The host builds the skill bundle in its own state directory, then writes that host path into the prompt as `Skill root: <path>`. The remote seam of that lane staged one asset only, the configuration seed. It staged no skills asset, so no skill file crossed into the sandbox. The agent then tried to read the skill file at the host path, and the read failed with a missing-file error. **Expected behavior** A remote run receives the skills the user selected, and the prompt names the directory that holds those skills inside the sandbox. **Steps to reproduce** 1. Select one or more skills for an agent that uses the Claude adapter. 2. Start a run for that agent in a remote sandbox on the Agent Client Protocol lane. 3. Ask the agent to read the skill file at the path the prompt names. The file is not there. **Agent adapter(s) involved** The Claude local adapter, on its Agent Client Protocol lane. The shared engine in `packages/adapter-utils` carries the prompt rewrite. **Additional context** The command-line lane of the same adapter already stages a `skills` asset and remaps onto the staged directory. This change reuses that mechanism instead of adding a new transport. One other adapter shows the same host-path shape on its own Agent Client Protocol lane. That lane is tracked separately and this pull request does not change it. ## What Changed - Return the host skill bundle directory from the Claude skill runtime step, and carry it through the remote managed-home context to the staging seam. The value is null for a non-Claude agent, for a run that selects no skill, and for a run whose selected skills all fail to materialize. - Stage that bundle as a `skills` asset on the Claude Agent Client Protocol remote seam, and only when the run selected a skill. - **Stage that asset with `followSymlinks: false`.** The bundle holds an owned copy of each selected skill, and the copy step never copies a symbolic link at the root or at any depth. So the bundle contains no symbolic link, and staging has none to follow. Refusing to follow one also stops a link planted in the bundle directory after the copy from pulling an unrelated host file into the sandbox. A regression test walks the real adapter sources and pins the reviewed `followSymlinks` value at every skills staging site, so a new or changed site fails the test. - **Drop a skill whose staged copy has no usable `SKILL.md`** from the prompt, the skill identity, the command notes, and the bundle, and log which skill was dropped and why. Without this, a skill whose copy failed, or whose `SKILL.md` is a symbolic link the copy step skips, stayed advertised in the prompt while its file was absent — the same missing-file symptom this change exists to fix. - Rewrite the `Skill root:` prompt line, the skill identity, and the command notes onto the in-sandbox directory. The rewrite runs in the engine, after the workspace placement returns the staged runtime. A compatible session resume reuses the cached staged runtime, so the rewrite runs on that path too. - Keep the session fingerprint on the host-independent skill identity. A change to the selected skill set still invalidates a warm session, and the volatile sandbox path stays out of the hash. - A local run, and a run with no selected skill, keep their current behaviour. ## Verification - `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` — 28 of 28 pass. - `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` — the new engine tests and the staging-site tests pass. - `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`, and the same check on the adapter package — both exit 0. - The end-to-end test drives the lane against a local sandbox stand-in. It reads the skill root out of the prompt the runtime received, and then opens the skill file at that path. That is the reported symptom, proved closed. - The new tests carry a sensitivity control. Restoring only the production files to their previous content fails 7 of the 9 new tests. The other 2 do not depend on production code: one is a parser unit test for the source scanner. ## Risks Low risk, and the change is a two-way door. A revert restores the previous behaviour exactly. - **Scope.** The change touches one adapter lane. It does not change the local lane, and it does not change any other adapter. No existing staging site changes its `followSymlinks` value. - **The staged bundle and the workspace.** The staged skills land under the runtime directory inside the workspace. The workspace restore excludes that whole runtime directory, so the staged skills never return to the host worktree. A test proves the exclusion end to end. - **Session reuse.** The rewritten path never enters the session fingerprint, so it cannot invalidate a warm session, and a compatible resume applies the same staged path. - **Direction of data.** Files move from the host into the sandbox only. The change adds no path that writes sandbox content onto the host. - **A dropped skill.** A skill with no usable `SKILL.md` is now absent from the prompt instead of named but unreadable. The run logs the skill and the reason, so the cause is visible. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking and tool use, through Paperclip agents. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/adapter-claude-local src/server/acp.test.ts` passes — 28 tests. - [x] `pnpm exec vitest run --project @paperclipai/adapter-utils src/acpx-engine/execute.test.ts src/skills-staging-follow-symlinks.test.ts` passes. - [x] `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit` exits 0. - [x] All continuous-integration gates are green. - [x] The automated review reports no open finding against the current head. ## Required Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used with version and capability details - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant bug report template - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run the targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect this change - [x] I have considered and documented the risks above - [x] All continuous-integration gates are green - [x] The automated review score is 5/5 with no open current-head findings - [x] I have addressed every reviewer comment that applies to the current head **Note on the branch history.** This branch first carried a different change: a filename-based admission filter that refused to stage files such as `.env` from a skill directory, together with a switch from symbolic-link bundles to copied bundles. That approach was rejected and **reverted** on this branch. It does not match the documented trust boundary, because the host already delivers credentials into the sandbox on purpose, and replacing the symbolic-link bundles broke live editing of a skill. The revert is in this branch's history. The file that work changed, `packages/adapter-utils/src/server-utils.ts`, is byte-for-byte identical to `master` here and is not part of this diff. Earlier review findings that name that file target the reverted code. All of them are resolved, and the automated review passes on the current head. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a05b828bcd |
Reduce run polling and workspace inspection amplification (#13174)
## Thinking Path > - Paperclip manages agent work and shows run progress to operators. > - Run lists, live events, transcripts, and workspace details must remain responsive as usage grows. > - Run-list redaction rereads the full context for every run. Hidden tabs can still trigger requests through live events and manual timers. > - Workspace detail reads repeat Git inspection even when concurrent callers request the same state. > - This pull request batches registry reads, pauses hidden-tab refreshes, and caches Git inspection for display. > - Cleanup keeps fresh Git checks, and redaction keeps company and run boundaries. ## Linked Issues **What happened?** Run-list responses perform one extra database read per run and parse full context JSON to obtain small secret registries. Hidden tabs continue transcript reads and event-triggered refetches. Workspace detail requests repeat Git scans. **Expected behavior** A run list reads registries once. Hidden tabs stop recurring run reads and reconcile when visible. Concurrent workspace detail reads share a short-lived Git result. **Steps to reproduce** 1. Open run lists and task transcripts in several tabs while agents run. 2. Hide some tabs and observe transcript and event-triggered requests. 3. Request a 200-run list and count redaction database queries. 4. Request the same workspace detail concurrently and count Git inspections. Related: #5255 adjusts polling cadence. This change addresses hidden-tab lifecycle, batched registry reads, and workspace inspection reuse. No duplicate with this scope was found. ## What Changed - Batch heartbeat and live-run redaction into one company-scoped registry query. Select only registry JSON for run and issue redaction. - Resolve duplicate secret values once per request. Preserve each run's registry and remove registry material from responses. - Suspend company event sockets and transcript reads while hidden. Refresh active queries and resume transcript offsets on return. - Prevent queued event invalidations and developer health polling from fetching in hidden tabs. Gate legacy run-log readers in both UI variants. - Exclude legacy plugin placeholder connections from remote health probes. Select only due connection IDs in SQL before the sweep limit. Preserve existing plugin records. - Cache concurrent Git display inspections for five seconds, with at most 256 entries. Leave close-readiness and cleanup checks uncached. - Add regression coverage and document the performance behavior. - Stabilize the existing Rust descendant-lineage fixture: allow a bounded 30 seconds for 300 durable notifications under concurrent test load, retaining every correctness assertion and adding timeout diagnostics. ## Verification - Regression coverage verifies one registry query for 200 runs, per-run isolation, request-local secret resolution, decryption failures, Git cache expiry/bounds, hidden-tab pause, and visibility recovery. - Real PostgreSQL redaction/run-route suites passed all 57 tests; workspace-service coverage passed. The health-sweep regression verifies plugin placeholders and chat connections remain untouched and do not consume the sweep limit. - Both legacy transcript viewers retain history and resume their byte offset after visibility changes. The related visibility/progress/chunk suites passed all 29 tests. Other focused UI suites and token gates passed. - Full `pnpm -r typecheck` and `pnpm build` passed. Affected-package typechecks/builds passed after review fixes. The concurrent Rust provider suite passed 84 tests (two ignored), and Rust formatting passed. - Full local `pnpm test:run` stopped after the general-server group: 10,538 passed, 65 skipped, four failed. Fresh chat-delivery and health-sweep reruns passed; building the debug runner fixture cleared the native-event test. One unchanged native-session recovery assertion still fails locally with a semantic-digest error instead of the expected settled-session message. The full local command is therefore not green. CI runs the later groups separately and skips the two native-session tests requiring a prebuilt runner binary (confirmed in its 37-test native-session suite). - All CI gates pass on final head `ee610e737`: typechecking, general and serialized tests, browser tests, runner verification, build, and canary dry run. One server shard passed on its single retry after exposure fixtures encountered port 42001 where they assumed 42000; that suite also passed locally (25 passed, three platform-specific skips). - Greptile reviewed the final head at 5/5 with no actionable findings. ## Risks - Workspace delivery display can lag local Git changes by five seconds. Destructive operations still inspect current state. - Hidden tabs do not receive company live-event notifications until visible. Active queries refresh on return. - This change preserves legacy plugin records and does not repair instance-specific workspace rows. There is no database migration. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository tools, code execution, and browser inspection. The exact model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted regressions; full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d1ba17eeca |
fix(adapter-utils): fail fast when the sandbox control channel is lost mid-turn (#13158)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Adapter utilities run agent turns and report their results to the control plane > - A lost sandbox control channel can leave an agent turn without a result > - The host then waits for the full adapter timeout instead of reporting the loss > - This pull request adds a push loss signal and a bounded host wait > - The benefit is a prompt failure terminal when the agent stops answering ## Linked Issues or Issue Description **What happened?** A sandbox control channel loss during an Agent Client Protocol turn left the host waiting for the four-hour adapter execution timeout. **Expected behavior** The host should detect the terminal channel loss, stop the turn, and report a safe failure without waiting for the agent. **Steps to reproduce** 1. Start an Agent Client Protocol turn through a sandbox adapter. 2. Close the duplex control channel while the turn remains active. 3. Observe the host response before the adapter timeout expires. **Paperclip version or commit** Test the pull request commit set at `10b6bbc5525a79fd575298607dd5a25ae448fc8a`. **Deployment mode** The change applies to sandbox-backed adapter execution. ## What Changed - Add `onLoss(listener)` to the duplex bridge handle. - Register the loss listener at turn start and read losses latched before turn start. - Cancel the turn on loss and arm a 30-second host deadline. - Close the stream locally when the deadline wins and create a host terminal. - Derive the public error from the closed `DuplexLossReason` enum. - Add tests for loss order, cancellation, timeout, and safe error output. ## Verification - Run `pnpm --filter @paperclipai/adapter-utils exec tsc --noEmit`. - Run `pnpm --filter @paperclipai/adapter-utils exec vitest run src/acpx-engine/execute.test.ts -t "run-disposition seam"`. - Confirm that the full pull request workflow passes. ## Risks The new deadline changes a lost-channel path from a long wait to a host-built failure after 30 seconds. Orderly completion keeps its existing behavior. The deadline race against a pending `turn.result` has no direct test. ## Model Used OpenAI Codex, GPT-5, with tool use and code execution. The runtime does not expose a more specific deployment version or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c1b55537ba |
fix(paperclip-runner): bump claude-agent-acp pin to 0.73.0 (#13162)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter can run agent turns through an ACP (Agent Client Protocol) server, `claude-agent-acp`, instead of the plain CLI > - Two separate packages each pin their own copy of that dependency: `packages/adapters/claude-local` (the server-side adapter) and `packages/paperclip-runner` (which builds the provider pack baked into every managed sandbox image) > - `claude-local` moved to `^0.73.0` in #12730, but `paperclip-runner` was never bumped past `0.70.0` — nothing keeps the two in sync when only one changes > - That split means a sandbox image built from `paperclip-runner`'s provider pack ships a `claude-agent-acp` the server-side adapter was never actually compatible with > - This pull request bumps `paperclip-runner`'s pin to `0.73.0`, the only version that satisfies both packages' declared ranges at once, and fixes the matching hardcoded version assertion in `docker/daytona-runner/Dockerfile` > - The benefit is one consistent, compatible `claude-agent-acp` version across both the server host and every sandbox image built from this source, instead of a silent split that only surfaces as a runtime failure ## Linked Issues or Issue Description No public issue exists for this specific split; opening directly per CONTRIBUTING.md path B, following the bug report template fields. **What happened?** `packages/paperclip-runner/package.json` pins `@agentclientprotocol/claude-agent-acp` at an exact `0.70.0`. `packages/adapters/claude-local/package.json` requires `^0.73.0` (added in #12730, 2026-09-02). Nobody re-synced `paperclip-runner`'s pin after that change — the two packages' dependency graphs are independent, so a bump in one doesn't propagate to the other. `paperclip-runner`'s copy is what the fleet sandbox image's provider pack actually ships, so every managed sandbox built from current source carries a `claude-agent-acp` version the server-side adapter's own declared compatibility range excludes. **Expected behavior** The two packages' `claude-agent-acp` pins should stay within a mutually compatible range, so a sandbox image built from this source always ships a version the server-side adapter actually supports. **Steps to reproduce** 1. Check `packages/adapters/claude-local/package.json`'s `@agentclientprotocol/claude-agent-acp` range (`^0.73.0`). 2. Check `packages/paperclip-runner/package.json`'s pin for the same package (`0.70.0` before this PR). 3. Note that `^0.73.0` on a `0.x` version only admits patch releases (`>=0.73.0 <0.74.0` per semver caret rules), so `0.70.0` falls outside it. **Paperclip version or commit** `master` as of this PR (paperclip-runner still at `0.70.0` prior to this change; claude-local's `^0.73.0` requirement landed in #12730). **Deployment mode** Any deployment that runs `claude_local` agents through the ACP engine against a sandbox image built from `packages/paperclip-runner`'s provider pack (managed cloud sandboxes in particular). Related PRs for context (not duplicates — none of these touch `paperclip-runner`'s pin): - #12730 — introduced the `^0.73.0` requirement in `claude-local` - #11873 — the last time `paperclip-runner`'s pin moved (`0.69.0` → `0.70.0`) - #13105 — separately made an unavailable ACP engine a hard failure instead of a silent CLI fallback, which is what turned this version split into a visible, run-blocking error rather than a quiet downgrade ## What Changed - Bump `@agentclientprotocol/claude-agent-acp` from `0.70.0` to `0.73.0` (exact pin, matching this package's existing pin style for its other agent-CLI dependencies) in `packages/paperclip-runner/package.json`. - Update the corresponding hardcoded version assertion (`test "$(claude-agent-acp --version)" = "0.70.0"`) in `docker/daytona-runner/Dockerfile` to `0.73.0`, so its own build-time check stays accurate instead of failing on the next build for an unrelated reason. - `pnpm-lock.yaml` is intentionally **not** included — `pr-trusted.yml`'s `Validate dependency resolution and regenerate stale lockfile` step already regenerates it for the merge tree and hands it to downstream `--frozen-lockfile` jobs as an artifact, so a manual lockfile commit here would just be stale the moment CI runs. ## Verification - `0.73.0` is a real published version on npm (confirmed via `npm view @agentclientprotocol/claude-agent-acp versions`), and it's the *only* version satisfying claude-local's `^0.73.0` range, so this isn't a guess at compatibility — it's the unique intersection of both packages' declared ranges. - `grep -rn "0\.70\.0" docker/ packages/paperclip-runner/package.json` after this change shows no remaining stale references to the old pin. - I did not run a full local install/test pass against a hand-updated lockfile, since regenerating one locally would conflict with leaving `pnpm-lock.yaml` untouched per the note above; CI's own lockfile-regeneration step is the intended verification path for a manifest-only dependency bump like this one. - Downstream/full verification (does a sandbox image actually built with this pin work end-to-end) is tracked separately in `paperclip-cloud` — an unrelated internal-only repo, so not linked here — where a sibling fix restores the ACP servers to the runtime `PATH` in the fleet sandbox image itself; both fixes are needed together for a working sandbox, but this PR is scoped to the version pin alone. ## Risks - Low risk: single-line dependency version bump plus a matching test-assertion update, no code changes. `0.73.0` is a patch release within claude-local's own already-declared-safe range, so there's no reason to expect it changes behavior tenants depend on. - The main risk is unknown breaking changes between `claude-agent-acp` 0.70.0 and 0.73.0 that aren't caught by the version-string assertion alone (that check only confirms the binary reports the right version, not that its behavior is unchanged). I have not audited that package's own changelog between those versions. - `docker/daytona-runner/Dockerfile` is a parallel/reference image (per its own header comment, meant to stay aligned with the private `paperclip-cloud/fleet-sandbox-image/Dockerfile`, which is out of scope here) — this PR does not touch that other Dockerfile. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), via Claude Code, with tool use (file edits, shell/git, `gh` CLI, `npm view` for version verification). No extended-thinking mode. Standard Claude Code context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — see Verification: a manifest-only bump with the lockfile intentionally left to CI's own regeneration step; no local test run applicable - [x] I have added or updated tests where applicable — version-pin bump only, no new behavior to test - [x] I have updated relevant documentation to reflect my changes — none applicable - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — pending CI run on this PR - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — pending review - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
889947c238 |
feat: add experimental native chat connectors (#13038)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also ask agents for work in their existing chat tools. > - Each external conversation needs one task and a current authorized source. > - Retries, Stop, and provider failures must not duplicate work or expose private data. > - The first chat PR establishes the opt-in provider and data contracts. > - This PR adds experimental channel integration and its durable control plane. > - Users can request work from connected channels and inspect delivery in Paperclip. ## Linked Issues or Issue Description Refs #13100 and #13092. This is the second of exactly two chat PRs. Foundation #13100 is merged and changed 143 files. Runner prerequisite #13092 is also merged. This PR changes 400 files against master, below the 500-file review limit. It contains no wireframe images or HTML galleries. ## What Changed - Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat connections. Keep chat disabled unless the operator enables experimental chat connectors. Preserve the production GitHub tool connection and its normal setup path. - Bind each provider bot identity to one immutable Paperclip agent. Bind each admitted external conversation to one task. Paperclip owns tasks, runs, permissions, and audit records. - Add durable admission, per-conversation queues, questions, task controls, progress, final replies, images, files, and delivery receipts. Board comments remain internal unless explicitly sent to the channel. - Check current identity, provider reach, resource access, credentials, runtime generation, and exact source before provider effects. Keep private responses private. Never send raw reasoning, private logs, credentials, or tool arguments. - Hold uncertain sends for explicit audited resolution. Make Board Send-to-channel atomic and idempotent. Keep reconnect and setup credentials in Paperclip secret storage. - Preserve current native-runner authority across retries, lost acknowledgements, and recovery. Keep immutable input and completion contracts separate from newer user input. Receipt reconciliation cannot launch a provider. - Reconcile chat close/new ordering and provider-effect lock order. Audit resource access changes in the same transaction. Submit only the selected resource from each UI toggle so stale pages cannot undo unrelated access changes. - Drain Codex stdout before certifying process exit. Bound the drain with the existing shutdown grace. Preserve observed terminal authority without treating an undrained process as successful or reusable. - Incorporate master `018ca5da` with its ACP Stop, mobile task layout, runner packaging, and official lock changes. Preserve dedicated chat-answer continuations in both directions when ordinary queued comments are adopted after Stop. - Fence late adapter readiness behind an earlier Stop for the same run. Preserve verified cleanup for registered adapters. Handle single Stop, agent pause, duplicate Stops, and failure release without creating a false cancellation receipt. - Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact failed-chat retry authorization and lineage, retired question-source suppression, and the block on generic recovery that would discard the admitted source. Fresh deferred input retains its separate promotion path. - Incorporate master `2a05b5ed3` and its queue-admission extraction, simplified transaction ports, and separate runner CI job. Preserve exact durable receipts, actor separation, and dedicated-answer isolation through the new module. A failed receipt insert rolls back the accompanying deferred-wake merge. ## Verification Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are resolved. This successor fixes two test-harness boundaries exposed by CI: per-case route-module preparation and actual durable-save completion before intentional runner termination. Production code and all existing test/turn deadlines are unchanged. [Exact-head Greptile review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594) is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable findings or open review threads. [Fresh exact-head CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341) passes **all 24 jobs**, including Build and both required aggregates. Normal exact-head guarded merge was attempted and rejected by the remaining branch approval policy: CODEOWNER review is required and no human approval is present. Normal **squash auto-merge is enabled** as of September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified; no approval bypass or self-approval was used. Earlier-head results below remain historical evidence, not qualification of this successor. - Final exact-head Linux evidence: 995/995 chat integration cases; 36/36 agent-skills routes; 35/35 runner live-session cases, including real process kill/resume; 1948 runner Vitest cases with three existing benchmark/platform guards; 870/870 API-authority cases; and 104 browser cases with four existing optional skips. Rust, conformance/replay, full repository build, typecheck, canary, all server/workspace shards, and both required aggregates pass with normal CI concurrency. Earlier failed attempts remain recorded below. - Latest test-only qualification: 141/141 route/permissions/authentication cases pass in separate cold forks, with plain server types and independent review clear. The real-runner suite passes 35/35, with plain runner types and independent review clear. A controlled premature-save acknowledgement fails as expected; matching ownership/effect/process evidence, rejected saves, real turn outcome, test abort, and pre-kill liveness are covered. No local reproduction of the original CI scheduling failure is claimed. The preceding [CI run](https://github.com/paperclipai/paperclip/actions/runs/34479680858) passes 21/24 jobs, including all 995 Linux chat cases and browser aggregate (104 passed, four existing optional skips); only Build, the skills serialized shard, and the required verification aggregate fail. Its exact-head Greptile review was 5/5. Both failed job logs are retained. - Final fixture qualification: all eight focused Discord cases and all 995 chat integration cases pass. The exact modal statement/PID is observed before taking the real connection lock; the test then proves its actual blocking relationship before mutation. Original SQL execution, provider behavior, negative assertions, and 1s/15s timeouts remain unchanged. Independent review is clear and test/production hashes remain frozen. The preceding [CI attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777) passed 22 jobs, including Build/runner, typecheck, canary, all other test shards, and browser aggregate (104 passed, four existing optional skips); the two fixture failures and failed verification aggregate remain recorded, not relabeled as a pass. - Current queue-module composition: 308/308 recovery/batching/queue/Stop tests; 995/995 full chat integration; 89/89 module tests, including real PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary tests; plain server and UI types. All four actual local process/ACP browser paths pass in 1.4 minutes. Fresh databases, no skips or retries, stable reviewed source hashes. The initial boundary failure is retained; its no-op service wrapper was removed without changing recovery context or weakening the check. An exploratory standalone test-directory typecheck fails because its new upstream transformation config is not a standalone typechecking project; standard CI/build does not invoke it, and no configuration was weakened to suppress those diagnostics. - The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed [all 24 CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958) and exact-head Greptile review at 5/5. Required CODEOWNER review prevented its normal merge before master advanced again. - Final extracted-module composition: 307/307 recovery, batching, queue and Stop-control tests; 995/995 full chat integration; 49/49 module tests including eight PostgreSQL adapter cases; and 19/19 issue-update tests. Plain server types pass. All four actual local process/ACP browser paths pass in 1.3 minutes. Fresh databases, no skips or retries in these cohorts, frozen source hashes, and independent review clear. - The preceding head `3e4e1c1c` passes [all PR CI jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820), including Build and required `ci / verify` and `ci / e2e`. Both the original Rust failure and the previously load-sensitive lineage fixture pass with unchanged Linux concurrency. Master advanced afterward and required this reconciliation. - Final master composition: 448/448 focused UI tests, 186/186 adapter tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI, server, shared, and adapter types pass. Token gates and diff checks pass. Independent server and UI reviews are clear. - Stop-registration regression: both real-service cases fail against exact `a95` source and pass with the fix. The full corrected recovery/control suite passes 265/265. Duplicate-owner and failed-Stop controls also pass. Plain server types pass. The readiness barrier prevents provider startup without adding an acknowledgment to an already terminal run. - Final qualification strengthens terminal-field equality and repeats both affected cases successfully on a fresh database. All four actual local process/ACP browser paths pass again in 1.3 minutes, without skips or retries. The final screenshot shows Cancelled, a paused subtree, retained input, and no error toast. - Two new actual-service regressions fail before the merge fix. They prove that queued-comment adoption could consume a dedicated chat answer or add unrelated input to that answer. The fixed four-case cohort passes, including ordinary upstream continuation and adapter Stop controls. Full recovery passes 257/257. All four actual local process/ACP Stop browser flows pass in 1.4 minutes, without skips or retries, on a fresh database. - The unchanged runner artifact was qualified with 171/171 transport tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11. Six controlled reader tests prove the exit/drain repair. Its local serial Rust workspace passed 546 top-level cases plus two invoked helpers; the later passing Linux CI supplies default-concurrency evidence. - Prior exact-source full chat integration passes 995/995. Settings regressions cover concurrent stale pages, 501 destinations, pending state, rejected updates, and explicit retry. These deterministic tests do not prove live provider behavior. - Retained failed attempts and their causes are in the [qualification log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md). The first merge adapter run timed out while macOS slept for 290 seconds. Its unchanged repeat passed with a temporary sleep guard. No assertion, deadline, or CI gate was weakened. Review commands include `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-process-recovery.test.ts src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh disposable databases. See the [browser runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md) for provider setup and separate live acceptance steps. ## Risks - This remains experimental. Deterministic tests and bounded live evidence do not establish every provider feature, tenant, permission layout, or media shape. Teams work-tenant qualification is still open. - Failed and uncertain provider effects remain visible and can require operator action. A transport receipt does not prove recipient visibility. - Native controller and runner artifacts must remain compatible. Preserve lease ownership, terminal authority, source binding, and quarantine during future changes. - Access and audit rows commit together, but activity notifications remain best-effort. This is not a new durable event outbox. - The PR operation does not deploy a live server, replace its runner, or change provider permissions. Remaining live qualification is documented in the [temporary handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md). ## Model Used OpenAI Codex assisted with implementation, tool execution, testing, and review. The work records `gpt-6-astra` assistance. The environment does not report a context-window size. No private reasoning traces are included. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |