mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 10:48:12 +02:00
f1a394bd30cb56fb9e479f98b9f50176fe921858
100
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9d2147171 |
fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud owns human identity and passes a verified identity to each tenant. > - The tenant can report no session while the Cloud session is still valid. > - The access gate and direct `/auth` route then show the instance password form. > - This pull request sends those users through the configured Cloud entry endpoint. > - Cloud can renew the tenant session or show its login page, then return to the original task. ## Linked Issues or Issue Description **What happened?** A Cloud tenant can display the self-hosted email/password form after an instance session check returns no session. This gives Cloud users the wrong login method. **Expected behavior** An active Cloud session renews tenant access automatically. A signed-out user signs in through Cloud. Staging and production use their own configured Cloud origins. Self-hosted instances keep their instance login form. **Steps to reproduce** 1. Open a Cloud tenant task or an `/auth?next=...` link. 2. Keep the Cloud session active but make the instance session check return 401. 3. Observe the instance password form instead of Cloud session recovery. **Deployment mode** Cloud-managed authenticated instances. No database or server API changes. Searched related authentication PRs. Native self-hosted OIDC support in #10411 is a separate feature; this change uses the existing Cloud entry contract. ## What Changed - Wait for deployment metadata before showing an instance login form. - Use the health response's Cloud origin and stack slug for session recovery. - Preserve the tenant path, query, and fragment. Reject external and recursive login return targets. - Limit automatic recovery per tab. Show a manual Cloud retry after failed recovery. Show service failures as errors. - Keep self-hosted login and local trusted access. Add focused tests, browser regressions, deployment documentation, and an unavailable-state design example. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - UI typecheck and `pnpm check:token-gates` passed after the final UI edits. - All UI tests passed: 639 files, 6,777 tests. - 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud links, and recovery coordination. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant UI and database with a simulated Cloud HTTP endpoint. They cover both Cloud origins, direct auth/task links, no password-form flash, preserved URLs, reload, and self-hosted login. - Hands-on browser test used the real Cloud gateway and a fresh tenant build with disposable local data. Active Cloud session plus a forced missing instance session returned to the task. An expired tenant cookie also renewed automatically and returned to the task. Removing both sessions reached the real Cloud email/social login UI. Persistent failure stopped at the retry screen; retry succeeded after removing the injected fault. The fixture used a loopback transport adapter and a simulated signed-out OIDC issuer. No production session or deployment was changed. - The default local browser startup hit the host's embedded PostgreSQL resource limit. The passing run used a separate disposable database on the test PostgreSQL process. - The full local `pnpm test:run` sweep was stopped after about 31 minutes once CI completed the full suite. It had reported 33 failures in the unchanged runner API unit/integration files; both files pass in isolation (1,749 + 28 tests). The local sweep did not reach the later workspace/serialized groups. CI completed all of those groups successfully. - Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at 5/5 with its only thread resolved. All CI gates are green on this commit, including all general/serialized server groups, workspace tests, Runner checks, typecheck, build, and all eight browser shards ([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)). ## Risks - Recovery depends on valid Cloud origin and stack metadata. Incomplete metadata shows an unavailable message instead of a password form. - Browsers with session storage disabled use the manual Cloud link, since automatic retries cannot be bounded across documents. - The external identity provider's email/social login was not completed in this local test. Existing Cloud authentication owns that flow. - No migration, credential format, membership rule, or production deployment changes. ## Model Used OpenAI GPT-6 through Codex. The exact served variant and context-window limit are not exposed in this session. Used reasoning, repository tools, code execution, and browser testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
270afd2fb8 |
feat(ui): show running commit in staging account menu (#14410)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The account menu shows the signed-in user's identity. > - Staging users need to know which server commit is running after a deploy. > - The health endpoint already returns that commit, but the menu does not show it. > - This pull request adds the short commit below the email on staging hosts. > - Users can open the menu to check a deploy without opening deployment tools. ## Linked Issues or Issue Description **What existing behavior does this improve?** The account menu on staging instances. **Current behavior** The menu shows the user's name and email. It does not show the running server commit. **Proposed behavior** On `*.staging.paperclip.app`, show `SHA 8751e2d` below the email. Use the current `/api/health` commit. Show the full SHA on hover and link to the commit on GitHub. Refresh the health query when the menu opens. Hide the label on other hosts and when commit metadata is unavailable. **Reason and benefit** A user can confirm which commit a staging instance runs after an automatic deploy. **Breaking changes** None. The server already returns the commit field. Refs #14060 for related account-menu work. This change adds deployment information only. ## What Changed - Add the existing health response commit field to the UI type. - Share the staging host check and a separate health query across both account-menu variants. A menu refresh failure leaves the access gate health state unchanged. - Show a short SHA below the email. Link to the full commit on GitHub and include the full SHA in its accessible name and hover title. - Document the staging label and cover staging hosts, other hosts, missing metadata, refresh on reopen, and request failure isolation. ## Verification - `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` — 24 tests passed. - `pnpm check:token-gates` — passed. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. The UI build and typecheck also passed again after review fixes. - Full test matrix — passed on the latest commit in [CI](https://github.com/paperclipai/paperclip/actions/runs/36447810742). The local `pnpm test:run` was stopped before completion after CI finished the same suites. The focused local tests, full local typecheck, and full local build passed. - Browser shard 6 passed on one rerun. Its first attempt lost part of the draft text in the existing attachment-receipt reload test. No code changed for the rerun. - Rendered the real account menu in a local browser fixture with a staging hostname condition and mocked health response. Confirmed the SHA fits below the email in the dark menu. - Greptile review — 5/5, all review threads resolved. - Manual check after deployment: open the menu on a staging host and compare the SHA with `/api/health`. Open the menu again after a deploy to refresh it. Confirm the label is absent on production and localhost. ## Risks - Low risk. Each menu opening on staging can make one additional health request. - The host check applies to `*.staging.paperclip.app`. Other staging domains will need an explicit update. - The label identifies the running server commit. It can briefly show cached data while the request completes. ## Model Used OpenAI Codex, GPT-6. The exact model variant and context window are not exposed in this session. Used code execution, repository inspection, and browser inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
14795136f5 |
fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3447609d22 |
fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.
## Linked Issues or Issue Description
**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.
**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.
**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.
Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.
## What Changed
- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.
## Verification
- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.
## Risks
- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.
## Model Used
OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
|
||
|
|
cea8dda472 |
test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users can delegate work through onboarding and Agent Chat. > - A completed task does not prove that its result reached the original conversation. > - Existing tests do not isolate completion after the source chat becomes idle. > - This pull request adds four explicit native-runner probes across Claude and Codex. > - The probes preserve the result and reply so we can separate delivery failures from inaccurate answers. ## Linked Issues or Issue Description Refs #13775. Refs #13813. These evals extend native-runner qualification. They measure completion updates before we choose a product change. ## What Changed - Add the opt-in `completion-updates` suite with two stories for each native provider. - Test completion in the existing onboarding task flow and after an Agent Chat handoff becomes idle. - Gate the chat worker on a brief inside its managed project workspace. Prove the source is idle before releasing the worker. - Check durable task completion, saved output, a subsequent source reply, and rendered access to the result. - Preserve replies, task state, screenshots, run events, and a separate semantic review rubric. - Add grader regression tests and update the documented eval contract. - Preserve the suites added on master and include four completion cases in the 306-cell catalog. Production behavior and prompts are unchanged. ## Verification - Passed all 565 eval support tests across 45 files after merging current master: `node node_modules/vitest/vitest.mjs run --config tests/runner-e2e/vitest.config.ts`. - Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p tests/runner-e2e/tsconfig.json`. - Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs tests/runner-e2e/launch.ts --list --suite completion-updates`. - Four-cell behavior campaign on source `ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`: https://github.com/paperclipai/paperclip/actions/runs/36072337485. - A screenshot-only follow-up waits for the restored source reply to render after result-link navigation. Its one-cell Claude onboarding verification passed on final head: https://github.com/paperclipai/paperclip/actions/runs/36075716141. Corrected report: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/. The original four-cell onboarding screenshots caught navigation loading; its saved reply evidence remains valid. The follow-up again found stale wording: "That work will run next" was posted 38 seconds after the child was Done. The four-cell campaign keeps its original source and measurements. - Published evidence: https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/. - Suite definition: `afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`, version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local execution, one attempt per cell. All four cleanup checks passed. Onboarding billing coverage is partial; reported zero cost must not be read as a free run. | Story | Automated delivery/access | Separate semantic review | | --- | --- | --- | | Codex onboarding | Pass | Pass: accurate completion reply with an accessible result | | Claude onboarding | Pass | Fail: reply says it will save the note once the task runs, after the note is already saved and the task is Done | | Codex idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | | Claude idle chat handoff | Fail | Worker completed and saved the note; no completion reply during the full observation window | Both chat cases positively recorded the source waiting and the worker at the brief gate before release. Both saved outputs include the brief-only start time. The opt-in campaign is red because it exposes current behavior. It is not a required merge gate. The PR does not fix that product behavior. Semantic review is a recorded human/agent assessment of retained evidence; it is not an automated prose-quality judge. - Second campaign: https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex chat reached the idle boundary and completed its task, then received no completion reply during the full window. Claude onboarding again returned a stale handoff answer. Claude chat exceeded the prior 110-second handoff setup budget; this revision raises that bounded setup window to 180 seconds. - Retained baseline: https://github.com/paperclipai/paperclip/actions/runs/36069427676. Onboarding passed delivery/access for both providers, but Claude gave a stale handoff answer. Chat cases stopped at fixture problems; they do not establish a completion-delivery failure. This revision fixes the workspace path and competing reference requirements. - On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR checks passed and two were skipped, including typecheck, tests, and build. Broad checks ran in CI, not locally. That head received Greptile 5/5 with no unresolved findings. The unchanged mobile repository-settings browser test passed on one targeted retry after a detached/disabled Save-button timeout. - Merged current master in `9b4491e1f` and resolved the catalog-count conflict. Eval support tests and eval TypeScript pass locally. All individual CI jobs passed on this merge commit, including build, typecheck, server tests, runner checks, and browser shards. The final aggregate check also passed: 54 checks passed and two were skipped. Greptile reviewed this exact commit at 5/5 with no unresolved findings. ## Risks - These explicit probes can expose current product failures. They do not change the default paid test selection. - Mechanical delivery and result access do not establish answer accuracy. The preserved reply still requires semantic review. - A fixture failure before the idle boundary or worker completion cannot establish a completion-update failure. - The handoff setup window lasts three minutes. The worker brief wait is bounded at four minutes. The observation window lasts two minutes after worker completion. It retains later replies without erasing earlier accessible delivery. ## Model Used OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository inspection, code execution, and GitHub tool use. The runtime does not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8751e2de46 |
fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The task board shows which recovery actions are active. > - A native run can stop while a person must repair its workspace. > - The board previously called that state “Recovery in progress.” > - The label implied that work would continue without operator action. > - This pull request derives the label from the recovery owner and live continuation. > - Operators can distinguish scheduled recovery from a repair that needs attention. ## Linked Issues or Issue Description **What happened?** A blocked task showed “Recovery in progress” after native finalization stopped and no automatic continuation remained. **Expected behavior** Show “Recovery needed” for an idle board repair or when no live recovery path exists. Show “Recovery in progress” while the recorded continuation can run, including an explicitly admitted export whose exact callback is executing even if the old recovery action remains board-owned. **Steps to reproduce** 1. Complete a native run whose workspace export cannot be recovered automatically. 2. Inspect the task recovery action and badge. 3. Compare the board-owned action with the old “Recovery in progress” label. Related work: #14314 serializes native workspace finalization and fences stale recovery outcomes. This change reports the recovery action that currently owns the task. ## What Changed - Show “Recovery needed” for board-owned active-run recovery unless the exact native export callback is positively verified as executing. - Require the native continuation run and a live or future continuation before showing progress. - Project native activity from the exact company, source issue, and run. Include active workspace export while the original heartbeat remains failed. - Require an executing callback before using a running export row as evidence. Preserve activity for long exports and clear it when the callback joins. - Give native resume its own card explanation. Preserve ordinary watchdog observation behavior. - Remove the redundant ownership sentence from all six recovery-card explanations that used it. - Document the labels and add regression cases for stopped, scheduled, and active recovery. ## Verification - Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests and token gates pass. No UI occurrence of the removed sentence remains. `pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54 successful checks, two optional Storybook checks skipped). Greptile is 5/5 with no open review threads. The duplicate local `pnpm test:run` was stopped after the full CI suite passed; it did not complete locally. - Original regressions: seven failures before the change, then 52 focused cases pass. - Review regressions: seven UI failures and nine database failures before the follow-up. All 75 database/API recovery tests, 167 UI tests, and 18 workspace lifecycle/finalizer tests pass. A further three RED cases cover explicit board retry activity; one RED case rejects orphaned running export rows after controller loss. Wrong company, issue, run, service, phase, and completed-operation cases remain inactive. - Final recursive typecheck, production build, token gates, and complete local suite coverage pass. Embedded PostgreSQL startup/socket failures passed in isolated retries with the canonical test environment; no expected behavior was weakened. The recovery regression added during the earlier full run passed in its final complete 75-case file. - Before the copy-only follow-up, all 56 CI checks passed on `c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no unresolved review threads. The final native-activity staging repeat passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`. - Verified on a separate staging instance: a real failed Daytona workspace export retains its board-owned repair action and displays “Recovery needed” in the task list. The repair card remains actionable without starting another provider turn. - Real staged export-only repair: the actual task list showed “Recovery in progress” while the original run had a positively identified running export operation, then Done after exact copyback of all 20,000 nonce-bound files. The accepted result and full provider session/turn/terminal envelopes remained unchanged. All 17 independent final checks passed; the browser downloaded the exact 19-byte result. The separate fixture was cleaned up with independent provider-absence verification. A control transport process restarted during repair; it did not submit another provider turn. - Final deployed-source repeat on `d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral Daytona allocation, actual browser export repair, and a saved full activity projection referencing the exact executing export operation with no scheduled retry. The task list showed recovery in progress, then Done; all 21 final checks passed, including 20,000 exact host files, unchanged provider provenance, and provider deletion only after committed copyback. The downloaded 19-byte result matched independently. This integrates #14334; no source change was required here. ## Risks - The label depends on the persisted recovery action. A separate runtime defect can still stop work; this change makes that condition visible. - Future recovery kinds must supply a valid continuation path before they can display progress. ## Model Used OpenAI Codex, based on GPT-6, with code execution, browser testing, and subagent tool use. The runtime does not expose an exact serving model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cbc5132e6c |
fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent adapters own local or remote provider processes. > - Run cleanup must know when the final provider process has stopped. > - A remote timeout or a lost transport does not prove that the provider stopped. > - Retried provider invocations also make an earlier stop signal stale. > - This pull request adds a verified final-invocation stop callback before workspace restoration. > - The dependent instruction revision change uses that boundary to preserve private instruction edits safely. ## Linked Issues or Issue Description **What happened?** Adapter completion did not expose a reliable point between provider shutdown and workspace restoration. Cleanup could lose provider-written files, or treat a remote timeout as proof that a process stopped. **Expected behavior** Cleanup runs once after the final provider invocation has a verified stop receipt and before workspace restoration. An incomplete remote command keeps collection pending. **Steps to reproduce** 1. Run a remote provider that returns a timeout without a numeric exit code. 2. Let the adapter return or retry the provider. 3. Attempt to collect provider-written files during cleanup. The adapter has no verified final-process boundary to use. This is the prerequisite for the stacked canonical instruction revision pull request. It has no database or UI dependency. Related #13291 verifies remote termination for later recovery; this change exposes the earlier adapter-owned stop boundary before workspace restoration. The scopes do not duplicate each other. ## What Changed - Add a stop callback to adapter execution context and a final-invocation fence. - Confirm local child closure and complete remote exit receipts. Reject remote timeouts, missing exits, transport failures, and SSH exit 255 as stop proof. - Invoke collection once before workspace restoration in eight CLI adapters and at the confirmed ACP stop boundary. - Preserve the stop observation when a later log flush fails. - Run bridge and workspace cleanup in `finally` even when collection rejects. ACP records a safe error without exposing a raw filesystem path. Grok keeps collection errors separate from workspace restore failures and preserves completed provider results when both cleanup steps fail. - Correct the existing Cursor test shell fixture so bounded remote file reads run against real fixture files. ## Verification - Three stop-boundary regressions failed before the callback implementation and passed after it. - Independent prerequisite branch: 377 tests passed across 24 adapter, process-target, and ACP suites. Two added collector-rejection tests failed before the cleanup fix and passed after it. - A third regression reproduced Grok misclassifying a collection failure as failed workspace restoration. Two additional cases covered completed and failed provider turns when collection and restore both fail. The Grok and restore-classifier suites passed 48 tests. - All nine affected package typechecks and affected package builds passed; Grok checks passed again after its classification fix. - Integrated instruction branch: native local, legacy local, and native Daytona each passed three browser tasks with exact persisted bytes, fresh-task readback, history/restore, and explicit conflict resolution. Unchanged warm Daytona passed three turns. Legacy Codex passed all five checks again after the exception-safe cleanup fix. - Alternate staging passed the same three-task native Daytona flow: exact stopped-run save, independent downloaded readback, browser history/restore, and explicit resolution of a real concurrent edit. The deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which covers the initial adapter callback. Later Grok cleanup failures are qualified by the adapter tests above. - The final combined native Daytona flow passed again on deployed `2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved ordinary instruction edits, independent readback, History/Restore, concurrent board conflict, and explicit candidate resolution. This native staging flow does not claim to exercise the Grok adapter. ## Risks - An unverified remote stop intentionally does not trigger collection. A later controller with verified stop evidence must recover it or report the copy unavailable. - The callback is optional. Callers that do not register it retain their existing behavior. - The callback runs before workspace restoration and can delay cleanup if its caller does not bound its own work. The dependent instruction collector uses bounded reads and retries. A rejected callback still permits bridge and workspace cleanup; it cannot claim an instruction save. ## Model Used OpenAI Codex, GPT-6, with tool use and code execution. The runtime does not expose the exact deployment variant or context-window size. Multiple Codex agents implemented and verified the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4b38db9622 |
fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed runs copy a selected workspace to an execution environment and restore its changes. > - Git snapshots select the files for that copy and for later recovery. > - A fixed output limit stops large generated trees before the run can start. > - Increasing the limit still keeps the complete filename lists in memory. > - This pull request stores those lists and merge baselines in disk manifests. > - Large snapshots can now complete with bounded filename buffers and explicit failure handling. ## Linked Issues or Issue Description Refs #14194. This is the streaming follow-up to the merged 32 MiB limit fix. Related: #13619 and #11621 cover workspace scan admission and demand. This change keeps the shared scheduler and changes the snapshot data path. ## What Changed - Stream changed, untracked, deleted, and ignored paths through the shared scheduler and the standalone adapter path. - Use SQLite manifests for file selection, duplicate removal, ignored-path lookup, baseline capture, and merge lookup. - Set a configurable 30-minute snapshot deadline. Keep the existing interactive scan deadlines. - Wait for each child process and pending sink write before removing temporary storage after failure or cancellation. - Use NUL archive lists and bounded deletion batches. Preserve unusual names, explicit selection, nested repositories, and source root checks. - Store manifest references in native recovery format v2. Check their location and digest before recovery reads. Keep v1 descriptors readable. - Remove temporary manifests at lifecycle completion. Use fixed-size temporary copy names for long basenames. - Admit each manifest with a SQLite page allowance based on current disk capacity. Keep a configurable free-space reserve and fail explicitly when either limit is reached. - Preserve a host file that replaces a directory deleted by the sandbox, and continue the rest of the restore. ## Verification - Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54 successful checks/statuses and two skipped Storybook jobs. No checks failed or remain pending. - [CI passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368): typecheck, build, all test shards, E2E, Rust checks, and the aggregate verify job. - [Greptile is 5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338) on the current head. All four review threads are resolved. Security checks passed. - 229 focused tests passed across Git sync, runtime staging, merge, manifest integrity, native recovery, and the scheduler (214 adapter/runtime tests and 15 scheduler tests). - A real 40,000-file fixture produces 43,428,890 filename bytes. The original standalone and scheduled scans fail. The new test passes all four filename paths, complete staging, exclusion of late files, unusual names, and deletion replay. - Recovery tests reject changed bytes, symlinks, and paths outside the controller state directory. Adapter-utils typecheck passed. - A test executor returned buffered output and caused two retry integration failures. The fixture now uses the shared streaming scheduler. All 13 tests passed with `corepack pnpm exec vitest run server/src/__tests__/heartbeat-project-repositories.test.ts`. The same CI shard now passes. - Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally. Each full local command hit SIGKILL/exit 137 in the 4 GiB container. These local commands did not pass. The current-head CI gates above provide the full verification. - A real-Git disk-capacity regression confirms a typed failure and removal of the incomplete manifest. Repeated writer attempts cannot exceed the permitted page count. - Follow-up real Daytona and separate staging qualification passed with the related archive validator (#14315) and exact-owner finalization fix (#14314). Three successive turns copied back all 60,000 files with 39,828,890 filename bytes and five unusual names. Independent host inventories verified every file and the pinned Git HEAD. Native, provider, session, and process identities stayed fixed; no retry remained. The task reached Done, and its browser-downloaded final proof matched exactly. The reusable regression is #14316, including an assertion of the effective environment idle policy. ## Risks - SQLite manifests use disk space. Each receives one quarter of the available capacity above the host reserve at creation. The reserve defaults to 256 MiB and has a 64 MiB configuration minimum. Disk capacity, filesystem quotas, per-path limits, Git resource use, and execution deadlines remain limits. - Each path and sink chunk has a 64 KiB limit. SQLite connections use a 1 MiB page cache. Invalid or incomplete records fail explicitly. - Restore transport keeps fixed and configured archive exclusions. A remotely created Git-ignored file can be transferred, but the host merge excludes it through the manifest. - Provider archive buffers, Git and tar memory, repository metadata, legacy v1 arrays, and the separate referenced-source resolver retain their own limits. Existing provider safety validators still buffer textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64 MiB. These separate transport limits can stop a sufficiently large restore before merge. This change does not claim bounded total process memory or unlimited transport size. - New descriptors use v2. Existing v1 recovery remains supported; a downgrade cannot read v2 descriptors. ## Model Used OpenAI GPT-6 through Codex. The exact deployment ID and context limit are not exposed in this run. The agent used code editing, terminal execution, tests, and GitHub tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
890d11137f |
fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.
## Linked Issues or Issue Description
**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.
**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.
**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.
**Paperclip version or commit**
Base commit:
|
||
|
|
bacc0e6a98 |
docs: remove automatic agent escalation from coordination skill (#14188)
## Thinking Path > - Paperclip manages work for AI agents. > - The coordination skill tells agents how to handle blocked work. > - The skill directs blocked work to managers and other agents. > - Those agents may lack the same access or authority. > - These extra assignments delay the required human action. > - This pull request removes automatic escalation advice. > - Agents must identify the missing capability and use the correct approval or human-input path. ## Linked Issues or Issue Description **Issue type** Incorrect information. **Where is the issue?** The critical rules in `skills/paperclip/SKILL.md` and blocker guidance in `skills/paperclip/references/api-reference.md`. **What's wrong?** The skill tells agents to escalate through `chainOfCommand`, ask another agent for help, and avoid human help. A manager title does not grant permission to fix a connection or complete an administrator action. **Suggested fix** Remove blanket escalation and agent-first rules. Keep normal delegation when the recipient has a concrete capability for a bounded task. Use saved human-input interactions or existing connection and approval flows for human-only actions. Searches for open PRs with “escalation”, “chainOfCommand”, and “ask another agent” found no duplicate skill change. Recovery-routing PRs change server behavior, which is outside this change. ## What Changed - Remove the chain-of-command escalation rule and both repeated agent-first directives. - Replace manager handoffs in the API reference with direct blocker handling. - Keep reporting fields, normal delegation, approval gates, and the ban on bypassing permission denials. - Keep the ban on cancelling cross-team tasks. Request a decision instead of automatically assigning the task to a manager. ## Verification - Final head `1483e82cc3fe10a7c910230b3779635d6d54c162`: Greptile 5/5, no outstanding findings, all CI checks pass (optional Storybook jobs skipped). - Created a fresh worktree at `.worktrees/skill-blocker-guidance` from the current master commit. - Ran `git diff --check`: passed. - Ran Node assertions against both documents: passed. Removed directives are absent. Human-input and capability-based delegation guidance is present. - Checked that checkout conflict, approval, dependency, and normal delegation instructions remain. - Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. None could run because this environment has no `pnpm` executable. - Refreshed both generated capability inventories and their derived contract after skill heading positions changed. Both generator and live-inventory checks pass. The eval and MCP baselines remain unchanged. - Ran `node --test packages/paperclip-runner/scripts/check-capability-inventory.test.mjs`: all four tests pass. - Ran `node packages/paperclip-runner/scripts/check-capability-inventory.mjs` and `node packages/paperclip-runner/scripts/generate-capability-contract.mjs --check`: both pass. - Added explicit requester routing for agent and human scope questions. Focused Node assertions pass. ## Risks - Agents can request human input earlier for actions that require human authority. - Existing runs or installed copies can keep old skill text until refreshed. - This change does not alter server recovery routing or permission checks. ## Model Used OpenAI Codex, with reasoning, tool use, and shell execution. The runtime does not expose a verifiable exact model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0f14d26123 |
fix(runtime): allow bounded large untracked workspace snapshots (#14194)
## Thinking Path > - Paperclip manages AI agents and their work. > - Remote runs need a snapshot of the task workspace before the agent starts. > - The snapshot lists untracked filenames through the shared Git scan scheduler. > - A generated directory with a few thousand long filenames can exceed the 1 MiB output limit. > - This stops setup and prevents the agent from continuing its task. > - This change gives that listing a 32 MiB bound and keeps the explicit file snapshot. > - Normal generated trees can now pass setup, while larger snapshots still fail at a finite limit. ## Linked Issues or Issue Description Refs #11572 and #12214 for the existing bounded scan and ignore-scan protections. **What happened?** An agent continuation failed during workspace setup with `Workspace Git scan exceeded its output limit`. A real Git fixture reproduces the untracked-file path: 5,000 long filenames in one generated directory exceed its 1 MiB output limit. **Expected behavior** The workspace snapshot must support ordinary generated trees with thousands of files. It must keep a finite output bound and select explicit files before staging. **Steps to reproduce** 1. Commit a base file in a Git repository. 2. Add 5,000 untracked files with long names in one new directory. 3. Call `readGitWorkspaceSnapshot` with the normal scan limits. 4. Observe the output-limit error before this change. **Paperclip version or commit** Base commit: `640dee1`. **Deployment mode** Remote sandbox execution from source. ## What Changed - Increase the untracked-file snapshot output bound from 1 MiB to 32 MiB. - Keep explicit file selection, the shared scheduler, the timeout, and the other scan bounds. - Test that 40,000 long filenames above the old 8 MiB bound reach the snapshot. Reuse those files with deeper paths to prove that output above 32 MiB still fails. - Test that files created after the snapshot, including an ignored secret, stay out of the overlay archive. - Check the workspace root identity and reject selected paths with symlinked parent directories before upload. - Test root replacement after path resolution, including root-level selected files. - Preserve existing workspace root aliases by capturing the resolved root before snapshot selection. Test that later alias retargeting cannot change the archive contents. - Stop staging on permission and I/O errors; continue to allow missing files. - Test these failures and preserve selected symlink entries. - Accept valid case-renamed directories by checking ancestor file types. A modeled case-insensitive regression failed before this correction and now passes. - Give the 40,000-file fixture enough time to remove its files. - Document the larger bound and staging behavior. ## Verification - Red: the 5,000-file regression failed with `stdout maxBuffer length exceeded` before the fix. - A separate check through the real server scheduler reproduced `workspace_git_scan_output_limit` on the original code. The revised code selected all 5,000 files. - The late-file regression failed against the first PR revision because the archive contained `drafts/late.secret`. It passes with the final explicit-file approach. - The staging regressions failed before the review fix: a substituted parent directory and permission/I/O errors were accepted. All three cases now stop before upload. - Green: 134 tests passed across `git-workspace-sync.test.ts` and `sandbox-managed-runtime.test.ts` with Vitest 4.1.11. This includes complete selection above 8 MiB and rejection above 32 MiB. - `pnpm --filter @paperclipai/adapter-utils... typecheck` passed after the revision. - The module-boundary check and `git diff --check` passed. - `pnpm -r typecheck` and `pnpm build` stopped in the Rust runner steps because this environment has no `cargo` executable. - The full `pnpm test:run` attempt ended with `SIGKILL` during the general server suite. It did not finish. That full-suite result belongs to the earlier revision. Fresh checks are required for this revision. - The new regression fails at the old 8 MiB bound with `stdout maxBuffer length exceeded`. All 134 focused tests pass with the 32 MiB change. - The affected typechecks and module-boundary check pass. The full local typecheck requires Cargo, which is absent in this environment. - Final verification for `b43e9bc95e55382c6a9bfe200487c164770c8be8`: 54 successful checks/statuses and two skipped Storybook checks. No checks remain pending or failed. - The [CI run](https://github.com/paperclipai/paperclip/actions/runs/36315569674) passes on attempt 2. The first attempt had one unrelated preview-fixture readiness timeout. That exact test passed locally; its CI shard passed on the single rerun. - [Greptile reports 5/5](https://github.com/paperclipai/paperclip/pull/14194#issuecomment-5852045501) on this revision. All review threads are resolved. - [Security review accepts the documented memory tradeoff](https://github.com/paperclipai/paperclip/pull/14194#discussion_r4115149465) for this finite mitigation. The separate streaming follow-up will remove full-list buffering. - This revision also passes affected local typechecks and the module-boundary check. Full local typecheck/build stop because Cargo is absent. The full local Vitest attempt was stopped after about 18 minutes once all remote gates passed; it did not complete locally. ## Risks - Each untracked-file scan can buffer up to 32 MiB instead of 1 MiB. The scheduler still limits concurrent scans and execution time. - Snapshots above 32 MiB still fail with the existing error. Tracked and ignored-file scan bounds stay unchanged. - A workspace that replaces a selected path’s parent with a symlink now fails staging. - Detailed logs from the reported host were unavailable. The exact command that exceeded its limit on that host is unconfirmed. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The runtime does not expose a more specific model build or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e0f051d3b |
Align task property icons with assignee avatars (#13319)
## Thinking Path > - Paperclip is the open source app that people use to manage AI agents for work. > - The task properties panel helps operators inspect and update task data. > - The status, assignee, and project rows used different leading icon sizes. > - The mixed sizes made the rows look uneven. > - The assignee avatar already provides the correct 24px visual size. > - This pull request sets the status and project visuals to that same size. > - The benefit is a tidy and consistent properties panel. ## Linked Issues or Issue Description **What happened?** The task properties panel showed a 12px status icon, a 24px assignee avatar, and a 16px project tile. **Expected behavior** The status, assignee, and project rows must use the same 24px leading visual size. **Steps to reproduce** 1. Open a task. 2. Open the Properties panel. 3. Compare the Status, Assignee, and Project rows. ## What Changed - Keep the Status glyph at 16px and use spacing to give it the 24px avatar footprint. - Set the Project row tile to the 24px small tile size. - Add a regression check for the Status row size. ## Verification - `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` (72 tests passed) - `pnpm check:token-gates` (all gates clean) - `pnpm --filter @paperclipai/ui typecheck` (passed) - Full repository typecheck and build reached the Rust runner step, but this environment does not contain `cargo`. - The full test command was stopped after unrelated chat integration tests did not finish. The focused UI suite passed. ## Risks - Low risk. This change only changes visual sizes in three property rows. - The wider status footprint can use slightly more horizontal space in a narrow panel. > This fix does not overlap with planned core work in `ROADMAP.md`. ## Model Used - OpenAI Codex, GPT-5.6, tool use and code execution enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d423fc4411 |
fix(ui): use large mobile selector modals (#14250)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The new task dialog lets an operator assign work on a phone. > - The mobile picker used fixed coordinates inside the Radix wrapper. > - iOS Safari already makes fixed coordinates relative to the visual viewport. > - The old code added the visual viewport offset a second time. > - This pull request removes the second offset and gives the mobile wrapper full viewport geometry. > - The benefit is a visible picker and search field when the iOS keyboard opens. ## Linked Issues or Issue Description **What happened?** The assignee and project pickers could move outside the visible screen in the new task dialog on iOS. **What did you expect to happen?** The picker and its search field must stay visible above the bottom edge and the software keyboard. **Steps to reproduce** Open the new task dialog on an iPhone. Select Assignee or Project. The picker can move outside the visual viewport when Safari pans the page. **Paperclip version or commit** The problem exists after the change in PR #13343. **Deployment mode** The problem affects the board UI in local and authenticated modes. Refs #13343 ## What Changed - Use visual viewport local coordinates for the new task dialog. - Give the mobile Radix popper wrapper full viewport geometry. - Render both shared selector components as large, titled mobile modals. - Keep desktop selectors as anchored popovers. - Cover the new-task assignee, new-task project, composer assignee, and generic searchable selector in Storybook. - Add iPhone-sized Storybook states for the assignee and project pickers. - Update viewport tests for the corrected coordinate model. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/NewIssueDialog.test.tsx src/components/InlineEntitySelector.test.tsx` - `pnpm check:token-gates` - Playwright used the iPhone 14 device profile. All mobile modals were inside the 390 by 664 CSS viewport. - Each mobile modal was `x=16, y=16, width=358, height=632`. - A desktop check kept the anchored popover at `width=320, height=211`. ## Risks - Low risk. The CSS only changes narrow mobile viewports. - The full-screen popper wrapper does not receive pointer events. The picker still receives pointer events. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.6, Codex agent with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task ID - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant Storybook documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
74508357fa |
fix(ui): keep mobile task pickers in view (#14249)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The new task dialog lets an operator assign work on a phone. > - The mobile picker used fixed coordinates inside the Radix wrapper. > - iOS Safari already makes fixed coordinates relative to the visual viewport. > - The old code added the visual viewport offset a second time. > - This pull request removes the second offset and gives the mobile wrapper full viewport geometry. > - The benefit is a visible picker and search field when the iOS keyboard opens. ## Linked Issues or Issue Description **What happened?** The assignee and project pickers could move outside the visible screen in the new task dialog on iOS. **What did you expect to happen?** The picker and its search field must stay visible above the bottom edge and the software keyboard. **Steps to reproduce** Open the new task dialog on an iPhone. Select Assignee or Project. The picker can move outside the visual viewport when Safari pans the page. **Paperclip version or commit** The problem exists after the change in PR #13343. **Deployment mode** The problem affects the board UI in local and authenticated modes. Refs #13343 ## What Changed - Use visual viewport local coordinates for the new task dialog. - Give the mobile Radix popper wrapper full viewport geometry. - Keep the mobile picker at the visible bottom edge. - Add iPhone-sized Storybook states for the assignee and project pickers. - Update viewport tests for the corrected coordinate model. ## Verification - `pnpm --filter @paperclipai/ui exec vitest run src/components/NewIssueDialog.test.tsx src/components/InlineEntitySelector.test.tsx` - `pnpm check:token-gates` - Playwright used the iPhone 14 device profile. Both picker boxes were inside the 390 by 664 CSS viewport. - The assignee picker box was `x=16, y=366, width=358, height=282`. - The project picker box was `x=16, y=366, width=358, height=282`. ## Risks - Low risk. The CSS only changes narrow mobile viewports. - The full-screen popper wrapper does not receive pointer events. The picker still receives pointer events. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`. ## Model Used - OpenAI GPT-5.6, Codex agent with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal task ID - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant Storybook documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
640dee1802 |
fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path > - Paperclip coordinates agent work across task turns. > - Warm runners need reusable sandbox leases. > - Lease acquisition read only the environment reuse setting. > - Startup then read the agent's new warm setting and rejected the ephemeral lease. > - This fix requests reuse before acquisition, with provider checks intact. ## Linked Issues or Issue Description **What happened?** Switching an existing native runner task from per-turn to warm fails with `runner_warm_environment_requires_reusable_lease`. Environment `reuseLease` defaults to false. Acquisition therefore creates an ephemeral lease, and capability narrowing correctly denies reuse. **Expected behavior** Changing lifecycle between turns requests a compatible lease without replacing the task workspace or changing the shared environment. **Steps to reproduce** 1. Start a native runner task with inherited lifecycle and environment reuse disabled. 2. Change the agent from per-turn to warm. 3. Continue the task on a reusable-capable sandbox provider. **Paperclip version or commit** Reproduced on `a6c4e7a`. Related: #12904 established warm workspace continuity; this fixes the earlier lease-selection mismatch. ## What Changed - Pass agent settings into acquisition and derive a run-scoped reuse request. - Respect explicit environment lifecycle overrides; leave other adapters unchanged. - Keep provider capability, ownership, cleanup and restore gates intact. - Add orchestration and database-backed transition tests; document the contract. ## Verification - Red: three new assertions failed before the fix. - Green: 156 tests pass in environment-run-orchestrator, environment-runtime, native-sandbox-lifecycle, and environment-execution-target-capabilities. - Database-backed regression proves ephemeral → reusable → resumed lease, same workspace, and unchanged stored environment. Provider RPCs are mocked, not live Daytona. - `git diff --check` passes. - Latest-head CI passes typecheck, build, tests, runner checks, E2E, and canary dry run. Greptile: 5/5; no open threads. - Recovery uses the persisted lifecycle. Direct regression: red before, six cases green after. CI runs the added Vitest cases. - Full local checks were unavailable (missing dependencies and earlier memory limits). ## Risks Warm mode now requests retained resources despite an environment's default `reuseLease:false`. Existing idle timeout and cleanup still apply. Unsupported providers remain denied. No migration, credential [REDACTED], or shared-environment mutation. ## Model Used OpenAI Codex agent; reasoning, code editing and test execution. Exact model ID and context size were not exposed to this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details available) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket [REDACTED] or instance-derived details - [x] I have run targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2ed0b65c4 |
fix(runner): enable API tools by default (#14186)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.
## Linked Issues or Issue Description
Refs #13003, which added the guarded API tools.
**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.
**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.
**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.
**Paperclip version or commit**
|
||
|
|
01d9a12185 |
fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings. Co-Authored-By: Codie <Codie@users.noreply.github.com> |
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bd20309323 |
fix(ui): keep mobile task pickers above the keyboard (#14022)
## Thinking Path > - Paperclip helps people manage AI agents and their tasks. > - The New Task dialog lets a person select an assignee and a project. > - A mobile browser reduces the visible viewport when the software keyboard opens. > - The dialog accepted invalid viewport data, and the pickers stayed inside a transformed container. > - This behavior could collapse the dialog or move a picker search field above the visible area. > - This pull request validates viewport data and puts mobile pickers in the visible viewport. > - The benefit is that a mobile user can see and use each picker while the keyboard is open. ## Linked Issues or Issue Description **What happened?** On a mobile device, the Assignee and Project pickers in the New Task dialog could move above the visible viewport. A short invalid viewport value could also collapse the dialog to a line. **Expected behavior** The dialog and each open picker must stay in the visible viewport while the software keyboard is open. **Steps to reproduce** 1. Open the New Task dialog in a mobile browser. 2. Open the Assignee picker or the Project picker. 3. Focus the picker search field so that the software keyboard opens. 4. Observe that the picker can move above the visible viewport. **Paperclip version or commit** `efce9356b5` **Deployment mode** Local dev and hosted browser UI. ## What Changed - Ignore zero, negative, and non-finite Visual Viewport measurements. - Keep the last safe dialog geometry until the browser gives a valid measurement. - Put entity pickers outside the transformed dialog container. - Size and position the mobile picker from the dialog Visual Viewport values. - Add unit tests for invalid viewport recovery. - Add Chromium tests for the Assignee and Project picker states. ## Verification - `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx` passed with 34 tests. - The Chromium viewport test passed 35 of 35 runs with five repeats and no retries. - `pnpm --filter @paperclipai/ui typecheck` passed. - `pnpm check:token-gates` passed. - `pnpm build-storybook` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - A full local Vitest run reached restricted workspace-runtime tests that require sibling worktree and runtime writes. Remote CI will run the supported test environment. ## Risks - Risk is low because the new picker layout applies only to mobile widths. - The layout depends on Visual Viewport data when the browser supplies valid values. - Unit and browser tests cover invalid data, mobile pickers, tablet layout, and desktop layout. > This change fixes a focused UI bug. It does not duplicate a core feature in `ROADMAP.md`. ## Model Used - OpenAI Codex with GPT-5. The agent used reasoning, repository tools, code execution, and browser automation. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dbffb4c04 |
fix(daytona): clean up sandbox allocation after create failure (#13979)
## Thinking Path > - Paperclip manages AI agents and their execution environments. > - The Daytona provider creates sandboxes before it returns a lease. > - Daytona can allocate a sandbox, then time out while waiting for it to start. > - The SDK throws without returning the allocated sandbox handle. > - Paperclip previously lost the resource identity and could leave a running sandbox behind. > - This pull request keeps an exact creation identity and deletes a matching failed allocation. ## Linked Issues or Issue Description Related: #13882. Searches for Daytona create cleanup and orphan fixes found no duplicate for this path. **What happened?** In [Grok qualification campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743), Daytona sandbox creation exceeded 300 seconds. No native runner or provider session started. The SDK threw, but an ownership-filtered provider lookup found the allocated sandbox still running. The disposable resource was then deleted separately. The original attempt and its failure remain retained. **Expected behavior** A failed create must retain enough identity to clean up an allocated sandbox. Cleanup must never delete another attempt's resource. If the provider cannot confirm cleanup during sandbox lease acquisition, the host must retain a durable cleanup record and retry the exact owned allocation after restart. **Steps to reproduce** 1. Have Daytona allocate a sandbox for an environment acquisition. 2. Make the SDK throw while waiting for startup, before it returns the sandbox handle. 3. Before this fix, acquisition fails without deleting the allocated sandbox. 4. The regression tests reproduce this failure and verify exact-owner deletion. **Paperclip version or commit** Observed at `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`. The same creation path exists on master; this fix starts from `efce9356b`. ## What Changed - Assign each cold create a unique provider name and creation-attempt label. - On create failure, look up that exact name and verify every ownership label before deletion. - Wait for provider deletion with bounded lookup and deletion calls. - Preserve the original create error after successful cleanup. Report both errors and the provider name when cleanup cannot be confirmed. - Treat a missing lookup after an uncertain create as unconfirmed cleanup: a delayed provider request might still create the sandbox. - Transfer only validated ownership fields through the acquire-lease RPC error; provider exceptions and arbitrary error data are not serialized. - Persist a pending-cleanup lease before the host retries deletion. Reuse the existing cleanup sweep and durable spool, retaining the original environment scope after its row is deleted. - Fence retries by company, environment, run, provider account, unique creation name, and every ownership label. Missing provisional-name lookups remain unresolved. - Journal an ownership-verified provider ID before host-driven deletion. Its absence then confirms cleanup after a lost deletion reply or final database update; failed journal writes block deletion. - Keep ownership labels separate from the mutable input Daytona modifies during create. - Report confirmed host-side deletion accurately while still rejecting the failed acquisition. - Add protocol, malformed-evidence, cross-scope denial, controller-restart, and deleted-environment regressions alongside the original bounded cleanup tests. ## Verification - SDK suite: 92 passed. Daytona plugin suite: 277 passed; six opt-in live tests skipped. Environment-runtime suite: 103 passed, then all four focused observation/recovery variants passed after adding a foreign-observation case. SDK and provider TypeScript checks and `git diff --check` pass. - Regressions failed before the respective fixes: missing durable ownership, misleading successful-cleanup message, and SDK mutation of the ownership labels. - Controlled real-Daytona proof at `d2ac1c95b32ca64daf039c0427137657f351c0e0`: create a real sandbox; inject an SDK failure and a failed immediate lookup; persist the validated ownership envelope; have a child journal the observed provider ID before deleting; deliberately drop its successful deletion receipt; reconcile absence from another fresh process and confirm it with a separate provider read. Passed, with no manual cleanup and zero model calls. The live proof uses an owner-only envelope file; actual host database/spool recovery is tested separately in the environment-runtime suite. - The preceding controlled attempt failed because the real SDK mutated the labels object. That failure is retained. The harness deleted its exact owned sandbox, and a fresh lookup confirmed absence. The fix copies the SDK input labels separately from the ownership snapshot. - [Repository CI](https://github.com/paperclipai/paperclip/actions/runs/36095925399) passes on this exact head, with a clean 5/5 review. The unchanged Cursor remote-command test passed its one bounded rerun after a 10-second timeout; the same test had already passed on the combined tree. The original failure is retained. Earlier `199f0852` CI and immediate-deletion proof are retained, without being promoted to the new source. - Combined-source verification uses temporary PR #13990 against the Grok feature branch. No Docker or Rust build runs on the developer laptop. ## Risks Creation failures now add up to ten seconds for lookup and fifteen seconds for deletion confirmation. The SDK still owns the initial creation timeout. If a request materializes only after the failed lookup, its durable pending record remains eligible for later cleanup; a never-observed creation name is never reported deleted from a 404. A previously observed provider ID can be reconciled as deleted. The host and SDK must be deployed together. Durable handoff applies to plugin-backed sandbox lease acquisition used by heartbeat and runner login. Probe and custom-image interactive setup retain immediate cleanup and explicit failure reporting, but do not use this lease-recovery path. A worker crash before delivering the failure envelope still cannot be recovered through this mechanism. A cleanup error never returns a lease. Successful creation behavior is unchanged except for the provider name and ownership label. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f56aee5423 |
fix(evals): capture complete durable run event streams (#13977)
## Thinking Path > - Paperclip manages AI agents and records their durable work outcomes. > - Product E2E checks those outcomes through the browser and public API. > - A successful long run can emit more than 1,000 durable events. > - The harness read one page and missed the later completion evidence. > - This pull request reads every page before it checks runtime invariants. > - Invalid or incomplete capture still fails. A missing page cannot produce a pass. ## Linked Issues or Issue Description Related: #13882 (Grok qualification). A duplicate search found no existing event-pagination fix. **What happened?** The local structured-question case in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537) completed the task and passed its six outcome matchers. It failed native runtime invariants because the capture contained exactly 1,000 events. The last captured event preceded the run's completion by more than a minute. The API caps each response at 1,000 rows. The harness did not request the next page. **Expected behavior** Read the complete durable event stream through the public API before checking semantic-result and terminal-event counts. Reject incomplete or malformed evidence. **Steps to reproduce** 1. Complete a native task that emits more than 1,000 durable events. 2. Place the semantic-result and terminal events after row 1,000. 3. Capture the run with the Product E2E harness. 4. Before this fix, the invariant checker sees only the first page. **Paperclip version or commit** Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same single-page capture exists on master. The original failed result remains unchanged; missing historical tail evidence is not reconstructed or graded as a pass. ## What Changed - Add a bounded event collector that advances through the public `afterSeq` cursor. - Use it for task success/failure evidence and shared chat run evidence. - Reject invalid pages, missing or non-increasing sequence numbers, repeated cursors, failed later requests, and an exhausted page limit. - Test completion events beyond the first page, exact page boundaries, and malformed evidence. - Correct the existing Everyday catalog test from 38 to the maintained 47 cells. The suite stays explicit-only. - Document the complete-capture requirement and its bound. ## Verification - Product harness typecheck passes. - The full credential-free harness suite passed 477 tests. After adding the chat integration regression, all 34 chat evidence tests pass. - Seventeen pagination tests cover the valid tail and malformed-evidence cases. - `git diff --check` passes. - [Full repository CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422) passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful checks and two intentional skips, including Rust, typecheck, build, server tests, and browser shards. Greptile gives this exact head 5/5 with no findings. No local Docker, Rust build, browser suite, or model invocation was used for this change. - [Live Grok requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743) is running on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler, and evidence fixes. Its full credential-free harness passes all 493 tests and typecheck. Live results are pending; the original campaign remains a failed measurement. ## Risks Long runs need more read-only API requests and larger private evidence files. Capture stops with an explicit error after 100 full pages. This changes neither production APIs nor provider behavior. It does not relax an invariant or change a historical grade. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c96c751f39 |
fix(heartbeat): claim task ownership with queued runs (#13973)
## Thinking Path > - Paperclip manages agents and their tasks. > - The scheduler can start several runs for one agent. > - Each task still needs one execution owner. > - Task creation and recovery can queue assignments for the same task. > - The scheduler previously marked both runs running before starting either executor. > - This change claims the run and task ownership in one transaction. > - Tasks with different IDs retain concurrent execution. ## Linked Issues or Issue Description Refs #13882. Related queue work: #11830 and #13965; those address different admission and cancellation paths. **What happened?** The Grok subscription Product campaign exposed two assignment runs for one task. Task creation and periodic recovery started runs within 200 ms. Both acquired Daytona sandboxes. One failed before provider work with `paperclip_runner_attachment_staging_not_authorized` because the other owned the task. Fixture cleanup then found an active session. The original failed result is retained in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537). **Expected behavior** Only the task execution owner may start provider setup. Competing queued work must wait. Different tasks may use the agent's available slots. **Steps to reproduce** 1. Set an agent's concurrent run limit to two. 2. Queue task-creation and recovery assignment runs for the same task. 3. Resume the queue while holding adapter execution open. 4. Before this fix, both runs become running. Only one has the task execution lock. **Paperclip version or commit** Reproduced on master `8781f06a8` and in the Grok campaign at `4196a4cd`. ## What Changed - Use the same company-scoped task ownership gate for assignments, direct comments, and queued comments. - Leave competing work queued while another live run owns the task. Transfer a terminal pointer only after the tracked executor and durable environment leases/finalization settle, including cleanup on another controller. - Commit the running state and task execution owner together. - Preserve the dedicated review path and the native replacement checkout guard. - Add sixteen database regressions for assignment/comment orderings, unrelated concurrent tasks, and cross-controller cleanup fences. ## Verification - Live subscription Product E2E planning passed on its first attempt in both Daytona (351,197 ms) and local execution (268,883 ms), with 6/6 matchers and cleanup passing in each, on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e` ([campaign](https://github.com/paperclipai/paperclip/actions/runs/36080870743)). The local screenshots and persisted state show the same plan revised, the first approval rejected, the revised approval accepted, and one completion. The full campaign remains unqualified because of a separate startup retry and EC2 Spot interruption. - The new same-task regression failed before the fix: two running rows instead of one. - All seven focused concurrency cases passed. Three comment cases failed before the shared gate was added. Five cross-controller cleanup cases failed before the durable gate was added. A warm-retention regression also failed before its successful release receipt was admitted; missing/failed receipts and a different retention policy stay blocked. - `node ../node_modules/vitest/vitest.mjs run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` from `server`: 48 passed. Native cleanup admission adds one passing test; task-drain release adds two. Total focused checks: 51 passed. - `git diff --check`: passed. - Current-head [CI](https://github.com/paperclipai/paperclip/actions/runs/36078495577) passes repository typecheck, tests, Rust checks, build, and browser suites. The unrelated repository-form browser shard passed one bounded rerun without assertion or code changes; its original navigation failure remains retained. No local Docker or Rust build was used. - The failed test's exact two sandboxes were checked: one was already absent; the other was deleted after its company, environment, and run labels matched. ## Risks The claim transaction adds a task row lock. It follows task-before-run lock order. A competing run stays queued until the current owner releases the task. The change does not alter attachment authorization, provider permissions, schemas, or public APIs. Review is 5/5 with all threads resolved on `23c0d13b6`. All CI checks pass on that exact head. Live Grok requalification will use the combined runner and scheduler source. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ### Live regression verification The Daytona plan/revise/accept lifecycle in [Grok campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743) passed on its first attempt at combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, in 351,197 ms, with all six outcome matchers and cleanup passing. The earlier competing-run failure remains retained. This is one live planning result; the complete Grok matrix and repeated qualification are separate gates, and this campaign also has separately retained infrastructure failures. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
efce9356b5 |
fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The browser must load its JavaScript before React can render a task. > - A failed import can stop that process before the React error boundary exists. > - The HTML entry then leaves an empty page with no recovery action. > - This PR adds a small recovery screen that works without React. > - The user can retry the same page and return to saved task content. ## Linked Issues or Issue Description Refs #13824 and #13895. This is a follow-up to their browser startup investigation. **What happened?** Interrupting the app bundle or a required import leaves an empty React root. A startup exception has the same effect. A React error boundary cannot handle these failures because React has not started. **Expected behavior** The page must explain the startup failure and offer a manual retry. A late successful load must dismiss the recovery message without a reload. **Steps to reproduce** 1. Open a saved task in the browser. 2. Abort the application bundle request, or make a required module return HTTP 503. 3. Observe the empty page before this change. With this change, use Reload page after the fault clears and verify the saved task and comment. **Paperclip version or commit** The failing regression baseline used master at `8781f06a8`. **Deployment mode** Local source build and compiled UI. Tests cover both initial navigation and a page controlled by the production service worker. The exact cause of the older intermittent Vite stall remains unconfirmed. Forty app loads and thirty replays of retained responses did not reproduce it. This PR fixes the missing recovery path; it does not claim to remove that historical cause. A normal HTTP 304 response is not a failure. ## What Changed - Add an inline startup guard and recovery screen in the HTML entry. It does not depend on the app module graph. - Show a manual reload action after a startup error or after 30 seconds without rendered root content. - Remove the notice, timer, observer, and error listeners when the app starts. Never reload automatically. - Keep the recovery screen outside the React root so it cannot satisfy app-readiness checks. - Add browser tests for interrupted imports, a stalled import, an evaluation error, service-worker-controlled retry, repeated offline retry, and cleanup after successful startup. - Return a static, uncached HTML retry screen when a service-worker-controlled navigation fails offline. It contains no task content. - Add a full-app test that retries an interrupted compiled bundle and checks the saved task, comment, composer, route, and absence of agent runs. - Document the coverage and the limits of the historical diagnosis. ## Verification - Red baseline: four recovery cases failed; the normal-startup case passed. After the change, all five recovery cases passed. The review found an offline retry gap; that additional case failed before the worker fix and passed afterward. - Full provider-free browser-support suite: 16 passed. - Compiled-app browser tests: four passed, including saved-task reload, interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar navigation. - Expanded service-worker, offline response, PWA, and worker build-ID unit tests: 37 passed. The two old plain-text offline expectations were reproduced as failures and updated for the HTML retry contract. - UI production build, full local repository typecheck (`pnpm -r typecheck`), runner-E2E typecheck, and design token checks passed. - Manual browser check: a temporary server failed the compiled bundle once. The recovery screen appeared. Clicking Reload page restored the same saved task, comment, and composer. - Full local `pnpm build` passed. - Full local `pnpm test:run` was attempted with a bounded deadline and stopped after it timed out. Workspace runtime/cleanup tests reported timeouts on this host. The monolithic local run is not a pass. The focused tests above and the complete Linux CI run provide the successful verification. - Final-head [CI run](https://github.com/paperclipai/paperclip/actions/runs/36072201966) passed. All 53 check runs succeeded; the two Storybook jobs were intentionally skipped. The legacy security status also passed. - Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5. Both review findings are fixed and resolved. ## Risks - The guard only handles startup before React renders root content. Existing React boundaries handle later rendering errors. - A slow startup can show the message after 30 seconds. A later successful render removes it; the page does not reload by itself. - The fallback uses native HTML when the app stylesheet is unavailable. - The worker changes only its offline navigation response. It returns static HTML with a reload button and `Cache-Control: no-store`. Its cache allowlist, private-response protections, task state, provider prompts, and grading rules stay unchanged. - This does not establish or fix the unknown cause of the historical intermittent Vite stall. ## Model Used OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not an exact served model ID or context window size. Used reasoning, code editing, shell tools, and browser testing. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8781f06a87 |
feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connections let those agents use external services with explicit access rules. > - Zapier, Arcade, Composio Connect, and Executor already have setup and runtime support. > - Their experimental switch still blocks discovery and setup by default. > - This pull request removes those gates and the Settings toggle. > - Users can connect these providers without enabling an experiment. ## Linked Issues or Issue Description Refs #13755. Refs #13941. **What existing behavior does this improve?** Apps browsing, inline setup, and agent connection search for the four MCP aggregators. **Current behavior** An instance must enable the MCP aggregators experiment before users or agents can start setup. **Proposed behavior** All four providers are available by default on local and managed instances. Old stored and managed values still parse but cannot disable them. ## What Changed - Remove the aggregator gates from Apps, inline setup, server setup, and agent search. - Remove the Settings toggle and its UI hook. - Retain the old setting key only for upgrade compatibility. Normalize it to true and ignore managed overrides, as Apps already does. - Replace opt-in fixtures with default-on coverage. Test old false values, all four setup flows, provider choice, and the removed toggle. - Update current connector guidance and remove the opt-in from the runner acceptance fixture. ## Verification - 306 focused tests passed across eight files: shared remote MCP contracts; server remote MCP lifecycle, aggregator fallback, settings normalization, and managed overlay; UI Apps browsing, setup, and experimental settings. - Server and UI TypeScript checks passed. - UI token gates and `git diff --check` passed. - The full local suite was not run, per the maintainer's instruction. All 54 CI checks passed; two checks were skipped. One unrelated workspace-preview readiness timeout passed on one failed-shard retry. - The setup fixtures use simulated MCP responses. This change does not claim new live provider acceptance. ## Risks - Existing instances now show all four providers, even if the old flag was false. This is intentional. - External provider choice, credentials, company isolation, agent grants, and tool policies still apply. Showing a connector does not authorize an external account. - No data migration is required. The compatibility key keeps old managed configuration documents valid. - Historical Zapier live acceptance remains incomplete in the existing evidence report. The maintainer explicitly requested the default-on rollout for all four existing providers; the report records that scoped exception. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository tools, and test execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
aa8fc86331 |
feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external services. > - Native connections should remain the first choice for a supported app. > - Other apps may be available through an external MCP provider. > - The user must know which external provider handles the connection and choose it before setup. > - This pull request adds ranked alternatives and server-authored instructions to connection search. > - Agents can follow the returned instructions while Paperclip validates saved choices and access. ## Linked Issues or Issue Description Related: #13879, which fixed inline MCP provider setup. This PR adds discovery and provider selection on top of that work. **Subsystem affected** Cross-cutting: shared connection contracts, server search and intent services, native runtime, CLI, inline setup UI, and evals. **Problem or motivation** An agent cannot offer a clear external-provider choice when Paperclip has no native connection for an app. Adding provider-specific branches to the core prompt would make those instructions harder to maintain. **Proposed solution** Prefer a native connection. Otherwise return verified alternatives in Composio, Arcade, Executor, Zapier order. Include an external-service disclosure, a question with None, and the next instruction in the search result. Validate the saved human choice before creating a selected fallback setup card. Reuse existing provider accounts and verify underlying app access separately. **Alternatives considered** Do not silently choose a provider. Do not claim that broad execution tools prove support for every app. Reuse existing questions and connection intents rather than add another connection model. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback capabilities. The MCP aggregators experiment remains the gate. No duplicate provider-routing PR was found in the public search. ## What Changed - Add a dated support index and authorized cached-tool evidence for external routes. - Return provider questions and next-step instructions from `connections_search`. - Preserve pending choices and declines across continuation. Validate company, task, agent, human, app, and current route eligibility. - Carry the selected app into new setup and account reuse, validate explicit provider requests against persisted human messages, and distinguish provider readiness from app authorization. - Sync native, MCP, REST, and CLI contracts. Keep core agent instructions provider-neutral. - Add production-component Storybooks, focused database tests, and three real-agent browser eval cases. - Record the plan, observed failures, fixes, passing evidence, and acceptance limits. ## Verification - Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and all review threads resolved. - After rebasing on master `18dac1e1e`: 64 focused shared, validator, route-contract, and database tests passed; server typecheck passed. - Embedded-browser test drive on the rebased head: native Jira card, HubSpot external-provider question, Arcade account reuse, one actual MCP read against a local synthetic fixture, reload persistence, and None preventing further calls. A real OpenAI-backed agent performed discovery and continuation. - UX observation: the agent initially combined mutually exclusive request fields; the server rejected it and the agent recovered without changing access. This extra retry remains visible in the transcript. - After rebase: 23 focused eval grader/catalog tests, affected TypeScript checks, token gates, production UI build, and Storybook build passed. Full local tests are intentionally excluded at the maintainer's request. - Before rebase: four browser/real-agent attempts passed: native Jira, None, and reuse of the second provider on two Codex profiles. - Browser evals used an isolated deterministic MCP fixture through the real Paperclip gateway. They do not prove production compatibility with all four providers. - Review `Apps / Connections / Provider choice` in Storybook. Choose Arcade, continue through Access, and verify the app name, external-service disclosure, and URL configuration. - The detailed verification report is `doc/connections/2026-09-23-aggregator-routing-verification.md`. ## Risks - The public support index is finite and can age. Account capability and app authorization still require verification after selection. - Existing installed-tool permissions remain in effect. Provider choice is not a new execution permission boundary. An early Mini attempt skipped search; clearer provider-neutral instructions made the targeted rerun pass. This is not a measured reliability rate. - Explicit requests skip provider confirmation only when a clear persisted human message or saved provider choice supports them. Other phrasing falls back to confirmation; the agent query alone is not consent. These routes do not add tool permissions. - No database migration or legacy Composio broker is added. Real-provider acceptance remains separate from fixture proof. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser tools. The exact deployment variant and context-window size are not exposed in this session. Product evals separately used the repository's primary Codex and Codex Mini profiles; those agents supplied test behavior, not independent provider compatibility proof. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b0155a681a |
feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack conversations use the same tasks and agents as the Paperclip board. > - A board reply must reach that Slack conversation and let the agent continue the work. > - An assigned agent also needs its Slack tools during normal tasks and scheduled routines. > - Both paths must keep the linked user's authority, delivery rules, and conversation history. > - This pull request adds those paths and reduces setup friction for Slack bots. ## Linked Issues or Issue Description **Subsystem affected** Server orchestration, Slack connector tools, shared contracts, and chat setup UI. **Problem or motivation** Replies entered in Paperclip did not provide a complete round trip to the linked Slack thread. Slack and board wakeups could select different model sessions for the same task. Agents also lacked their assigned Slack tools outside Slack-origin work, which prevented a routine from sending its responsible user a briefing. Inviting a bot could leave the new channel disabled. **Proposed solution** Mirror human board messages with author attribution and route the agent result to the same thread. Use the same session key across both entry points. Supply Slack tools to the connection's assigned agent in normal tasks and routines, using the current responsible user's verified link. Enable newly invited channels while preserving explicit disabled choices. Add a browser-agent setup prompt to the Slack wizard. **Alternatives considered** A separate Slack scheduler or task dispatcher would duplicate existing Paperclip workflows. Reusing the connection owner's identity would grant the wrong authority. Replaying old channel history could start unintended work. This change uses ordinary task wakeups, routine dispatch, and Slack's original invitation mention event instead. **Roadmap alignment** Extends the shipped Scheduled Routines and governed Apps capabilities. It does not add a separate task lifecycle. Related work: #13828 and #13809. Related test stabilization: #13877. The existing plugin Slack-control proposals are separate from this built-in connector change. ## What Changed - Queue human Paperclip messages for the original Slack thread with display-name attribution and stable delivery identities. Require the author’s current linked Slack identity and recheck access before delivering messages or agent replies. - Apply pause, dependency, cancellation, and closed-workspace guards before explicit Board sends request work and again when the durable outbox dispatches it. - Route agent results back to Slack and preserve model-session continuity, including replies that reopen completed tasks. - Resolve assigned Slack connections for normal agent tasks and routines. Recheck the responsible user's link, membership, and permissions at execution. - Add `slack_open_dm` for the responsible user's bot DM and request the `im:write` scope. - Enable newly discovered invited channels. Keep explicit OFF choices and normal admission and deduplication rules. - Add a copyable Slack setup prompt for a computer-use agent, with Storybook coverage. Share the prompt-button component with GitHub. - Update Slack tool documentation and runtime instructions. - Stabilize the mobile project browser test by waiting for the final canonical route before editing, preserving all persistence assertions. ## Verification - Live staging: invited the bot after the first mention. The channel became enabled and the bot answered that original mention. - Live staging: a normal Paperclip reply appeared in Slack with author attribution. The agent completed the calculation and replied once in the original thread and in Paperclip. - Live staging: a codeword entered in Slack was recalled from Paperclip. A following Slack calculation used the result from the Paperclip turn. Run metadata confirmed the same model session for both entry points. - Live staging: a scheduled routine used `slack_open_dm` and `slack_post_message` to deliver one DM. The existing app was reinstalled with `im:write`. The test routine was paused after verification. - Before the master merge: 397 focused feature tests passed. The continuity fix passed all 76 issue comment/update route tests and six focused route/integration cases. Typecheck, build, and token gates passed. - Review fixes: 47 focused integration cases passed, covering link revocation/replacement, private membership removal, guarded outbox dispatch, concurrent workers, lost scheduler responses, a real one-connection pool, exact reply provenance, and attachment retries. All 99 issue-comment route tests and the Slack catalog browser test passed. - Full local typecheck, production build, token gates, and module-boundary checks passed. The full local test command passed 25,915 tests before a 15-second timeout in `issue-thread-interaction-routes.test.ts`; that entire suite passed on isolated rerun (81 tests). Remaining serialized coverage is provided by the current-head CI shards. - An unchanged Cursor adapter test hit its 10-second limit in CI; all five tests in that file passed on a local rerun in 3.11 seconds, and the failed CI shard passed on its single retry. - The preview-server readiness test passed a local rerun (28 tests). The mobile-project readiness fix passed three repetitions of both browser tests (6/6). - Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates passed, including all eight browser shards, all server/chat suites, typecheck, build, runner checks, and security checks. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - Human messages on a Slack-linked task now publish to its Slack thread. The task banner states this behavior. Incoming Slack messages and internal agent bookkeeping must not echo back. - Normal tasks and routines can now use the assigned bot. Authority remains bound to the current responsible user's link; it does not fall back to the connection owner. Revocation, private-context limits, and queued-write checks still apply. - Existing Slack apps need `im:write` and a reinstall to open DMs. Other existing capabilities remain available without that scope. - New invited channels default to enabled. Explicit disabled choices remain disabled. Channels created by bot tools still require a person to enable responses. - No database migration or new provider credentials are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI, and browser tools. The runtime does not expose a more specific model build identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db8f8fe5b7 |
fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path > - Paperclip manages agent work through shared runner contracts. > - Direct protocol evals qualify provider behavior against a mock control plane. > - Grok supports API keys and company subscription credentials. > - The hosted protocol workflow selected an API key for every Grok cell. > - Product subscription support did not enable subscription protocol runs. > - This change adds explicit subscription selection and checks the recorded authentication mode. ## Linked Issues or Issue Description Refs #13878, #13882, #12618. The direct Grok protocol roster cannot run with subscription authentication through the trusted default-branch workflow. Add an explicit selector while keeping API-key dispatches compatible. Keep the actor allowlist, protected environment, immutable source revisions, and publication gates. ## What Changed - Add `grok_authentication` with `api_key` and `subscription` choices. Keep `api_key` as the compatibility default. - Deliver the protected `GROK_AUTH_JSON` secret only to a subscription-selected Grok cell. Do not provide an API key to that cell. - Read authentication mode from the pinned eval program's actual roster summary. Retain it in the cell, catalog, campaign roster, and result. - Reject missing or mismatched authentication evidence during aggregation. Preserve cell metadata and an allowlisted failure reason before failing a cell, so malformed evidence cannot hide the retained attempt. - Document credential setup, source separation, and temporary-secret cleanup. ## Verification - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed. - Validated all 39 Grok cells at eval revision `3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests only the subscription credential. Validation made zero provider calls. - Negative coverage rejects invalid selectors and missing or API authentication evidence in an otherwise passing subscription attempt. - `git diff --check` passed. All 53 current-head checks passed; the unchanged callback-drain timing test passed its bounded rerun, and the failed attempt is retained. Greptile reviewed `ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining findings. - No Docker or broad builds ran on the developer machine. CI performs repository checks on the configured fleet. ## Risks Grok runs require an eval revision that records `authenticationMode` in the roster summary. Missing evidence fails closed. The credential contains account access and refresh tokens; an owner must approve its delivery to the protected environment before a live run. The change adds no PR trigger or authorization bypass. Live subscription protocol qualification remains pending this workflow reaching master. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6448cd060 |
fix(runner): tolerate Codex account notifications during active turns (#13902)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Rust runner connects task runs to Codex app-server sessions. > - Codex sends account notifications on the connection during authentication updates. > - These notifications have no task, thread, or turn identity. > - The runner treated a valid account update as an invalid task event and stopped the run. > - This pull request classifies account updates as connection information while preserving task identity checks. > - Sandbox images can also contain an older runner with compatible metadata. The server must compare its bytes before reuse. > - Slack conversations use the deployed runner and can complete when Codex refreshes account state. ## Linked Issues or Issue Description Related: #13853. That change fixes the supported Codex version range. This fixes a separate notification failure after the version check succeeds. **What happened?** An active Codex task stopped with `thread_binding_mismatch` when the provider emitted `account/updated` without a thread ID. Slack showed that the agent stopped before completing its turn. **Expected behavior** Connection-level account updates must not stop a task or acquire task authority. Account details must not appear in task output. **Steps to reproduce** 1. Start a task with the native Codex runner. 2. Emit `account/updated` during the turn, with `authMode` and `planType` but no thread ID. 3. Observe that the old runner rejects the notification and fails the turn. **Paperclip version or commit** Reproduced after #13853. The fix is based on `b648d8cdd`. **Deployment mode** Cloud staging with a remote Codex runner and a Slack chat connection. ## What Changed - Classify `account/updated` and `account/login/completed` as connection information. - Reject account notifications that contain execution identity fields. - Reuse the existing bounded diagnostic path without publishing account payloads. - Test both notification types during two consecutive turns. Preserve existing identity rejection tests. - Compare preinstalled sandbox runner bytes with the controller artifact. Stage the deployed binary after a mismatch, failed checksum, or timeout. Keep exact retained artifact reuse. ## Verification - `cargo test --release --locked -p paperclip-runner-core`: passed, 582 test executions; two existing ignored cases. - `cargo test --release --locked -p paperclip-runner-core --test codex_provider`: 89 passed; two existing ignored cases. - `cargo fmt --all` and `git diff --check`: passed. - Repository-wide `pnpm -r typecheck`: passed. Server typecheck also passed after the artifact-selection change. - Server executor suite: 456 tests passed, including preinstalled digest match, mismatch, checksum failure, and timeout. - Repository-wide Vitest and build are running. - Deployed `5648d90d5dd762c5a8b697face269f6eea9cd59d` to one staging stack and verified its serving SHA. - Retried the failed conversation through the Slack UI. The agent replied and the task completed. - Sent a fresh Slack mention asking which bots belong to the channel. Verified successful `slack_members` and `slack_user` calls, a successful run, and the delivered Slack answer. - Sent a follow-up in the same Slack thread without another mention. The agent returned the requested names from context, with one final reply. ## Risks - Only two known connection-level notification methods change behavior. Unknown authoritative methods and malformed or conflicting identities still fail closed. - An older sandbox runner now requires one upload when its bytes differ from the controller artifact. Remote OS and architecture checks still apply. - No schema, credential, permission, Codex version-range, or UI changes. The Codex minimum remains 0.149.0. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code editing, terminal tools, and browser verification. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f55759942b |
fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b41ccf097f |
fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agents request connections through cards in task threads. > - MCP aggregators need provider-specific URLs and authentication. > - The task dialog used the generic setup form and omitted these fields. > - This pull request uses the same provider setup controller in tasks and Apps. > - Users can configure a connection without leaving the task. ## Linked Issues or Issue Description Related: #13755, #13855. **What happened?** An inline Executor request opened a very wide dialog with an empty credential step. Connect failed because the MCP URL was missing. The other aggregator cards also bypassed their provider setup. **Expected behavior** Each card shows its provider instructions, URL field, and authentication options in a bounded dialog. Completing setup grants access only to the requesting agent. **Steps to reproduce** 1. Enable experimental MCP aggregators. 2. Have an agent request Zapier, Arcade, Composio, or Executor from a task. 3. Open the card and continue past Access. ## What Changed - Route page and task setup through the same provider controller. - Bound the task dialog width and preserve the requesting agent's access scope. - Support existing accounts, saved drafts, URL/token setup, and task-bound OAuth. - Keep a sign-in link available when the browser cannot open a popup. Verify completion through the existing durable callback path. - Add inline Access, configuration, and narrow Storybooks for all four providers. - Document the shared setup requirement and correct Executor's URL instructions. ## Verification - Focused Vitest selection: 24 passed. Covers all four inline forms, requester access, existing accounts, saved drafts, OAuth retry, callback validation, popup cleanup, and generic reconnect endpoint preservation. - UI typecheck, UI build, design-token gates, and Storybook build passed. - Live local browser: new Executor, Arcade, and Composio connections completed provider consent from task cards. Each appeared Connected with the requester selected. - Real Test calls returned Executor output `4`, an Arcade public GitHub star count, and Composio tool-discovery results. An ungranted agent was denied access. Real Paperclip process-agent runs discovered each provider catalog with only the requester’s connection installed. - Deployed implementation commit `5d722e89c` to the isolated staging tenant. The original failing Executor card now completes, discovers seven actions, limits access to its requesting agent, and resumes that agent. Its continuation completed real Executor calls and the provider resume flow, then returned an upstream Airtable authorization link. A real staging Test call returned `4` in 1.5 seconds. Later PR commits add regression coverage and popup-unmount cleanup. - Zapier fresh-token browser test remains pending a provider clipboard handoff. Its URL and token flow passes focused tests. - Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook deployment and visual regression jobs skipped. Greptile 5/5, both review threads resolved. Canceled runners and unrelated chat timeouts passed the single retry on unchanged code. - The full local test suite was not run, as requested by the maintainer. CI runs the repository gates. ## Risks - OAuth popup behavior differs by browser. The explicit sign-in link and durable server completion checks provide recovery. - Saved task drafts store only a connection ID in browser storage. Credentials remain in the existing server vault. - No database or server protocol changes. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code execution, and browser tools. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
60c7c9cd1a |
fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E campaigns test the native runner in local and Daytona environments. > - Each campaign resolves one target lockfile and verifies its downloaded artifact. > - The Daytona image job did not pass that artifact digest to Docker. > - Docker used an older default digest and stopped before any selected task ran. > - This pull request passes and validates the campaign digest at the image build boundary. > - The image keeps its checksum check and frozen package installation. ## Linked Issues or Issue Description **What happened?** The merged-master [qualification campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409) stopped in the Daytona image build. The resolved target lock digest was `e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`. Docker used its default digest, `57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The checksum check rejected the mismatch. All 12 selected model cells were skipped. This follows the image provenance work in #13814. **Expected behavior** The image build must check the same lockfile artifact that the campaign restored and verified. A changed lockfile must still fail the checksum check. **Steps to reproduce** 1. Start a Product E2E campaign with a Daytona cell on master `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. 2. Resolve a target lockfile whose digest differs from the Dockerfile default. 3. Observe the provider-pack image stage reject the lockfile before model execution. **Paperclip version or commit** `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. **Deployment mode** GitHub Actions Product E2E campaign with a Daytona image build. **Install method** Built from source with the campaign lockfile artifact. **Agent adapter(s) involved** Native Codex and ACPX Claude cells were selected. No model cell ran in this failed campaign. **Database mode** Not involved. The failure occurs during image creation. **Access context** The authorized default-branch paid workflow. The build receives no provider credentials. ## What Changed - Read the image checksum from the existing target-lock job output. - Require a 64-character lowercase hexadecimal digest before image inspection or build. - Pass the digest as the existing Docker build argument. - Add regression checks and document the campaign checksum handoff. ## Verification - The Daytona image regression fails with the original workflow and passes with the fix. - All six Daytona image contract tests pass. - All 450 Product E2E unit tests pass. - Product E2E typecheck passes. - Actionlint passes for the changed workflow. - A context-shaped resolution probe preserves the downloaded lockfile bytes and digest. - All latest-head CI checks passed on `67d41fd9439b2a9a809ddb05765f8617585072c5` ([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)). - Greptile gave 5/5 on this head; its test-scoping comment is addressed and resolved. - A hosted Daytona image rebuild and the three remote qualification cells remain pending after merge. ## Risks The campaign digest comes from the existing trusted target-lock job. The restored artifact checks, Docker checksum check, frozen install, content identity, image signing, and verification remain in place. The standalone Docker default remains available. This change does not alter task behavior, prompts, credentials, or dependency versions. ## Model Used OpenAI `gpt-6-astra` through Codex performed diagnosis and review with code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation, implementation, and verification. Context window limits are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4721f55803 |
fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - MCP aggregator setup can return from provider consent to a saved draft. > - The branded setup did not explain failed or cancelled authorization. > - It also offered identity changes that the server does not apply when a saved connection resumes. > - This pull request explains OAuth return outcomes and keeps the displayed identity consistent with the saved policy. > - Users can understand the outcome and retry the same connection. ## Linked Issues or Issue Description Related: #13755 introduced the MCP aggregators. #13758 retired the legacy Composio broker. #13584 proposes changes to the provider handoff window; this fix retains the current handoff behavior. **What happened?** Cancelling Composio consent returned to the setup form with no explanation. The same controller ignored failed OAuth callback outcomes. Returning to Access on a saved draft also offered personal/shared choices, although the server retains the saved identity. This could make the OAuth request disagree with that identity. **Expected behavior** Explain cancellation or failure, preserve the draft, and offer Try again. Display the retained credential identity and start OAuth with the policy returned by the server. **Steps to reproduce** 1. Enable MCP aggregators and start a Composio connection. 2. Continue to provider consent and cancel it. 3. Observe the return screen. Before this change, it showed the form without cancellation feedback. 4. Go back to Access. Before this change, the form offered personal/shared choices even though resume retains the original identity. **Paperclip version or commit** Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`. The original report of successful consent leaving setup unfinished did not reproduce. This PR addresses the recovery defects observed during that investigation. **Deployment mode** Isolated local development instance and authenticated staging deployment, with real Composio consent and provider calls. ## What Changed - Read the OAuth callback outcome in branded MCP setup and show cancellation or failure feedback. - Retry the same saved draft without displaying untrusted callback error text. - Keep saved personal/shared identity fixed in Access and select OAuth identity from the returned credential policy. - Add six focused regression cases and authorization-failure Storybook states for Arcade, Composio, and Executor. - Document return-screen recovery and retained identity in the connector playbook. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth return' --maxWorkers=1`: 6 passed, 143 skipped, including after integration with current master. - `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed for the implementation commit. - Browser recovery: cancelled real Composio consent, observed the new feedback, returned to Access, retried the same personal draft, completed consent, and ran a real tool call. - Staging on implementation commit `84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal connections each completed on the first consent attempt and loaded 11 tools. Real discovery, execution, and schema calls succeeded. Both connections stayed Connected after reload. A real agent used the shared connection through the Paperclip gateway and returned the public repository documentation hierarchy with one success and zero errors. - All current-head CI gates passed on `662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and agent-chat browser shards each had an initial timeout; both passed on one targeted rerun without code changes. Greptile reviewed this exact head at 5/5, with no open review threads. - No full local suite was run, as requested. CI provides the broader checks. The PR adds a master merge and documentation after the live-tested implementation commit. ## Risks - This shared setup controller also serves Arcade and Executor. Their callback rendering and retry behavior have focused test coverage; this investigation used Composio for live provider testing. - Saved identity remains fixed during resume. A different identity requires a new connection, consistent with server behavior. - No database, protocol, credential storage, or gateway policy changes. ## Model Used OpenAI GPT-6 via Codex, with code editing, shell tools, and browser testing. The exact runtime model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
106d89314f |
fix(evals): keep raw diagnostics out of public history (#13850)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E campaigns retain evidence from live provider runs. > - The public publisher copied per-attempt files based on their extension. > - Credential redaction does not remove hidden reasoning or provider session IDs. > - This pull request keeps those raw diagnostics out of public evidence bundles. > - Public reports retain normalized grades and declared fixture screenshots. ## Linked Issues or Issue Description **What happened?** The Product history publisher accepted all JSON, log, Markdown, and text files under an attempt directory. API snapshots and process logs can contain provider reasoning even after credential redaction. **Expected behavior** Keep raw attempt diagnostics in retained Actions artifacts. Publish normalized results and declared fixture screenshots. Preserve the original grades and attempt evidence. **Steps to reproduce** Place a credential-redacted API snapshot with a reasoning event under an attempt's snapshots directory. The previous public path check admitted it because it ended in .json. ## What Changed - Remove extension-based admission for per-attempt text files in both S3 and Pages staging. - Require PNG paths to appear in the existing screenshot declaration allowlist. - Preserve root normalized results, grading, billing, provenance, and reviewed screenshot behavior. - Add seven denial cases for raw diagnostics, renamed files, malformed JSON, and false screenshot declarations. - Update report copy and the publication security contract. ## Verification - Focused history and report tests: 47 passed. - Replayed the filter against a copy of retained Grok evidence: six diagnostic files removed, one declared screenshot retained, zero provider calls. Original artifacts remain unchanged. - git diff --check passed. - Product typecheck in the reused local dependency tree reports an unrelated plugin type mismatch for organizationSwitcher in server/src/services/plugin-capability-validator.ts. CI will verify a fresh install. - No local browser or Docker execution. ## Risks Public reports no longer link raw per-attempt logs or API snapshots. Those files remain in the original Actions artifact. This change affects future publication only; it does not withdraw or rewrite existing immutable public campaigns. Normalized result content and marked screenshot review retain their existing boundaries. No workflow authorization, provider credentials, application behavior, or database changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ac854a9b59 |
ci: build Daytona eval images on the authorized fleet (#13847)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E campaigns test the browser, server, and runner together. > - The trusted workflow selects an authorized execution fleet. > - Daytona image builds still use a fixed GitHub-hosted runner. > - This pull request applies the existing fleet selection to image builds. > - Campaign builds and tests then use the configured EC2 fleet. ## Linked Issues or Issue Description Refs: https://github.com/paperclipai/paperclip/pull/13845 **What existing behavior does this improve?** The location of Daytona Product E2E image builds. **Current behavior** The image job uses ubuntu-latest even when RUNNER_E2E_AWS_ENABLED selects EC2 for the rest of the campaign. **Proposed behavior** Use the existing authorized runner output for the image job. Preserve the GitHub-hosted fallback when the EC2 switch is off. ## What Changed - Route Daytona image builds through the existing authorized runner selection. - Assert that the image job depends on authorization and receives no provider credentials. - Document the build routing and credential boundary. ## Verification - Product E2E workflow security: 11 tests passed. - git diff --check passed. - Reviewed actor checks, workflow triggers, target commit selection, signing, and package permissions. They are unchanged. - Required CI checks pass on the latest head. The first CI attempt had failures in unchanged chat tests; one diagnostic rerun passed. Both attempts remain in Actions. - Greptile reviewed the latest head at 5/5 with no inline findings. - Live EC2 image verification follows after this trusted workflow change is merged. No local Docker execution. ## Risks EC2 image builds depend on the fleet having working Docker and sufficient disk space. Existing authorization, image digest checks, signing, and the GitHub-hosted fallback remain in place. No provider secrets are added to the image job. No application or database changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
950ccb8eef |
ci: enable Grok qualification in the trusted paid workflow (#13845)
## Thinking Path > - Paperclip manages AI agents and their work. > - Product E2E tests verify tasks through the browser, server, and runner. > - These paid tests use a trusted workflow from master and an isolated target commit. > - The Grok target branch selects XAI_API_KEY, but the trusted workflow does not deliver that credential. > - Its artifact test also needs the pinned Python verifier before execution. > - This pull request adds both bindings inside the existing paid boundary. > - The tests can then run on the configured EC2 fleet without laptop Docker. ## Linked Issues or Issue Description Related evaluation infrastructure: https://github.com/paperclipai/paperclip/pull/11297. No duplicate Grok paid-workflow change was found. **What existing behavior does this improve?** Branch-targeted Grok Product E2E qualification in the existing paid workflow. **Current behavior** Grok cells cannot receive their selected API credential. The Grok build-revise case also misses the artifact-verifier setup step. **Proposed behavior** Deliver XAI_API_KEY only when the selected matrix credential is XAI_API_KEY. Prepare the existing pinned verifier for the Grok qualification suite. **Reason and benefit** Run the controller, browser, runner and artifact checks on the EC2 fleet. Preserve default-branch workflow authorization and protected environment secret access. ## What Changed - Bind the selected XAI credential only in the paid test step. - Install the checksum-verified Grok binary for local cells before provider access. - Include Grok qualification in the existing pinned artifact-verifier preparation. - Add an optional max_parallel input that can only lower the configured campaign concurrency. Use 1 for the Grok test key. - Extend security assertions and document setup. ## Verification - Ran the Product E2E workflow-security tests: 11 passed. - Checked the diff for whitespace errors. - Reviewed credential selection, setup ordering, numeric actor gates, target commit pinning, and trusted report checkout. - Live Grok execution follows after this workflow is available on master. This PR does not claim completed Grok qualification. ## Risks The paid test step can use the selected XAI credential and incur provider charges. The credential remains in runner-e2e-paid and is absent from setup, build, and reporting jobs. The default-branch gate and existing environment restrictions remain in place. No database migration or product behavior changes. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool execution. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f447990d77 |
fix(claude): recognize ACP quota fallback errors (#13831)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Claude adapter reports failures to the recovery system. > - Recovery must distinguish account quota from other provider limits. > - The Claude ACP bridge has a default message for an account with no quota. > - The current classifier misses that message and returns `acpx_turn_failed`. > - This pull request recognizes that exact typed message and uses the existing quota wait. ## Linked Issues or Issue Description Refs #13651. This is a narrow follow-up to its typed quota classification. Related PRs #13549 and #10276 address broader quota classification and reset handling. This change covers the bridge's exact default quota message. **What happened?** A typed Claude ACP `limit` failure with the title `The Claude account has no available quota.` returns `acpx_turn_failed`. The recovery result has no quota label. **Expected behavior** Return `provider_quota`. Use the existing one-hour quota backoff when the provider gives no reset time. Keep context, turn, rate, and configured budget limits out of the quota path. **Steps to reproduce** 1. Use the local ACP fixture to return the exact title above with category `limit` and severity `error`. 2. Run the Claude adapter in oneshot or persistent mode. 3. Before this change, the new regression cases receive `acpx_turn_failed` instead of `provider_quota` on ACPX 0.12.0 and 0.13.1. **Paperclip version or commit** Reproduced on `a959e4750`. Rebased onto current `master` before submission. **Deployment mode** Local source checkout with isolated ACP child-process fixtures. No live provider calls or customer-stack changes. ## What Changed - Recognize the exact Claude bridge quota fallback only for typed `limit` failures. - Add real-process regression cases for both pinned ACPX versions and both session modes. - Verify that unrelated categories, positive quota wording, and historical generic limit errors do not imply quota exhaustion. - Document the fallback and its existing recovery backoff. ## Verification - Red/green reproduction: all four new fallback cases failed before the classifier change and passed afterward. - 108 targeted tests passed across Claude ACP, quota, parser, and server recovery suites. - Claude adapter typecheck and build passed. - Real-process tests verify that provider text stays out of results and logs. - Full local `pnpm -r typecheck` and `pnpm build` passed. - The full local `pnpm test:run` attempt stopped after embedded PostgreSQL could not load a missing library symlink. The dependency setup was repaired in the worktree. The isolated database test then passed. The complete test suites passed in GitHub CI. - All 54 latest-head checks passed on `3d4d4e65386c5b6023ba34e6a2abcd94cff9a9bc`. Two optional Storybook jobs were skipped. - Greptile: 5/5, with no inline review threads or requested changes. ## Risks - Low risk. An exact match is required inside an existing typed `limit` failure. - If upstream changes this wording, this fallback can stop matching. Existing quota-message detection remains in place. - No schema, credential, or permission changes. Historical generic errors remain ambiguous and are not reclassified. ## Model Used OpenAI Codex, GPT-6. The session identifies the model family as GPT-6 but does not expose a more specific model ID or context-window size. Used reasoning, repository inspection, code editing, and terminal test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a10702a878 |
feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people start and continue agent tasks from other services. > - A Slack conversation needs access to its surrounding discussion and Slack collaboration tools. > - The agent must use the linked requester's access and keep private material within its permitted audience. > - This pull request adds Slack tools through the existing connector contribution and approval framework. > - People can ask an invited bot to read a discussion, create follow-up tasks, and collaborate in Slack. ## Linked Issues or Issue Description **Subsystem affected** Chat connectors, connector runtime, tool gateway, and connection Settings/Access. **Problem or motivation** Slack-origin tasks can receive messages but cannot inspect the rest of a channel or act through the originating bot. People must paste context or configure a separate integration. **Proposed solution** Supply typed Slack tools and a bundled skill only to the originating task and assigned agent. Resolve the linked requester on the server. Check bot and requester access before reads and writes. Use existing durable actions and approvals. Retrieved messages remain source material. **Alternatives considered** Slack's user-OAuth MCP server does not replace the customer-created chat bot. An unrestricted Web API proxy would not provide suitable permission or publication boundaries. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps work with a provider contribution. It does not add a task dispatcher or a separate Slack task lifecycle. Related: #11144 covers generic per-user MCP grant execution; this change binds Slack bot operations to chat-origin tasks. ## What Changed - Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and shared native/HTTP execution. - Bind tools to company, endpoint, task, run, assigned agent, and admitted linked requester. Check membership and revocation on each call and before queued writes. - Add paginated reads, bounded history search, source links, messages, file uploads, reactions, pins, bookmarks, topics, canvases, lists, and approved channel operations. - Restrict private-source publication, including automatic replies and uploaded deliverables. Keep other people's bot DMs inaccessible. - Reuse action receipts, idempotency, approvals, and reconciliation. Suppress an identical explicit-send/final-reply duplicate. Return governed results through their verified originating conversation. - Add endpoint-bound personal search OAuth storage and lifecycle. Keep native real-time search disabled until a runtime meets Slack's transient-result requirements. Current runtimes use bounded history search. - Show capabilities, scope upgrades, and personal search authorization in Settings/Access and Storybook. Document provider and runtime limits. ## Verification - Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no unresolved findings. One unchanged rapid-callback timing test passed on a single CI retry. - Approval presentation regressions cover board-comment precedence and exact Slack publication; the expanded database assertion passed in CI. The local PostgreSQL startup probe later became unavailable, so that final assertion was verified in CI. Slack setup and failed-run retry browser tests also passed locally. - Full workspace typecheck and build passed. Server typecheck/build passed again after the approval routing fix. - Broad local suites passed in separate groups: server 12,958 tests, UI 6,555, shared 770, skills catalog 20, and other workspace packages 2,652. CLI and serialized server checks passed after environment/timeout retries. These are composite results, not one uninterrupted green full-suite invocation. - PostgreSQL authority regression covers admitted identity, cross-company/task/agent rejection, recovery, retained-session revocation, OAuth refresh/disconnect races, approval execution, exact publication lineage, retries, uncertain sends, and duplicate suppression. - Gateway/response regressions cover separate-origin approval batches and durable continuation. Focused provider, access, search, native runtime, route, and AgentMail regressions pass. - Storybook capability, missing-scope, OAuth configuration, authorization, and disconnect states were inspected in the browser. - Live staging: read a channel decision and full thread, create exactly two assigned backlog tasks, add a reaction, paginate discovery to exhaustion, and return bounded search matches with source links and coverage. - Live staging: create/edit/read a canvas and list, inspect the canvas in Slack, post/edit one message, and create a channel only after approval. New channels remain disabled for responses. - Live staging: read a response-disabled channel from the requester's DM; writes to that channel were denied. The test setting was restored. - Final live retest passed: explicit file upload and exact content read-back; approved deletion of only the disposable bot message; continuation confirmation returned to the original Slack thread without repeating the action. - Optional OAuth, private multi-user boundaries, native RTS, and CLI provider execution are not fully live-qualified. The staging agent initially supplied malformed tool arguments; valid arguments succeeded, and the tool/skill descriptions now emphasize UUID write keys. ## Risks - Existing Slack apps must add scopes and reinstall for new capabilities. Provider plans and document permissions can still restrict operations. - Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET` for governed tool actions. The staging instance was configured with explicit operator approval; fleet provisioning is a separate gap. - Native RTS is not exposed on current transcript-retaining runtimes. Bounded history scans are deliberately reported as incomplete. Inline file reads support text/canvas content up to 256 KiB; other types return metadata. - Private document edits fail closed when the full audience cannot be verified. Uncertain effects other than posts/uploads require inspection instead of blind retries. - Shared approval-delivery code now separates outcomes by source run to preserve origin boundaries. No database migration is required. - A separate completion-validator gap remains when the agent cites a prior run's registered artifact during finalization. It asked for registration again even though Slack delivery was confirmed. This change does not add a connector-specific task-completion policy. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and browser testing. The exact deployed model identifier and context-window size were not exposed in the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a959e47508 |
fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps give agents controlled access to external services. > - Google Chat uses OAuth profiles and reviewed MCP tools. > - Those profiles request membership and read-state access that the supported feature set does not need. > - Removing read-state access also requires us to block unread search filters, including on existing connections. > - This pull request reduces both OAuth methods and enforces the reduced search contract before dispatch. > - Users retain conversation lookup, message history, ordinary search, and approved message sending. ## Linked Issues or Issue Description **What happened?** Both Google Chat profiles request membership and read-state scopes. The supported tool set does not include membership listing or read-state updates. Message search still advertises an unread filter. Related work: Refs #12619. **Expected behavior** Managed and customer-owned OAuth request only the scopes needed for supported features. Unsupported unread filters fail clearly before any provider call. Existing cached catalogs and broader grants must not bypass that policy. **Steps to reproduce** Start Google Chat OAuth from either connection method and inspect the requested scopes. Inspect the message-search tool schema, then submit a search with `searchParameters.isUnread` set to true or false. **Paperclip version or commit** The scope change is based on master at `110d176fc`. **Deployment mode** Managed Cloud and self-hosted instances with Google Chat Apps enabled. ## What Changed - Remove `chat.memberships.readonly` and `chat.users.readstate.readonly` from shared profiles and all four Chat connection methods. - Hide unsupported read-state fields and instructions in agent and board Test tool schemas. - Reject explicit unread filters, including false, null, snake-case fields, and encoded filter objects, before provider dispatch. - Recheck previously approved calls and support existing profile-bound and URL-only Chat connections. - Add scope, signed broker request, OAuth URL, allowlist, schema, and dispatch regression tests. - Document coordinated app/broker rollout, existing-grant reconnects, and the remaining deployment checks. ## Verification - All seven focused OAuth and Chat gateway test files pass: 498 tests on the rebased branch. - `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4. - The complete sharded CI test matrix passes, including general server, Chat, workspace, serialized server, Runner, and all eight browser e2e shards. The duplicate unsharded local `pnpm test:run` was stopped after CI passed; it did not complete locally. - `git diff --check` passes. - `node scripts/ingest-app-definitions.mjs` succeeds and leaves the branch unchanged. Google Workspace JSON is the durable reviewed input used by the generator. - All current-head CI checks pass at `7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build, and canary dry run. Greptile is 5/5 with zero unresolved threads. The generator concern was withdrawn after review of the source and regeneration evidence. - No production deployment or live Google consent test was performed. After coordinated deployment, verify reduced consent scopes, normal search/history, message sending, and rejection of unread filters. ## Risks - Coordinate deployment with the companion Cloud broker scope change. Mixed versions can reject exact-scope requests. - Existing tokens are not narrowed or revoked. Grants with old scopes need new consent. Do not revoke a shared Google client to migrate one profile. - Explicit unread filters now return an error instead of being sent to Google. Ordinary search and the approved send tool remain available. - No database, UI, lockfile, or workflow changes. ## Model Used OpenAI Codex, a GPT-5-based coding agent, with tool use and code execution. The exact runtime model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9a76390cfa |
test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.
## Linked Issues or Issue Description
Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.
**What happened?**
Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.
**Expected behavior**
Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.
**Steps to reproduce**
Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.
**Paperclip version or commit**
Based on master at
|
||
|
|
1ccae464c5 |
feat(ui): add task artifact media gallery and full-row links (#13825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks collect the files and work products that agents create. > - The Artifacts tab shows these outputs in rows, which makes videos hard to compare. > - Small text links also make artifact rows harder to open. > - This pull request adds image and video tiles with previews and makes the full artifact row clickable. > - Users can compare outputs and open the existing media viewer with one click. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task Artifacts tab and media previews in task chat. **Current behavior** Video outputs appear as file rows or icons. Users must click a small link to open a work product. A task with eight video outputs gives little visual context. **Proposed behavior** Show images and videos in a responsive gallery. Show a paused video frame as the thumbnail. Keep documents, links, and other files in rows whose entire area opens the item. **Reason and benefit** Users can compare generated media without opening each item. Larger click targets also make the sidebar easier to use. **Breaking changes** None. This uses the existing artifact URLs, media viewer, run grouping, and attachment filters. No API or database changes. Related work: #11226 added the task sidebar output surface, and #7361 added rich attachment previews. #3524 concerns a separate reviewed-assets panel. This PR improves the existing task artifact components. The duplicate search found no active PR for this change. This is polish for the shipped Artifacts & Work Products roadmap item. ## What Changed - Add a shared media tile for work products and agent attachments. - Reuse video and image previews in task artifacts and chat. Seek up to one second into videos and reset preview state when the source changes. - Use the existing task gallery for playback and downloads. Preserve grouping and attachment deduplication. - Extend native links and buttons across work-product rows, including keyboard focus indicators. - Add eight offline Storybook examples for video outputs, mixed media, clickable rows, narrow and wide panels, missing previews, empty state, and light mode. - Register the component in the design guide and document its use. ## Verification - 108 focused component tests pass, including thumbnail seeking, source changes, gallery activation, and attachment deduplication. - `pnpm build`, `pnpm -r typecheck`, Storybook build, token gates, and `git diff --check` pass locally. - Reviewed the production components in the embedded browser. Checked all eight video thumbnails, mixed media, narrow layout, light mode, blank-area row clicks, keyboard gallery activation for generic-MIME images, and playback from chat video thumbnails. - Storybook: open **Tasks / Artifact Gallery** and select **Eight Video Outputs**, **Mixed Media And Files**, or **Whole Row Clickable**. The small local clips are synthetic fixtures. - All build, typecheck, unit, runner, and end-to-end CI jobs pass for `b8ace289b5e07df5b9f2c319b159f3923ce427b4`. Greptile gives 5/5 with both review findings resolved. All 54 PR checks pass, including the external security scan. The full test suite passed in CI. The duplicate serial local test run was stopped after CI finished; the 108 focused tests, full build, and recursive typecheck passed locally. ## Risks - Video thumbnails require the browser to load metadata and a frame. A slow server or unsupported codec can leave the fallback visible; opening and downloading still use the existing viewer. - Full-row click targets change pointer interaction with work-product cards. Native link and button semantics remain in place. ## Model Used OpenAI GPT-6 in Codex, with code execution and browser tools. The exact deployment ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a68f3d8e35 |
fix(runner-e2e): align Daytona image and provider pack provenance (#13814)
## Thinking Path > - Paperclip is an open source app people use to manage AI agents for work. > - The Daytona runner image provides the native runner and its provider package. > - The controller also sends a provider package to native Daytona cells when the image package does not match. > - A stale image and a package from another source revision caused a 1.8 GB upload before Claude could run. > - This pull request refreshes the reviewed lock checksum and documents how to reuse the exact package from an immutable image. > - The benefit is a reproducible setup path and clear evidence when image and package provenance do not match. ## Linked Issues or Issue Description **What happened?** A Daytona native Claude run used image source revision `45c99a0d06cbd5b04982b06b79de321149930ac5` with a controller provider package from revision `294853dc...`. Runtime verification rejected the image package and staged a large package upload before model execution. A hosted campaign also failed during image setup because the Dockerfile expected lock checksum `d7d96cf0...` while the resolved lockfile checksum was `4b796c312833ebf2be4c38228babc0292c76774fb40d43bd18b54bd6b205753d`. Related public work reviewed: [#12795](https://github.com/paperclipai/paperclip/pull/12795), [#12862](https://github.com/paperclipai/paperclip/pull/12862), and [#12887](https://github.com/paperclipai/paperclip/pull/12887). **Expected behavior** The Daytona image and controller provider package must come from the same verified build. A package extracted from the immutable image must pass the existing manifest, source revision, lockfile, binary, bridge, and artifact checks before a native Claude run starts. **Steps to reproduce** 1. Set `PAPERCLIP_E2E_DAYTONA_IMAGE` to the old immutable image digest. 2. Set `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH` to a package built from a different source revision. 3. Run a native Claude Daytona cell. 4. Observe provider package verification failure followed by the large staging upload. 5. Build the image with the stale Dockerfile lock checksum and observe the checksum failure. **Paperclip version or commit** `3b8df3dcd6f99e99277faa45f3351989edb5c239`. **Deployment mode** Daytona native runner E2E. **Install method** Built from source. **Agent adapter(s) involved** Claude Code through the native ACPX runner. **Database mode** Not database-related. **Access context** Not applicable to the setup failure. ## What Changed - Refreshed `PAPERCLIP_RUNNER_LOCK_SHA256` in `docker/daytona-runner/Dockerfile` to the resolved lockfile checksum. - Added a local guide for extracting the provider package from an immutable verified image with Docker. - Documented the required provenance checks and the expected manifest-matched runtime log. - Documented that cold package upload coverage must remain separate from recovery coverage. - Pinned the preview-service test guest to the test runner’s Node executable and logged guest startup and bound ports for readiness diagnostics. ## Verification - 442 focused E2E tests passed. - Typecheck passed. - `pnpm build` passed. - Five provider-pack reuse tests passed. - Six Daytona image contract tests passed. - Fixed Daytona campaign [35740613581](https://github.com/paperclipai/paperclip/actions/runs/35740613581) passed. - The overall hiring campaign [35739993219](https://github.com/paperclipai/paperclip/actions/runs/35739993219) failed because of an unrelated Mini metadata failure; its Claude cell passed. - Full local `pnpm test:run` was attempted but did not complete. The isolated Postgres install was repaired and its 15-test probe passed. - After the fixture change, all seven preview reservation tests passed locally and CI server shard 6/12 passed on `20234f75f`. This removes login-shell Node resolution variance; the exact cause of the earlier CI-only timeout is not established. - CI run [35766034635](https://github.com/paperclipai/paperclip/actions/runs/35766034635) passed on `20234f75f`. All server, runner, browser, build, and typecheck gates passed. - The signoff browser case initially failed waiting for an approver run. All five signoff tests passed locally without changes; the one allowed CI retry passed all 19 shard tests. This is recorded as an intermittent failure, not a demonstrated product fix. - Greptile reviewed `20234f75f`: 5/5, no actionable findings. - [Follow-up report](https://pages.paperclip.ing/runner-daytona-hiring-20260922/) includes timings, evidence links, and the remaining hiring configuration failure. - Review the immutable image source revision and extracted `provider-pack.json` before another paid recovery run. ## Risks - The Dockerfile checksum gate intentionally fails when the resolved lockfile changes. A future dependency change must refresh the reviewed checksum with the image change. - The local extraction guide requires Docker and a pullable immutable image. - An image built from an older source revision can still fail runtime manifest verification. The guide does not bypass that check. - The change does not alter runner prompts, approval policy, or recovery behavior. ## Model Used OpenAI Codex using the primary GPT-6 backend; the exact backend deployment ID is not exposed. Repository analysis and code execution used tool access. Assistance also came from OpenAI gpt-5.6-luna. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2788f20fc0 |
fix(ui): show ancestors in the task detail Tasks panel (#13823)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks form a hierarchy that explains why each piece of work exists. > - The streamlined Tasks panel shows child tasks and work created by the current task. > - It does not show ancestors, although the task response already includes them. > - This pull request adds linked ancestors above Subtasks in root-to-parent order. > - Users can now move up the task hierarchy from the same panel. ## Linked Issues or Issue Description Related implementation: #13241. A search found no duplicate ancestor-panel PR. **What happened?** The Tasks panel omitted the current task's ancestors. The streamlined header also hides hierarchy breadcrumbs. **Expected behavior** The Tasks panel should show the ancestor chain and let users open each ancestor. **Steps to reproduce** 1. Open a task with a parent and grandparent in the streamlined UI. 2. Open the Tasks tab in the side panel. 3. Observe that it shows child and created tasks but no ancestors. **Paperclip version or commit** Reproduced on master at |
||
|
|
3d78e3a4ec |
fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path > - Paperclip manages AI agents and their work. > - The native Runner keeps a live provider process between task turns. > - Managed GitHub access used a token tied to one run. > - A new run forced Paperclip to replace that process to replace its token. > - This PR gives the session a stable credential transport and binds each operation to the active run. > - The agent can keep its process while Paperclip checks current identity and grants. ## Linked Issues or Issue Description Follow-up to #13738. Related credential-rotation work: #11770 and #8208 use process replacement for other adapter credentials; this change applies to managed GitHub access in the native Runner. **What happened?** A configured GitHub connection forced a warm native provider process to close at each new run. The saved conversation survived, but the live process did not. **Expected behavior** Keep the warm provider process. Resolve GitHub access for the current run when each command starts. Deny access while idle or after the run ends. **Steps to reproduce** 1. Configure managed GitHub access for a native Runner agent with a warm session. 2. Complete a turn, then send another message to the same task. 3. Observe the provider process close with the reason `warm native session configuration changed`. **Paperclip version or commit** Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before submission. **Deployment mode** Local and remote native execution, including the sandbox callback bridge. ## What Changed - Move configured native GitHub transport and launcher ownership from the run to the provider session. - Bind the broker only after the executor acquires session ownership. Clear that binding when the run exits. - Keep the shared live-run, identity, grant, and trust-policy checks for each credential request. - Reject wrong scopes, idle requests, and credential responses that arrive after their run binding changes. - Retire transport and launcher files with the provider session. Keep anonymous commands available if bridge startup fails. - Add red/green executor tests, real subprocess and callback-bridge tests, and database checks. Update the runtime documentation. ## Verification - Before the fix, both new local and remote warm-session reuse tests failed. - After the fix, 435 targeted tests passed across the executor, broker, launcher, token, and database suites. - A real long-lived test process kept the same PID and original environment across two runs, including through the production callback bridge on local test processes. - Server typecheck and TypeScript compilation passed. - Full workspace typecheck and build passed. Server typecheck passed again after the review fix. - The fallback-logging regression failed before the fix; all 9 broker tests pass afterward. - The exact chat sidebar browser scenario passed locally. The initial CI timeout showed failed Vite module downloads; all eight browser shards pass on the latest commit. - All 53 latest-head checks passed, including the full CI test matrix and security checks (two unrelated conditional checks skipped). - The duplicate full local test run was stopped after CI passed; it is not claimed as a completed local pass. Targeted local tests, workspace typecheck/build, and the browser scenario passed. - Greptile reviewed the latest commit at 5/5 with no unresolved findings. - No fresh paid provider or Daytona campaign has run for this change. ## Risks - The broker now lives as long as the provider session. Tests cover idle denial, late cleanup, late responses, shutdown, and failed startup. - Its in-memory authority does not survive a controller restart. Existing checkpoint and process-recovery rules still apply. - Raw GitHub credentials remain confined to individual command processes. The session transport token cannot select a different task, agent, company, or run. - No database migration or public API change. ## Model Used OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
74a9730acb |
fix: continue native agent chats after worker loss (#13813)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses native workers to run Claude and Codex conversations. > - A worker crash leaves a cleanup hold because its provider did not acknowledge suspension. > - A new user message must not reuse that unverified session or repeat old tool calls. > - The existing continuation path can preserve history and start a fresh session, but local cleanup ownership remained held. > - This pull request verifies the stopped local owners and releases only their cleanup hold for a new user turn. ## Linked Issues or Issue Description Refs #13775. **What happened?** After a native worker crashed, both providers retained cleanup quarantine. A saved plan survived, but the conversation could not produce another answer. **Expected behavior** Once the old worker and provider process groups have stopped, a new user message can continue in a fresh session with the saved work and prior action history. **Steps to reproduce** Run the opt-in `agent-chat-qualification` suite with case `worker-crash-retry` on native Codex and native Claude. The fixture saves a plan, kills the exact worker through a Linux pidfd, releases a local read-only brief, and sends a new message. ## What Changed - Verify the exact local worker stop receipt, provider identity receipts, released leases, and retained state before retiring a native cleanup hold. - Recheck process liveness and state before admission. Keep the old run and durable session files intact. - Use the existing explicit conversation continuation path. Generic Retry remains blocked for cleanup quarantine, including on the old failed-run marker after a successful continuation. - Extend the live oracle to require a successful fresh session, correct predecessor context, unchanged plan, one original message, and one answer containing a reference introduced after the crash. - Add physical-proof and database-backed admission tests. Document the precise qualification scope. ## Verification - Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database, and public APIs. - Core recovery proof source: `3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8: `9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`. - [Core recovery campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990) · [Public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/). - Both cells verify the real crash boundary, blocked generic Retry, unchanged saved plan, one original prompt, one fresh successor with predecessor context, and one run-attributed answer containing the post-crash reference. Billing coverage is partial because the crashed runs did not report complete usage; missing cost is not zero cost. - [Master baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443): both providers stopped at quarantine, with cleanup passing. - [Follow-up baseline without the fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866): Codex produced a complete red result. Claude reached the same error, but its artifact upload was canceled. - [Complete Claude baseline with the same version-8 definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177): red at cleanup quarantine, cleanup passed, source `6479a90c5754044356b39a9278b9a3e92ce8e55e`. - Earlier candidate attempts remain retained: [first](https://github.com/paperclipai/paperclip/actions/runs/35738449214) passed Codex and found a Claude fixture wait race; [second](https://github.com/paperclipai/paperclip/actions/runs/35740408065) exposed the normalized session-open receipt mismatch. Both corrections are in the final source. - Targeted server suites: 570 passed before the final two additional receipt regression cases. The physical-proof suite, including those cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and Product E2E typechecks passed. - **All 54 PR checks passed on final head `db6f775df`**, including repository typecheck, build, tests, browser shards, and canary dry run. The canary job required one retry after its runner received a shutdown signal. On the earlier core proof head, two timing-sensitive tests passed in isolation and on a single CI retry. - Final UI regression checks: 5 passed; UI typecheck and token gates passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed. - Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup passed**, including the browser assertion that the quarantined historical run never regains Try again. [Final campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013), source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash `bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`. [Final public evidence report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/). Final screenshots and retained state inspected for both providers; cost coverage remains partial. ## Risks - Missing or conflicting stop evidence keeps the conversation blocked. This change does not kill an unverified process. - This qualifies new user input after local worker loss. It does not enable automatic replay, exact-session recovery, remote crash recovery, or native onboarding defaults. - Old action outcomes remain part of the continuation. A process exit is not proof that an action did not happen. - No schema migration or production prompt change. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used reasoning, repository inspection, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8725d6ce09 |
fix: make answered Slack conversations idle (#13809)
Settle published, successful Slack turns as Idle; resume the same conversation on an admitted message. Preserve unfinished work, delivery errors, and explicit dispositions. Verified through focused lifecycle/API/UI tests, full CI, and a real staging Slack conversation in the embedded browser. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5842185e4f |
fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat needs reliable native execution before native runners become the onboarding default. > - Existing stories covered idle reassignment and controller restart, but not an executing worker handoff or worker process loss. > - Status answer tests also need to reject stale claims and invented facts. > - This pull request adds six opt-in full-stack cells with independent state assertions and retained evidence. > - The probes exposed a misleading Retry across server projection and recovery-banner paths; the fix reports the blocked recovery honestly. > - The tests preserve failures without changing recovery policy, production prompts, or onboarding defaults. ## Linked Issues or Issue Description Refs: #13762. Related: #13765 (Retry targets the latest failed attempt), #13753 (task context ownership), #13746 (native recovery work). ## What Changed - Add active reassignment with saved draft and plan preservation, old-worker cancellation, and successor completion checks. - Preserve recovery-needed projection when native cleanup fails before its coordinator exists, refuse a generic retry that would immediately fail again, and replace the recovery banner's misleading Retry with Inspect run. - Add verified local worker process loss with a required successful continuation; retain a failing qualification result when recovery is unavailable, while independently verifying the UI/API refuse doomed retries. - Add two-turn factual answer checks for current blockers, stale claims, inactive backlog work, and unknown facts. Retain prose for separate semantic review. - Add positive and negative oracle calibration and document fault isolation, cleanup, billing, and qualification limits. ## Verification - Eval TypeScript check passes. - All 442 eval support tests pass locally. The 89 focused server tests and server typecheck pass. Six recovery-banner UI tests and token gates pass. - Initial new-cell campaign: https://github.com/paperclipai/paperclip/actions/runs/35657128077. All six results are retained; four failed on fixture-contract issues and two exposed real worker cleanup quarantine. - All 26 existing native onboarding cells: https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26 passed on master |
||
|
|
846336e5a0 |
test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path > - Paperclip lets people manage agents through ongoing conversations. > - Chat users can change instructions while a provider is already working. > - Existing chat evals wait for each turn to settle before the next message. > - They cannot prove delivery during active work or the saved effect of a correction. > - Existing fixtures also enable Agent Chat through the API rather than the settings UI. > - This PR adds bounded browser workflows and checks their persisted outcomes. ## Linked Issues or Issue Description Refs #13741, #13752, #13750. **What happened?** The chat suites cover planning, delegation, status, and recovery. They lack active-turn follow-ups and the experimental settings lifecycle. A sequential conversation can pass even if messages sent during work are lost. **Expected behavior** A follow-up submitted during a provider turn survives and affects the final reply. A changed launch day appears in the saved plan. Disabling Agent Chat rejects new messages while preserving history; re-enabling resumes the same conversation. **Steps to reproduce** Run the explicit `agent-chat-stories` suite. It selects three local cases for each native Claude and Codex profile. An ordinary provider command waits for a fixture brief file so the browser can send the follow-up at an observed active-run boundary. ## What Changed - Add six opt-in Product E2E cells for settings, active follow-ups, and plan corrections. - Drive experimental settings through the UI and verify disabled sends are rejected by the public API. - Use a bounded file wait in the actual isolated agent workspace, with provider-written readiness and an undisclosed brief reference. - Grade persisted user messages, final replies, native run outcomes, and exact saved plan fields. - Accept active-turn steering or one queued successor; reject lost input, duplicate input, and stale outputs. - Allow one steered run or two sequential runs throughout the shared harness, while preserving exact counts for other cases. - Require a single marker-bearing response attributed to the final provider run. - Unload the development browser client before restarting the server, avoiding reconnect/navigation races without weakening the post-restart memory check. - Add browser regressions for restart isolation and asynchronously saved settings switches. - Document prepared-agent setup, native onboarding limits, and the separate API-tool rollout gate. ## Verification - Eval TypeScript check passed. - Eval support suite: 436 tests passed in 39 files. - New oracle calibration: six tests passed, including plausible invalid outcomes. - Browser support regressions: seven tests passed; the restart regression was observed failing before the fix. - Catalog discovery selects exactly six local native cases and leaves default paid selection unchanged. - [Consolidated existing native chat report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/): master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a browser navigation timeout across restart; the page request returned 200 and the chat rendered. - [Nine targeted restart/replay cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/) passed on `1fe2fe275`, including the original failure, across native Claude/Codex and local/Daytona; all cleanup passed. - [Initial six-story campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817) retained all six failures: asynchronous switch assertions, unavailable fixture paths, and rich-text escaping in raw command comparisons. The corrected fixtures preserve the same behavioral assertions. - [Six-story campaign v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/) on `8232773a0`: 4/6 passed (both settings cases and both Claude interruptions). Codex could not see the host-temp fixture outside its workspace; this failed before follow-up delivery was exercised. - [Four affected interruption cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/) all passed, including cleanup, on definition v3 / `ad6ac0545`. Files live inside the actual agent workspace and the observed run workspace is verified. Both providers saved Friday in the real plan with the undisclosed brief reference; follow-ups persisted while the original run was active. Together with both unchanged settings cases from v2, all six new scenario variants have passing live evidence. - Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful checks, two intentional skips, zero pending/failing checks; mergeable and clean. Fresh Greptile 5/5, zero unresolved findings. - Full typecheck, tests, build, and browser CI passed remotely. One earlier head encountered a signoff-policy browser timing failure; the final head passed that shard. - Local pnpm wrapper could not fetch its version/signature metadata in the restricted environment; local eval checks used the installed Node executables. Repo-wide validation was completed by GitHub Actions. ## Risks These are eval-only changes. The file wait is a timing fixture in the isolated agent workspace, not a production runner hook. Native Codex host-filesystem isolation stays unchanged. It has a two-minute limit and is released in `finally`. The prepared-agent settings case is not full native onboarding: the wizard currently offers legacy adapters. The disabled-entry assertion uses full document navigation, which clears the prior React Query cache; preserved history is checked through the public API and re-enabled chat. No production prompt, rollout default, adapter behavior, or credential policy changes. Active-task reassignment and worker-crash recovery remain outside these new cases. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8813a50105 |
feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path > - Paperclip manages agent work as tasks and runs. > - GitHub chat brings repository conversations into those tasks. > - A review bot needs the assigned agent, its authority, and governed provider tools. > - The existing channel connection did not supply that review workflow or a complete setup journey. > - This pull request adds GitHub App setup, account access, event prompts, task-bound review tools, and exact-commit checks. > - Operators can inspect each review through the same task, run, and activity systems. ## Linked Issues or Issue Description **Subsystem affected** GitHub chat, governed connection tools, task execution, shared/database contracts, and connector setup UI. **Problem or motivation** Operators need a GitHub review bot that runs their assigned Paperclip agent. Mentions and PR events must preserve task ownership and requester authority. Provider publication must use the bot App identity and enforce the configured permissions. **Proposed solution** Extend the existing GitHub chat connector with resumable App onboarding, linked-member and sponsored-guest access, editable event prompts, and governed review operations. Validate structured assessments on the server and compute a stable Paperclip Review check for the exact head commit. **Alternatives considered** A separate review scheduler would duplicate Paperclip execution and permissions. Reusing personal GitHub credentials would change the bot identity and credential boundary. **Roadmap alignment** This extends the existing Connected Apps and governed-tool infrastructure. The project owner requested and approved this design. Related PR #8645 imports external Codex review feedback; this change runs an assigned Paperclip agent and publishes its results through the existing chat connector. ## What Changed - Include the current Paperclip instance origin in the copied setup prompt. Storybook uses its configured Paperclip origin; callback parameters and URL credentials are excluded. - Add a Claude/Codex copy button in the real setup and Storybook opening step. Its detailed prompt asks four setup questions and guides embedded-browser setup, verification, and optional required checks. Clipboard failure exposes selectable instructions. - Add a tutorial that explains why App installation, review scheduling, and required checks are separate choices. - Add manifest registration, an existing-App path, separate installation and repository selection, repository refresh, and explicit account confirmation. - Add low-trust agent guidance, effective capability verification, member selection, and explicit restricted guests with a sponsor. - Add configurable PR events, prompts, repository overrides, rating thresholds, and separate formal-review permissions. - Give the assigned agent governed App tools to read PRs, comment, begin an assessment, submit findings, and optionally submit a formal review. - Bind review history, root PR events, and inline replies to ordinary tasks. Deduplicate deliveries/findings and reject stale publication. - Link check Details to the underlying task on the current trusted hostname, or to Reviews before task creation. - Add schema migration 0283, API contracts, production UI, and 49 interactive Storybook states. - Repair local lease recovery. Keep the Cloud Dockerfile identical to master; no provider-pack layer or runtime-default environment variable is added. - Retry only rolled-back wake-admission transactions after transient endpoint-lock contention. A deterministic held-lock regression proves one accepted wake. ## Verification - Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero diff against master. Final workspace typecheck and build passed. The new PostgreSQL migration regression passed and preserves existing relation and constraint identities after replay. - Greptile reviewed this exact head at 5/5. There are zero unresolved review threads and no merge conflicts. - All current-head checks are green: 54 passed and two conditional Storybook jobs skipped. This includes complete server/workspace test suites, build, typechecks, policy checks, Runner suites, browser suites, and security status. One timing-sensitive callback-ordering test passed in isolation and its CI shard passed one retry. The duplicate local full-suite run was stopped after CI completed; it is not counted as a local full-suite pass. - Before the final Slack rebase and migration renumbering, 186 focused GitHub tests, 14 native bootstrap cases, token gates, and Storybook build passed. The final rebase retained the new Slack communication guidance. - The embedded-browser setup test copied the full detailed prompt, including the configured Paperclip instance URL. Desktop and narrow layouts were checked. Component tests cover successful copying and clipboard failure with selectable text and retry. - Live local and hosted GitHub acceptance evidence refers to application revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks exercised issue mentions, automatic PR reviews, inline findings, repeated mentions, task continuation, and failing-to-passing checks after a push. The Storybook agent generated, built, and browser-rendered pages; missing acceptance text failed, matching text passed, and broken JSX produced an incomplete result. - Live cases also covered independently disabled push events, prompt injection, duplicate signed deliveries, rapid pushes, stale-result rejection, finding deduplication, and restart recovery. Formal reviews were denied while disabled and published only after explicit enablement. Check Details links pointed to the underlying task on the trusted hostname. - Those hosted native Claude runs used the provider-pack layer now removed from this PR. They do not prove native Claude works on the standard Cloud image. A replacement hosted native Codex run is not yet verified: the disposable QA tenant has only an Anthropic AI connection. No new staging or production deployment was made for the packaging removal. - Required-check merge enforcement could not be tested because the private disposable repository's GitHub plan rejected the rules configuration. Published success/failure/incomplete check states were verified directly. ## Risks - Latest master allocated migration 0282 to Slack. The GitHub migration is regenerated as 0283 with replay-safe table/index/constraint creation; a PostgreSQL regression verifies existing relations and constraints are preserved. Existing preview tenants remain subject to the fleet migration-history compatibility preflight; no bypass is introduced. - Migration 0283 adds company-scoped configuration, registration, review, and publication records. Existing connections retain their behavior until reviews/tools are enabled. - Signed webhooks and expiring registration state remain required. Hosted installations also need the companion narrow Cloud gateway exemptions. - Agent assessments can be incomplete or wrong. The server enforces coverage/result structure, current-head publication, rating policy, and separate formal-review permission; it does not replace code-review judgment. - No Cloud image packaging changes are included. Remote native ACPX/Claude and OpenCode retain their existing operator-supplied provider-pack prerequisite. Native Codex and Codex with managed MCP tools do not require that pack. Earlier staging deployment evidence refers to its stated revision, not this packaging-removal head. Production rollout and merging remain outside this change. ## Model Used OpenAI GPT-6 through Codex, with repository, code execution, API, and embedded-browser tools. The exact serving model ID and context-window size were not exposed by the environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d9b3a5653e |
feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors let people use the same tasks and agent tools from external conversations. > - Agents need communication guidance that fits the conversation medium. > - That guidance belongs in the original task context, without repeated instructions on each turn. > - Connection owners also need clear settings and a consistent way to remove a connection. > - This pull request adds initial Slack guidance, optional connection instructions, and chat connection menus. > - The benefit is clearer Slack replies with the existing Paperclip workflow and permissions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent replies in Slack and chat connection management in the Apps catalog. **Current behavior** Slack tasks do not carry a saved communication profile. The catalog shows a separate Manage button and does not offer removal on every chat connection row. **Proposed behavior** Save Slack guidance when a new conversation creates a task. Restore that original guidance when a model session is rebuilt. Do not append it to ordinary follow-ups. Expose optional additional instructions in Slack Settings. Put Manage and Remove connection in a three-dot menu for all chat providers. Keep Finish setup visible for drafts. **Reason and benefit** Small answers fit in Slack. Substantial deliverables use ordinary document or artifact tools with a useful Slack summary. Connection settings apply to new tasks and cannot change permissions. Users can remove both active and unfinished chat connections from the catalog. **Breaking changes** Two additive database columns store endpoint preferences and the initial conversation snapshot. Existing endpoints default to empty preferences. Existing conversations keep their original behavior. Non-Slack guidance is unchanged. Related public context: https://github.com/paperclipai/paperclip/pull/13741 improves native chat recovery. This change adds communication context to those existing execution paths. A search found no duplicate communication-guidance PR. ## What Changed - Add a provider-guidance registry, enabled for Slack first. - Persist optional endpoint communication instructions and capture an immutable snapshot when a conversation creates a task. - Resolve guidance from the verified company-scoped connection. Restore it for fresh native and legacy sessions without per-turn reminders, extra model calls, or extra context queries. - Add the Slack Settings field, validation, audit coverage, and Storybook save/error states. - Add Manage and Remove connection menus for all seven chat providers. Keep the draft setup button. Require removal confirmation and allow retry after failure. - Add regression coverage, an active/draft menu story, and connector documentation. ## Verification All CI checks are green for |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3790ca2f13 |
fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Runner tasks must continue after approval and stop when the user presses Stop. > - Live evals found races at approval delivery and provider startup. > - Browser readiness and CI setup errors also hid the actual task results. > - This pull request fixes those races and the related test infrastructure. > - Regression tests and saved live reports show which cases now pass. ## Linked Issues or Issue Description Companion eval definitions PR: https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused and provider/environment infrastructure). Related: #13741 now supplies the late-startup Stop fence and warm-attachment recovery; this PR retains that fence and extends startup tracking and regression coverage to both native backend paths. #13539 introduced queued approvals during active runs. #13738 fixes child assignment, task replies, and warm process continuity and is already in the base. #13291 concerns automatic continuation of interrupted legacy sandbox runs; this PR fixes native startup cancellation and does not change that recovery policy. **What happened?** An accepted service approval could wait after its source run stopped. Stop could return success before the provider handle existed. Work could then start after Stop, or a cancelled run could be recorded as failed. Some E2E tests also failed on unloaded browser content or irrelevant reply wording. Runner CI could fail before model work because of dependency or sandbox setup. **Expected behavior** Deliver each settled approval once after its source run stops. Do not start work after an acknowledged Stop. Preserve the audited cancellation. Test the intended product behavior with a ready browser and verified runtime dependencies. **Steps to reproduce** 1. Approve a service request while its source run is active. Let the run finish. Check that its result starts one continuation. 2. Delay provider startup. Press Stop before its handle is available. Check cancellation, then submit `/new`. 3. Run the browser, warm-workspace, and Stop-and-redirect cases from the linked report. **Paperclip version or commit** The branch includes master at `9d19f98b5`. The report records the original source for each focused attempt. **Deployment mode** Isolated local development instances and disposable Daytona sandboxes. ## What Changed - Deliver settled tool-action results for the exact company and source run during final cleanup. Keep the existing idempotent receipt and periodic recovery sweep. - Wait for startup to hand off its provider handle before acknowledging Stop. Reject first-turn admission after cancellation. Preserve a matching audited pending or acknowledged cancellation. - Wait for mounted task history and connector controls in browser tests. Record failure evidence. Grade workspace contents and process continuity separately from exact reply wording. Require each warm-turn marker once and in order, allowing surrounding prose. - Stop-and-redirect now checks that the source file exists and work is active before Stop. - Resolve target dependency locks in an uncredentialed CI job. Verify the lock artifact hash. Keep orchestration and publication on the trusted workflow revision. - Materialize the pinned OpenCode executable and configure the exact Codex executable's user-namespace profile before provider credentials are available. - Compress Daytona directory uploads with gzip. Preserve files, executable modes, symlinks, empty directories, and confinement checks. - Classify file-transfer RPC deadlines as infrastructure. Keep unrelated runner RPC failures visible. ## Verification - [Focused live report with screenshots and original attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14 of 15 selected Product E2E cases pass across the recorded revisions. Claude and Codex Stop → `/new`, Claude service approval, delegation, both hiring/reuse cases, and native Daytona warm continuity pass. - Two credentialed Runner smoke cases pass. These are not full protocol coverage. - E2E harness after the master merge: 429 tests pass. E2E and server TypeScript checks pass. - Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build passes. The compression test fails against the old code and passes with the change. - Runner backend/runtime regression group: 161 tests pass. Cancellation/startup selection: 26 tests pass. Approval delivery: 34 real-database tests pass. - Workflow security: 7 tests pass. Both edited workflows pass actionlint. Runner TypeScript and Rust builds pass. - After merging master, all 389 native executor tests pass, including both native backend paths and late startup after the Stop deadline. - Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic local `pnpm test:run` was interrupted to integrate master and is inconclusive. The [hosted CI test partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738) pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox callback schema test initially received HTTP 503. It passed five isolated local runs, its full local test file, and one failed-job CI retry. No assertion was weakened. ## Risks - Stop can wait for the bounded startup handoff. If it cannot settle, the existing pending-recovery state remains instead of a false acknowledgement. - Immediate approval delivery must remain idempotent across cleanup and recovery sweeps. Tests cover duplicate delivery and company/run boundaries. - The workflow changes still need hosted Linux verification. They retain the trusted workflow and credential boundaries. - Gzip reduces the observed provider upload from about 1.8 GB to 663 MB. It does not yet fix the remaining Claude Daytona transfer timeout. That recovery test never reached Claude, so recovery remains unverified. Use a matching image with the verified provider package preinstalled for the next recovery test; retain cold-upload coverage separately. - The report preserves diagnostic runs with missing source metadata and marks them as such. It does not claim a new full-suite pass. - This PR adds no new prompt policy or historical status reconciliation. ## Model Used OpenAI GPT-6 through Codex performed the primary implementation and review. The exact primary backend model ID is not exposed in this session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work and verification. The agents used repository tools, code execution, and browser tests. The exact backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 <noreply@openai.com> |
||
|
|
0103cb2292 |
test: use neutral references in chat hiring evals (#13752)
## Thinking Path > - Paperclip helps people manage agents and their work. > - The agent chat hiring eval checks delegation, saved output, and worker reuse. > - It asks a worker to include a unique reference in each document. > - The open wording let the worker choose a credential-style label. > - The existing redactor removed that reference and failed the coordination check. > - This PR specifies a neutral reference line while keeping the same grading checks. ## Linked Issues or Issue Description Refs #13741. **What happened?** The Codex hiring eval saved a checklist with `Tracking token: [REDACTED]`. The runner treats the chosen label as credential syntax. The previous prompt only asked for the identifier and did not choose its label. This tests a redaction boundary unrelated to hiring and reuse. **Expected behavior** The coordination fixture requests ordinary business content with a neutral reference label. The grader still requires the exact identifier in the saved output. **Steps to reproduce** Run `agent-chat-hardening.runner-codex.local.hire-delegate-reuse` with the previous fixture. The retained failing attempt is in [campaign 35617045456](https://github.com/paperclipai/paperclip/actions/runs/35617045456). ## What Changed - Request `Reference: ...` in the checklist and review documents. - Increment the hardening suite definition to version 5 and record the reference format. - Document the fixture boundary. Production redaction and all grading checks stay unchanged. ## Verification - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - The catalog lists exactly the Codex and Claude local hiring cells for the selected case. - [Live campaign 35620731321](https://github.com/paperclipai/paperclip/actions/runs/35620731321) passed both selected local cells on `d8d7afe21`: native Codex (`gpt-5.6-sol`) and Claude (`claude-sonnet-5`). Both passed on attempt 1, including cleanup. [Published eval report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35620731321-1/). - Inspected saved evidence: each provider preserved both exact reference lines, used the same hired worker for checklist and review, and completed the final status query. Definition version 5 and hash `2c1c9f754c049c3e7cafd6c8d0a137e7ca06d546ce1e7b07ad5b7823f0f28db9` distinguish these results from the prior prompt. - [PR CI 35620741101](https://github.com/paperclipai/paperclip/actions/runs/35620741101): full build, type checks, and test partitions passed. One unrelated Telegram integration test initially failed because two random fixture company IDs produced the same seven-character issue prefix. The failed shard passed on one retry; no test code was changed. All latest-head merge checks are green. - Greptile reviewed `d8d7afe21` at 5/5 with no review threads. ## Risks The live models can still fail the coordination workflow. This change does not qualify or change credential-redaction policy. Earlier failed attempts remain part of the evidence; the new fixture has a distinct definition version. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9d19f98b50 |
fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Agent chat uses native runner sessions to plan, delegate, and track that work. > - A user can press Stop while the native session is still starting. > - The server can acknowledge that Stop without dispatching it, then let the session submit a turn. > - This leaves chat recovery waiting for an execution that the user expected to stop. > - This PR waits for the startup handle, dispatches cancellation, and prevents a late startup from submitting a turn. > - New full-stack evals check the resulting records and outputs across Claude and Codex. > - Those evals also exposed missing ACPX readiness fields, unbounded polling, and an old-run identity check that rejected valid warm handoffs. ## Linked Issues or Issue Description **What happened?** Stop during native startup could record an acknowledged cancellation with `dispatched: false`. The provider could then begin work. A subsequent `/new` stayed queued. A remote Claude follow-up also exhausted the command journal while probing warm-session readiness: ACPX never returned the readiness fields required by the shared transport. Once readiness worked, attachment incorrectly compared the next run descriptor against the old run ID. The 25 ms polling loop could issue 4,800 commands during its two-minute wait, beyond the 500-command bound. The existing chat eval treated lifecycle logs as proof of an active provider turn, so it did not distinguish startup cancellation from active-turn cancellation. **Expected behavior** A Stop during startup must reach the pending session. A late session must not submit a prompt after Stop. Recovery must retain control when startup exceeds the bounded wait. Chat evals must check saved task state, document contents, worker identity, account binding, and duplicate effects. **Steps to reproduce** 1. Start a native Claude or Codex chat turn. 2. Press Stop after process startup is requested but before the provider turn starts. 3. Send `/new`, then send a fresh message. 4. On the affected base, cancellation can be acknowledged without dispatch and the reset stays queued. **Paperclip version or commit** The live Claude baseline reproduced this on `29d6b3509`. The branch also includes master commit `0f5fafe16`. Related work: #13678, #13686, #13693, #13291, #13738. A separate runner reliability branch also contains a startup-wait fix. Its overlap must be reconciled before merging; this branch additionally prevents prompt submission after a late startup. ## What Changed - Wait for a pending native startup before acknowledging a run-scoped Stop. Preserve the existing recovery error when that wait expires. - Keep a Stop guard on startup. Cancel a late handle before it can submit a provider turn. - Add regression tests for normal handle publication and publication after the Stop deadline. - Back off blocked warm-attachment probes. Keep the fast two-snapshot barrier, fail closed, and record changed blockers. - Add red/green tests for delayed readiness, persistent blockers, alternating readiness, and readiness near the deadline. - Publish ACPX readiness and blockers. Preserve the old authority’s event acknowledgement barrier; only settled sessions can proceed to attachment. - Bind warm ACPX descriptors to the validated next authority while retaining old-run event correlation until activation. Preserve session identity and provider profile checks. - Exercise two consecutive run rotations through a qualified fake sidecar, verifying checkpointing, provider identity, pre-activation rejection, and new-run work admission. - Separate startup and active-turn cancellation checkpoints in the browser eval. - Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells across Claude and Codex. - Cover hiring and reuse through managed AI accounts, source-based review, current blocked-task status, request replay after a lost HTTP acknowledgement, server restart continuity, and Stop/reset continuity. - Use ordinary production agent instructions. Enable API tools only for the two coordination cases that need them. - Calibrate the matchers with invalid records and outputs. Require remembered context after restart and a structured status snapshot that distinguishes the current blocker from history and task status from active execution. Compare the public issue mutation contract and relationships during read-only reporting. Preserve before/after source records in failed eval evidence. - Fix the lost-ack browser harness and verify it against a real HTTP server. Check the chat composer after restart instead of waiting for an unrelated document lifecycle event. - Document the scope and limits of each case. ## Verification - The startup regression failed on the unfixed executor and passed after the fix. - `pnpm test:e2e:runner:typecheck` passed. - `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files. - `pnpm exec vitest run server/src/services/native-runtime/native-session-executor.test.ts` passed: 385 tests. - [Baseline live campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868): Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed. Codex delegation was blocked by provider capacity. - [Eval-only startup campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786): both providers failed as expected. Both persisted `dispatched: false` and left `/new` queued. - [First fixed campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706) on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both providers. Failed cases exposed eval harness defects and remote continuity failures. All attempts remain available. - [Original workflows and stronger memory checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649) on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and startup Stop passed for both providers; Codex remote restart passed. Claude remote restart exposed the missing readiness contract. Two Codex planning cells hit provider capacity. - [Unchanged-model retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548): Codex planning and backlog creation both passed. - [18-cell campaign with ACPX readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963) on `6a98ef743`: 16/18 passed, including all local/remote Stop and committed-send cases. Claude remote continuity exposed the next-authority check, now fixed. Codex hiring produced its checklist, but the runner redacted the requested marker after it appeared as “Tracking token: …”. That content-redaction policy is unchanged and remains an explicit limitation. - [Structured status grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725) on `50448c228`: both providers passed on their first attempt, including cleanup. - [Complete read-only state grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011) on `551e13892`: both providers passed. - [Final ACPX handoff and hiring retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456) on `cbd637587`: all three Claude Daytona cases passed (restart continuity, active Stop/reset, and lost-ack replay). Codex hiring reproduced the content-redaction failure: the saved checklist contained `Tracking token: [REDACTED]` instead of the required business marker. All four cases completed cleanup successfully. [Published report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/). The only subsequent commit adds the qualified-sidecar integration test; production code is identical to this live proof. - `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without paid models. - Runner TypeScript typecheck passed. All 5 warm-readiness tests pass; two failed with the prior fixed-rate loop, and the late-readiness test failed before the pacing correction. - ACPX readiness and warm-identity regressions each failed before their fixes. All 292 runner-core Rust library tests passed. The qualified-sidecar integration test passes. Rust formatting is checked. - Status-grader regressions for misleading historical mentions and previously unchecked mutations each failed before tightening the oracle and pass now. - [Latest-head CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307) passed on `a4093c8f1`: full build, type checks, test partitions, browser E2E, and native runner checks. Two unrelated tests initially failed (Sentry fixture release attribution and local-service fixture readiness); both passed locally together (35 passed, 5 optional SDK tests skipped) and on the failed-job retry. No changes were made to those tests. - Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed and all review threads are resolved. - The paid live suite is not fully green: the reproducible content-redaction case remains red. This is separate from the passing PR merge checks. No production content-redaction, prompt, model, or completion-policy change is included. - Managed-account hiring and review cases explicitly enable API tools; these do not qualify default new-user onboarding. ## Risks - Stop can wait up to 30 seconds for startup, then use the existing pending-recovery path. This does not prove that remote cleanup has finished. - Blocked warm readiness adds up to 750 ms between later probes with the two-minute remote budget, or about 32 ms with the default five-second budget. Ready sessions retain the short second barrier. - Paid evals can fail because of provider capacity or agent decisions. Each failure needs evidence-based classification. - The HTTP request replay case checks comment idempotency and duplicate effects. It does not prove replay safety for an ambiguous provider tool call. - The new suite is opt-in. It does not increase the default paid campaign. - No production prompts or model selection change. Review-handoff behavior and content-redaction policy remain separate product decisions. The latter can remove harmless business content that looks like credential syntax; the failing attempt is retained. ## Model Used OpenAI Codex, GPT-6, with repository tools and code execution. The exact deployment model ID and context window are not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e9bc2efdaa |
docs: turn DeepWiki into a contributor handbook (#13745)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Contributors need to understand its control plane and find the code that owns each behavior. > - The public DeepWiki has many pages, but it lacks a guided path through task execution and system boundaries. > - A refresh alone cannot define that reading path or the evidence each page must show. > - This pull request adds a 30-page contributor handbook configuration with source anchors and clear page boundaries. > - The handbook helps contributors trace work, investigate failures, and choose the right code and tests. ## Linked Issues or Issue Description **Issue type** Unclear or confusing documentation; outdated generated documentation. **Where is the issue?** [Paperclip DeepWiki](https://deepwiki.com/paperclipai/paperclip). The repository did not have a `.devin/wiki.json` configuration. **What's wrong?** The wiki needs a contributor reading path. Readers need to connect product terms to implementation, follow a task through both execution paths, and find the state and tests that explain failures. Index freshness also depends on regeneration. **Suggested fix** Define six chapters and 30 pages in `.devin/wiki.json`. Give each page source anchors, required questions, and coverage boundaries. Use shared notes for evidence standards, diagrams, terminology, and freshness rules. Keep existing generation effort settings. Searched public issues and PRs for DeepWiki, wiki, and contributor handbook work. No duplicate or related DeepWiki change was found. Checked `ROADMAP.md`; this is a documentation configuration change. ## What Changed - Add only `.devin/wiki.json`, using the [documented DeepWiki configuration](https://docs.devin.ai/work-with-devin/deepwiki#steering-deepwiki). - Define 30 pages under Understand Paperclip, Contribute to Paperclip, Orchestrate Work, Run Agents, Extend and Integrate, and Govern and Diagnose. - Add 68 generation notes with 244 verified repository source paths. Require source and test citations, useful parent pages, and explicit ownership boundaries. - Require a task tour with separate direct-adapter and experimental Runner paths, run and task state diagrams, and a connection authorization flow. - Separate run logs, operator-configured observability, and first-party telemetry. Require defaults and feature gates to be resolved from the indexed code. ## Verification - Passed JSON and supported-field validation. Checked unique titles, valid parents, an acyclic hierarchy, the exact page tree, note limits, and all 244 source paths. - Reviewed coverage against the six reader questions in the plan. Confirmed each page has source anchors, test anchors, and coverage boundaries. - Passed `git diff --check`. - Passed `pnpm -r typecheck` and `pnpm build` on the implementation base, `c65fc9e3c81c41aafe421aa90a00514b84343285`. - Ran `pnpm test:run`. The broad run hit a 15-second timeout in `server/src/services/native-runtime/remote-deliverable-file.test.ts`, in “reads verified remote bytes without touching controller paths”. Stopped the broad run after the failure. Full-suite verification remains incomplete. - Reran that test file in isolation: all 30 tests passed in 6.45 seconds. - Rebased onto current `master`, `57fd8b70d`. Repeated configuration, source-path, and diff validation after the rebase. No application tests were added for this configuration-only change. - After merge, request DeepWiki regeneration. Confirm its indexed commit contains the configuration. Then check the generated task tour and one page from each chapter against implementation and test citations. Regeneration and generated-page review are pending. ## Risks - Generated text can still contain errors. The instructions require citations and honest treatment of conflicting evidence, but the generated pages need review. - The new page structure can change navigation and page links. - Source paths and implementation can change between generations. Notes require replacement anchors, current defaults, and clear labels for experimental behavior. - The public index stays stale until someone requests regeneration. This PR adds no refresh service and changes no application behavior. ## Model Used OpenAI Codex, based on GPT-6. The exact runtime model identifier, context window size, and reasoning setting are not exposed in this session. Used repository inspection, web research, shell tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
57fd8b70d2 |
feat: add agent avatar download to Slack setup and settings (#13740)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Slack connections let a team talk to those agents in Slack. > - Agents now have a saved avatar, but Slack setup did not offer that image. > - A matching avatar helps a team recognize its agent. > - This pull request adds an optional avatar step and a download in connector Settings. > - Users download a PNG and upload it directly in Slack with clear instructions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Slack connector onboarding and its Settings page. **Current behavior** Setup does not offer the assigned agent's avatar or explain how to upload it in Slack. **Proposed behavior** After Slack connection verification, users can download a 512 × 512 PNG of their agent's saved avatar. They can upload it in Slack, confirm, or skip. Settings keeps the download and upload instructions available after onboarding. **Reason and benefit** The same avatar helps people recognize the agent across Paperclip and Slack. Users who skip the optional step can return to it in Settings. **Breaking changes** None. No schema, authentication, Slack scope, or provider API change. Completed connections keep their existing completion state. Searched existing Slack avatar and Cliptoon PRs; no matching implementation was found. ## What Changed - Add an optional avatar step before personal Slack account linking. Keep the numbered sidebar and shared footer. - Resolve the selected agent's saved appearance for the preview and PNG download. - Add the same download and expandable upload instructions to connector Settings. - Remember uploaded or skipped per company and endpoint in browser storage. Treat uploaded as user confirmation, not provider verification. - Reject failed or non-PNG download responses and allow retry. - Reuse the production avatar components in onboarding and Settings stories. - Test wizard progression, resume, Settings, download recovery, storage isolation, and terminated assigned agents. - Exercise real PNG downloads in the Slack browser flow and keep default app names consistent with app creation. - Fetch the assigned agent directly so its saved avatar remains available after termination. ## Verification - Focused chat suites: 48 passed; the two affected suites passed again after the final naming fix (32 tests). - Slack browser E2E passed through setup, avatar download, account linking, and Settings download. PNG signature and 512 × 512 dimensions verified. - UI token gates passed. - Browser: downloaded the real 512 × 512 PNG; checked confirmation, return, mobile layout, and Settings instructions. - Full workspace typecheck, application build, and production Storybook build passed. - All latest-head CI checks passed (54 passed, 2 skipped), including all browser, chat, general, and serialized test groups. The unrelated Sentry test failed once and passed on the single CI rerun; its suite also passed locally. - Local full-suite attempt encountered a rapid Slack callback ordering failure under concurrent build load; that test passed in isolation, and all three chat shards passed in CI. The remaining local run was not used as the merge gate. - Review the Connections / Slack / Add avatar and Avatar in Settings stories. ## Risks - Slack upload is manual. Confirmation does not claim to verify the Slack icon. - Optional step progress is browser-local. Clearing storage or changing browsers can show it again. Setup still works when storage is unavailable. - The existing avatar API remains the image source. Download failures show a retry message. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fd071748ee |
fix: stop repeated notifications for finalized run failures (#13739)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native run reconciliation repairs saved execution outcomes after
interruptions.
> - The sweep also visits runs whose final results are already
committed.
> - An unchanged failed run still received a new status delivery ID on
every sweep.
> - The browser treated each delivery as a new failure after its short
duplicate window expired.
> - This pull request makes unchanged projections a no-op and suppresses
repeated or historical run toasts.
> - Operators receive fresh failure alerts without repeated alerts for
old work.
## Linked Issues or Issue Description
**What happened?**
An old failed run repeatedly produced failure toasts while the browser
remained open. The task could already be cancelled. Reconciliation
rewrote the same failed outcome and queued another status broadcast.
**Expected behavior**
An unchanged committed run must not queue a new status notification.
Repeated deliveries must still refresh cached state without another
toast.
**Steps to reproduce**
1. Finalize a native run with a failed result and successful workspace
finalization.
2. Deliver its pending execution status and cancel its task.
3. Replay finalization and status delivery on each periodic sweep.
4. Observe another failure broadcast for every sweep before this fix.
**Paperclip version or commit**
Reproduced against source commit
|
||
|
|
0f5fafe16b |
fix(runner): preserve task replies and warm process continuity (#13738)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - The Runner connects provider sessions to task state, replies, and delegated work. > - Full-stack tests found lost final replies, rejected helper calls that stopped the parent, and unnecessary process restarts. > - A completed child could also receive a new assignment wake that the scheduler then cancelled. > - This change fixes those boundaries and gives agents clearer teammate instructions. > - The tests retain strict completion and process-continuity requirements. ## Linked Issues or Issue Description **What happened?** A generated attachment comment could suppress an agent's final reply. A known Codex helper could stop its parent when it requested a Paperclip tool. Native Daytona processes restarted between turns because Paperclip minted an unused GitHub broker token. Reassigning a completed child queued a run that immediately cancelled. Revision instructions also allowed agents to do work assigned to a named teammate themselves. **Expected behavior** Keep the final reply. Reject helper tool requests without borrowing parent authority or stopping the parent. Keep an unconfigured sandbox process alive between turns. Treat assignment-only changes to completed tasks as metadata changes. Preserve explicit teammate assignments during revisions. **Steps to reproduce** Run the retained Runner E2E cases for file handoff, teammate reuse, Daytona warm continuity, and Legacy Claude interview/plan acceptance. The focused regression tests reproduce the reply, helper, process-lifetime, and assignment-wake defects without provider calls. **Paperclip version or commit** The live lifetime and completion campaign used `db3857807`. This PR replays the changes on master `c65fc9e3c`. See Verification for the limits of that evidence. **Deployment mode** Isolated local instances and native Runner sessions in Daytona sandboxes. Related work: #13546 handles a different queued-run issue after an issue-lock compare-and-set failure. This PR prevents the unnecessary assignment wake earlier. #13410 covers retained user services; this PR covers the provider process. No duplicate fix was found. ## What Changed - Exclude generated deliverable-binding comments from final-reply deduplication. Preserve the attachment and explicit user-facing replies. - Reject Paperclip tool and input requests from known Codex helper threads without terminating the parent. Keep unknown-thread rejection intact. - Explain how to hire or reuse a persistent teammate and preserve named delegation on revisions. Update generated protocol fixtures. - Use stable, token-free GitHub wrappers for unconfigured native sandboxes. Preserve credential isolation, configured-account rotation, and cleanup after partial staging failures. - Do not queue assignment-only wakes for done or cancelled tasks. Keep explicit reopening behavior. - Make warm-continuity fixtures create real review cards. Read the persisted final response selected by production presentation logic. Missing selected evidence still fails. ## Verification - Before rebase: 560 focused route, native-executor, and launcher tests passed. The new regressions were reproduced before their fixes. - Live E2E: Legacy Claude interview/plan acceptance passed 3/3 repetitions. Daytona warm continuity passed 2/3 full repetitions. Each successful run retained one process and provider session for all three turns. - The remaining Daytona repetition stopped after a same-URL browser reload left the page blank. Both completed turns retained the same process. Its failed verdict remains unchanged; this PR does not claim the blank-page cause is fixed. - Reports: https://pages.paperclip.ing/runner-e2e-lifetime-race-20260920/investigation.html and https://pages.paperclip.ing/runner-e2e-behavior-followups-20260919-results/investigation.html - Post-rebase `pnpm build` and `pnpm -r typecheck` passed. All 414 Runner E2E harness unit tests and its typecheck passed. Codex protocol tests: 88 passed, 2 ignored. - Latest-head CI: 55 successful checks and 2 intentional skips. Greptile: 5/5 with no review threads. The unchanged workspace exposure tests hit a fixed-port collision on the first CI attempt; their local suite passed (25 tests, 3 platform skips), and the CI shard passed on one retry. - The duplicate local `pnpm test:run` was stopped after the full hosted general and serialized test shards passed. It did not finish locally and is not counted as a local full-suite pass. ## Risks Configured GitHub accounts retain run-scoped credential rotation and can still restart warm processes. That limitation requires a separate design. Known provider helpers cannot use Paperclip coordination tools directly; they must return findings to the parent. The delegation prompt is an instruction, not an enforced guarantee; Codex Mini hiring/reuse failures remain open. No schema or workflow changes are included. ## Model Used OpenAI GPT-6 through Codex, with repository tools, code execution, and parallel coding agents. The exact deployment suffix and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub issues and shared reports) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full hosted test shards also pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9f30eb10dd |
fix: reduce chat latency and preserve managed session reuse (#13710)
## Thinking Path > - Paperclip manages AI agents and keeps their work attached to tasks. > - Chat connectors carry user messages and agent replies between a provider and those tasks. > - Each extra startup and context reset delays a reply. > - Managed account metadata was lost during adapter decoding, so compatible follow-ups started fresh. > - This branch fixes the reset and measures the remaining preparation, execution, and delivery costs. > - The changes must preserve account isolation, authorization, durable output, and recovery ownership. ## Linked Issues or Issue Description Refs #13699. The related service lifecycle work in #13410 and #13408 is separate; this branch focuses on task-bound chat response latency. **What happened?** Managed AI follow-ups started new provider sessions even after their configuration fingerprint stayed stable. The Codex codec removes unknown fields. The resume check then read the removed credential identity and treated it as a credential change. **Expected behavior** Compatible follow-ups resume the correct provider session. Changes to credentials, responsible users, permissions, or task configuration retain their reset behavior. **Steps to reproduce** 1. Use a Slack connector with a managed AI connection. 2. Send a message, then send a same-thread follow-up. 3. Inspect the configuration reset reason and the provider session identity. **Paperclip version or commit** Reproduced on `2a99de80ec52db01eead901f28323926ceaf3c1d`. **Deployment mode** Cloud staging with a native Codex runner. ## What Changed - Read saved credential identity before adapter decoding discards it. - Remove the internal credential identity from adapter-facing session params. - Test the real Codex codec and missing, changed, or unmanaged identity cases. - Preserve configured warm Codex runners and flush refreshed credentials after every turn. - Fence detached or closing session handles from successor credential ownership. - Stage current Codex launch credentials after restoring durable session history, uploading launch assets only once. - Reuse a runner binary already in the retained sandbox only when its SHA-256 matches the controller-owned artifact; still verify required capabilities before launch. - Lock the task before the run when saving results, preventing deadlocks with task updates. - Scope reusable projectless sandboxes to the company, environment, task, agent, and runtime configuration; verify Daytona sentinels for that scope. - Admit a new authorized chat message after a fully committed failed run and verified process cleanup. - Send compact deltas for verified plain-text Slack continuations. Match the actual prior run and current comment identity/body; exclude edited historical comments and prior agent output, preserve genuine brief edits and the full bootstrap fallback. - Keep attachments, omitted input, questions, approvals, recovery, and other providers on their existing framing. - Document managed session compatibility, credential lifecycle, and compact continuation boundaries. ## Verification - Workspace/session coverage: 156 tests passed. - Native session and credential ownership coverage: 390 tests passed, including exact artifact reuse, mismatches, failed probes, timeouts, and explicit artifact overrides. - Explicit continuation and durable chat authorization coverage: 172 tests passed. - Session resume and launch preparation coverage: 416 tests passed. - Result persistence coverage: 15 tests passed. The new concurrency test reproduced a PostgreSQL deadlock before the lock-order fix. - Environment lifecycle coverage: 92 tests passed, including projectless reuse and task/agent isolation at both selection and atomic handoff. - Daytona plugin coverage: 237 tests passed; 6 gated tests skipped. Standalone plugin build passed. - Compact Slack continuation and native resume coverage: 69 tests passed, including full-bootstrap retention, matching message authors/bodies, current-delivery selection, rejection of duplicate identities and historical comments, brief edits, and attachment/recovery fallbacks. - Final frozen-head `pnpm test:run` on repository-supported Node 26: 668 suites passed, 3 skipped, 1 failed; 12,797 tests passed and 82 skipped. The sole failure was a local `socket hang up` in `issue-recovery-actions.test.ts`, not an authorization assertion mismatch. All 57 tests in that suite passed three fresh reruns, and the suite passed latest-head CI. The full local invocation is therefore not claimed green. - An earlier Node 24 full run exposed an unrelated macOS symlink-cleanup failure; that 11-test catalog suite passes on Node 26 and in CI. No test behavior or timeout was relaxed. - Full local typecheck and build passed. Latest-head CI is green; Greptile is 5/5 with no unresolved review threads. - Two real Slack baseline replies took 25.1 and 24.6 seconds (24.9-second mean). Three same-thread signed probes on this head took 23.8, 23.9, and 22.7 seconds (23.5-second mean). This is a small sample and a modest wall-clock improvement, not a large or statistically established speedup. - In that same thread, uncached provider input fell from 8,514 tokens before compact input to 694–765 tokens afterward. The current delivery uses a 362-character delta; the full 19–21k-character bootstrap remains available for failed resume. Verified runner artifact preparation fell from about 1.2 seconds to 0.6 seconds. - A fresh thread created a separate task, sandbox, and provider session with full bootstrap (24.4 seconds). Its follow-up reused its own sandbox/session and compact input (28.3 seconds, including 16 seconds of model execution). Model variability and process startup remain substantial. - A signed duplicate webhook produced exactly one user comment, one successful run, and one final Slack reply. Slack's API independently confirmed the actual replies and a public task URL without an internal or pool hostname. - Earlier signed probes verified recovery after a failed run and reuse across a server deployment. The final idle test observed Daytona report the sandbox as stopped, then delivered a new reply in 19.9 seconds using the same sandbox/provider-session identity and compact input. Slack’s API confirmed that reply. - Live probes use signed synthetic inbound webhooks and real outbound Slack delivery, read back through Slack’s API. The final browser recheck found the Mac locked and the Slack tab blocked by another extension, so this is not claimed as full UI E2E proof. - This is a review branch. Do not merge until the maintainer reviews it. ## Risks - Incorrect session reuse could mix account or task context. Missing or changed identities continue to reset, and existing authorization checks remain in place. - Warm mode remains opt-in. Remote warm mode requires a reusable sandbox lease. Retained processes keep credentials until they close, so idle expiry and ownership fences are required. - A fresh user message may continue after a committed provider failure. Approval, current authorization, process termination, and prior-result checks remain required. - Projectless sandbox reuse is task- and agent-scoped. Missing or mismatched ownership cannot replace an existing lease; existing workspace-scoped leases keep their scope. Opt-in reuse retains a sandbox per task/agent, so provider auto-stop and deletion policies still determine idle compute and storage costs. Fleet defaults are unchanged. - Compact prompts apply only after proven resume and a matching prior-run delta. Missing or specialized context falls back to full input; fresh sessions always receive the full bootstrap. - No schema or migration changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, tool use, and test execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2a99de80ec |
fix: supply public task links and preserve managed AI sessions (#13699)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors deliver agent replies to external conversations.
> - Agents need a public task link when a user asks to open the task.
> - The prompt and task tools lacked that link, so an agent could invent
an internal address.
> - Managed AI credential directories also changed the session
fingerprint on each run.
> - This change supplies public task URLs and excludes only those
temporary directory values from the fingerprint.
> - Follow-up messages can reuse compatible sessions while real
configuration changes still reset them.
## Linked Issues or Issue Description
Refs #13680 and #13694 for the related Cloud-origin fixes. No duplicate
open PR was found.
**What happened?**
An external chat reply could contain an invented internal task URL. The
publication filter then removed the link. Follow-up runs also lost their
saved provider session because each managed credential home used a
different temporary path.
**Expected behavior**
Agents receive the current public board URL for a task. Temporary
credential directories do not reset an otherwise compatible session.
Account, credential, model, permission, and custom environment changes
still invalidate it.
**Steps to reproduce**
1. Use a chat connector with a managed AI connection.
2. Ask for the current task link.
3. Send a follow-up message with the same agent configuration.
4. Inspect the task URL and the session reset reason.
**Paperclip version or commit**
Reproduced on the source at
|
||
|
|
b193077582 |
fix: report Slack callback health correctly behind Cloud proxies (#13694)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack connections turn messages into governed agent runs and return replies to Slack. > - Remote runs must replace an incompatible sandbox runner with the controller's packaged binary. > - The fallback used a package-relative path that does not match the vendored server layout. > - Slack callback health also compared the internal proxy address with the public callback address. > - Master now contains the runner fallback fix; this pull request fixes callback health and extends missing-runner regression coverage. ## Linked Issues or Issue Description Refs #13677, #13680, #13691, and #13686. **What happened?** A packaged controller stopped a Slack-triggered remote run with `runner_remote_artifact_unavailable` when the sandbox runner needed replacement. Working Slack callbacks also showed a stale URL warning behind the Cloud gateway. **Expected behavior** The controller stages its packaged runner when needed. Callback health uses the observed public address and still detects real address changes. **Steps to reproduce** 1. Run a packaged server with a sandbox that has an older runner or no runner. 2. Send a Slack mention to an agent that uses that sandbox. 3. Route signed Slack callbacks through a claimed Cloud gateway that rewrites the upstream host. 4. Check the run and the Slack callback health panel. **Paperclip version or commit** Reproduced on master at `aeef493f4a7603b7b1254421b80fb00212982390`. The fix branch also includes #13691. **Deployment mode** Packaged server with a Cloud gateway and a Daytona sandbox. ## What Changed - Extend the controller-owned runner fallback tests with a missing sandbox binary case; retain the fix now merged in #13686. - Prefer dedicated gateway diagnostic headers that survive provider rewrites of standard forwarded headers. Use validated host hints only for callback-health evidence on claimed Cloud instances after provider acceptance. Preserve request bodies, routing, authentication, and configured callback URLs. - Cover current, stale, and missing sandbox runners, all Slack callback surfaces, rejected callbacks, malformed proxy hints, real host and port changes, and the existing self-hosted behavior. - Document the artifact lookup and callback-health boundaries. ## Verification - Native session executor and binary resolver suites: 379 tests passed after merging current master (`45c99a0d0`). - Targeted callback integration suite with disposable PostgreSQL: 4 tests passed before rebase. - Full `pnpm -r typecheck` and `pnpm build` passed after merging current master. The full local suite passed 12,668 tests; one suite failed to start its disposable PostgreSQL. Rerunning that suite alone passed all 31 tests. - The built server resolver selected the executable under `server/dist/vendor/paperclip-runner/bin/`. - Final callback regression: all 4 targeted integration tests pass, covering provider header rewrites, default ports, uppercase/trailing-dot hosts, and ignored self-hosted hints. - Live staging proof of the runner fix: a previously failed Slack thread recovered, a new mention received its requested response, and the account-connect command succeeded. Unsigned callbacks returned 401. - Live browser and Slack acceptance passed: generated callback URLs, account linking, a new mention, an interactive question and answer, and all three callback-health indicators. The connector was activated through the onboarding UI. - All 54 latest-head checks pass, two optional checks are skipped, and Greptile is 5/5 with no unresolved threads. ## Risks - Runner fallback must select a binary for the remote platform. This preserves explicit remote artifact overrides and the existing capability checks. - Proxy headers are not identity proof. They are used only for diagnostics on claimed Cloud instances after the provider accepts the request. Self-hosted instances ignore them. Wrong public hosts and ports still warn. - No database migration or new public API contract. ## Model Used OpenAI GPT-6 via Codex. Used reasoning, repository tools, code execution, and browser/native-app testing. Exact model variant and context window were not exposed by the session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
45c99a0d06 |
fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Legacy adapters launch provider CLIs and expose connected tools. > - Existing defaults did not consistently grant full automatic permission. > - Remote Claude used a fixed tool list that omitted MCP tools and future tools. > - Direct Codex launches and OpenCode configuration also used narrower defaults. > - This change gives all these paths the same full-auto default as native runners. > - Explicit restrictive settings continue to work. ## Linked Issues or Issue Description Refs #13686. This PR is stacked on that native-runner and task-reassignment PR. Merge #13686 first. Related: #831 (constructed Claude agents), #1935 (adapter-switching permission defaults). ## What Changed - Use actual Claude permission bypass for local and remote runs and probes. Remove the fixed tool list so MCP and future provider tools are included. - Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`. Do not mark ordinary host execution as a sandbox. - Default direct Codex execution to approval and sandbox bypass, matching agent creation. Preserve explicit false, CLI profiles, sandbox modes, approval policy, and network restrictions. - Set OpenCode's full-auto runtime permission to `allow` for every tool and connection. Preserve the existing explicit opt-out. - Default Gemini probes to the same YOLO mode as execution. Make the legacy ACP `default` alias use `approve-all` for fresh and resumed sessions. - Add default, opt-out, remote, probe, connected-tool, and resume regression tests. Update adapter configuration documentation. - Other adapter paths already request full automatic permission or have no provider approval gate. ## Verification - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server and legacy ACP tests passed, including defaults, explicit opt-outs, remote launches, connected tools, and fresh/resumed sessions. - Greptile reviewed current head `8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security reviewer acknowledged the documented full-auto requirement. Acknowledged discussions are resolved. - Current head has **54 passing checks**. [PR checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The process-adapter signoff browser shard passed on one retry after its first attempt exceeded a three-second issue-run wait. - **Six native Claude/Codex real-provider cases passed on their first attempt, with cleanup passing**, against the combined branch: plans, reassignment, and backlog creation/status. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). This does not claim a real-provider run of every legacy adapter. - The live-tested revision is `a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs only in the corrected heartbeat test expectation; application code is identical. - The campaign result-enforcement job passed. The separate report publisher failed during frozen dependency installation because the trusted workflow's patched-dependency configuration does not match its lockfile. Passing case evidence remains downloadable from the workflow. - Full-suite coverage comes from CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. ## Risks - Missing permission settings now grant all provider operations, including connected tools. OpenCode full-auto also overrides ambient provider permission rules. An explicit Paperclip permission opt-out preserves restrictive behavior. - Claude refuses full bypass as root outside an identified sandbox. Ordinary host deployments must run Claude as a non-root user. Managed sandbox launches include the required marker. - These defaults do not grant additional Paperclip roles, connections, or company access. Existing controller authorization and governance still apply. - This PR depends on #13686. Retarget it to master after that PR merges. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7bc03e0acd |
feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agent Chat uses native runners to save plans and coordinate tasks. > - Provider defaults differed across harnesses and could stop unattended work at a second permission gate. > - Agents also lacked a dedicated tool to move existing work to another agent safely. > - This change defaults native providers to full automatic permission for provider tools and connected tools. > - A guarded reassignment tool preserves task identity, stops the previous run, and schedules the new owner once. > - Codex and Claude chat acceptance tests now use production permission defaults. ## Linked Issues or Issue Description **Subsystem affected** Native runner, ACPX Claude permission policy, task authority, and Agent Chat acceptance tests. **Problem or motivation** A user can authorize an agent to save a plan or create a task, but Claude's default provider gate can still stop that action. Reassignment needs a dedicated operation that preserves context and avoids concurrent owners or unintended recovery runs. **Proposed solution** Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`. Apply the defaults at configuration, execution, fresh-session, resume, driver, and proxy boundaries. Keep explicit permission settings and server-side company, claim, task-mode, and approval checks. Add `reassign_task` with version checks, durable idempotency, audited cancellation, and guarded successor scheduling. **Alternatives considered** A Paperclip-only allowlist still blocks provider tools and other connections during unattended work. Full automatic permission is the requested product default. Recreating a task discards its identity and history. Updating assignment without stopping the previous run can leave two agents working on the same task. **Roadmap alignment** This extends the existing planning, delegated work, governed tool access, and recovery features. It adds no new service or schema migration. Recent related tasks and open PRs were checked for duplicate work. **Additional context** Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner startup). The stacked legacy-adapter companion is #13693. This also fixes the deployed-server artifact fallback needed to stage the current runner binary. ## What Changed - Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to `never`, including missing settings at direct driver and proxy entry points. These defaults cover provider tools and connected tools. Preserve explicitly configured restrictive modes. - Include assigned approval reads using canonical side-effect classifications, so verifying a recorded approval does not trigger another provider gate. Paperclip approval decisions still enforce controller authority. - Carry the new permission mode through server configuration, execution contracts, recovery identity, TypeScript, and Rust. Keep `approve-paperclip` as an optional restricted mode, with exact SDK rules and closed unknown requests. It is not a default. - Add `reassign_task` to the semantic catalog, controller, mock authority, and generated contracts. - Guard reassignment with company authorization, expected owner and version, protected-state checks, and durable retry receipts. - Honor explicit backlog task creation atomically with the initial plan, without scheduling a wake. Preserve backlog holds regardless of dependency readiness. - Stop active work before changing ownership. Restore the prior owner through a guarded, idempotent wake if final handoff validation fails. Keep intentional reassignment stops out of failure recovery. Preserve backlog and blocked states without waking them early. - Add authorization, concurrency, replay, stop, and permission boundary regressions. Add Codex and Claude chat reassignment cases and run native chat cases with production defaults. - Clarify shared runner guidance: save plans and Paperclip documents directly with `write_document`; create and register a local file only when a downloadable file is requested. - Document provider defaults and the operator choices for existing agents. ## Verification - Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks passed**, with two intentional skips. [PR checks](https://github.com/paperclipai/paperclip/pull/13686/checks). - Greptile reviewed that exact head at **5/5**. The security reviewer acknowledged the intended full-auto default, and the acknowledged discussions are resolved. - Full workspace `pnpm -r typecheck` and `pnpm build` passed locally after rebasing onto current master. Targeted adapter/server, runner, API, default/resume, and heartbeat configuration tests passed. - **All six real-provider acceptance cases passed on their first attempt, with cleanup passing:** plan handoff, task reassignment, and backlog creation/status, each on native Claude and Codex. Evidence records Claude's effective `approve-all` mode. [Campaign and downloadable evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548). - The live campaign tested combined revision `a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only a heartbeat test expectation correction; application code is unchanged from that live-tested revision. - The campaign's result-enforcement job passed. Its separate report publisher failed because the trusted workflow's `patchedDependencies` configuration differs from its frozen lockfile. All six results and screenshots remain available as GitHub artifacts. The overall manual workflow is red for this publishing failure. - Full-suite coverage is supplied by the passing CI partitions. The separate unsharded local run was stopped after the corresponding CI partitions passed; it is not counted as a completed local run. - Reassignment tests cover stale state, cross-company access, denied authority, cancellation failure, compensating wake, and idempotent retries. Backlog tests verify the original creation audit, saved plan, exact task count, and absence of task-bound runs. ## Risks - Agents with no explicit permission mode now receive full provider tool permission, including connected tools. This is a deliberate broad default. Existing explicit restrictive modes still apply. Controller authorization, company isolation, workspace boundaries, and Paperclip governance remain in force. - Reassignment crosses run cancellation and task ownership transactions. Durable stop intent, revalidation, audit receipts, and guarded queue dispatch cover interruptions and retries. - The new permission enum requires a current runner artifact. The remote artifact fallback uses the same resolved controller binary for upload and execution. - Live provider behavior remains subject to the selected model. Targeted live results do not qualify the full catalog. ## Model Used OpenAI Codex, based on GPT-6, with code execution and repository tools. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
04546c82d5 |
fix(runner): reconnect Daytona sessions after controller restart (#13691)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner can execute a task inside a Daytona sandbox. > - The sandbox can keep running when the Paperclip controller restarts. > - Recovery treated sandbox process IDs as local process IDs and selected the wrong recovery path. > - Live verification also found races between startup, shutdown, and queued task cleanup. > - This pull request verifies the existing remote owner and orders those transitions. > - Users can continue the same task and provider session after a controller restart. ## Linked Issues or Issue Description **What happened?** The Daytona `recover-controller` cases failed with `runner_state_identity_mismatch`. Remote process IDs can be absent on the controller or collide with unrelated local processes. Recovery then looked for remote state in the local runner directory. Later turns could also start before the previous executor released its sandbox resources. **Expected behavior** Reconnect to the original sandbox and authenticated runner. Preserve the task, provider session, and queued comments. Reject a replacement sandbox or mismatched identity. Do not start another provider during reattachment. **Steps to reproduce** Run the `everyday-workflows` `recover-controller` case for `runner-codex` or `runner-acpx-claude` in Daytona. The browser creates a Python tool, requests a revision, restarts the controller during execution, and queues another revision. It then downloads and tests the final ZIP. Related: #13682 is the preceding operational fix. #13291 addresses legacy sandbox conversation recovery, a different execution path. #13666 includes broader run-capacity work; this change guards cleanup of an existing native task executor. ## What Changed - Add remote runner recovery without interpreting sandbox PIDs on the controller. - Verify the original provider lease, remote workspace, durable state, process marker, and authenticated PRP authority before adoption. - Compare the process marker with live Linux boot identity and start ticks to reject PID reuse. Read virtual proc files through the guaranteed Node runtime; unavailable proof blocks adoption without blocking a fresh launch. - Make the E2E supervisor own the actual server process so forced restart cannot leave a late database closer behind. - Scope the chat delivery lease test to its own fixture instead of draining other tests’ pending deliveries. - Preserve provider-attempt counts and recorded evidence during reattachment. - Serialize an idle-session checkpoint with admission of the next native turn. - Wait for an in-progress startup to acknowledge restart detachment. Fail after a bounded deadline if it cannot. - Keep a queued comment waiting until the previous native task executor releases its resources. Allow unrelated tasks to continue. - Update the Daytona image's resolved lock digest to match current dependency manifests. - Add classifier, ownership, process, startup, checkpoint, and queued-admission regression tests. Document recovery behavior. ## Verification - 415 focused tests passed across native execution, restart recovery, workspace synchronization, queued admission, and real-process restart tests. The final Node-based fingerprint change passed all 375 native-session tests. - Runner harness unit tests: 394 passed. Chat integration shard 2: 335 passed after fixture isolation. - The exact fingerprint command succeeded twice in a disposable Daytona sandbox and returned the same identity; the sandbox was deleted. - 11 real-process restart integration tests passed, including absent and colliding remote PIDs. - Repository typecheck and final build passed. Broad local checks found machine-dependent database startup and timing failures; focused retries passed. The final-revision PR pipeline is green. One unrelated browser shard hit a five-second blank-page timeout on the first run and passed its targeted retry. - Final-revision local headed browser E2E: `everyday-workflows.runner-acpx-claude.daytona.recover-controller` passed on attempt 1 in 4.7 minutes, **40/40 checks**. Manual browser inspection confirmed Done, all three ZIPs, and delivery of the queued follow-up. All three runs succeeded using the same provider session. The harness downloaded and independently tested the final artifact. - Final-revision Daytona campaign: https://github.com/paperclipai/paperclip/actions/runs/35463999611 — **Codex passed first attempt (4.8 minutes); ACPX Claude passed first attempt (6.1 minutes)**. Campaign aggregation/publication is finishing; both test jobs succeeded. - Greptile reviewed `beb08d8493b3286f5bb988dead369ff8c96a395d`: **5/5**, no open findings. - Staging browser verification is pending selection of a disposable staging instance and removal of a Chrome extension UI block. ## Risks - Recovery now depends on the original sandbox remaining available. A replacement or mismatched identity still blocks adoption. - Shutdown waits up to 30 seconds for a native startup to reach a safe detach point. An unfinished startup returns a clear failure instead of a false detach receipt. - Queued native work on the same task waits for cleanup. Unrelated tasks remain eligible. - The image digest update rebuilds the Daytona runtime image. No database migration or public API change is included. ## Model Used OpenAI Codex, GPT-6, with repository inspection, code execution, and browser tools. The runtime does not expose the exact deployed model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da257c3069 |
Warn when routine webhook URLs may not be publicly reachable (#13684)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Routines can start that work when another app sends a webhook. > - Local and private URLs often cannot receive events from public services. > - HTTPS alone does not make a Tailscale address public. > - This pull request explains these limits during setup and editing. > - Users can still finish setup for senders on their own network. ## Linked Issues or Issue Description Refs #13637. Webhook setup needs a clear warning when the generated URL appears local, private, or unencrypted. The warning must explain how to make the endpoint reachable without blocking private-network use. ## What Changed - Add a shared warning banner to the Connect, Check connection, and Edit webhook views. - Distinguish localhost, private network addresses and domains, HTTP, and Tailscale hostnames. - Explain the difference between Tailscale Serve and Funnel. Link to the Paperclip HTTPS guide. - Add five full-page Storybook examples, design guide examples, and documentation. - Add URL classification tests and a regression test that finishes setup despite the warning. ## Verification - Passed 41 focused URL and trigger-flow tests. - Passed workspace typecheck, workspace build, token gates, and Storybook build. - Browser-tested the Tailscale story through Check connection, Finish setup, and Edit webhook. The warning stays visible and does not block setup. - Open Product / Routines / Webhooks stories 11–15 to review the warning states. - All 54 PR checks passed, including the full test matrix and eight browser shards; two optional Storybook jobs were skipped by workflow policy. - The server supervisor readiness test timed out once in CI, then passed on rerun and locally (6 tests). - The duplicate local full-suite run was stopped after the complete CI matrix passed. - Greptile: 5/5 on the current commit, with no unresolved review comments. ## Risks - URL checks are hints. They do not test DNS, firewall rules, or actual reachability. - A Tailscale hostname can serve either private Serve traffic or public Funnel traffic. The warning explains this uncertainty and permits both. - No API, schema, authentication, or webhook delivery behavior changes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
64895f187b |
fix(runner): clarify completion errors and restart test failures (#13682)
## Thinking Path > - Paperclip manages work across persistent agent sessions. > - The runner validates completion calls before accepting their results. > - Generic validation errors can leave the agent unable to repair a rejected call. > - Restart tests also exempted every later failure on an intentionally interrupted run. > - This change gives bounded schema feedback and limits the test exemption to expected interruption outcomes. > - Failures become easier to repair and diagnose without changing authorization or task prompts. ## Linked Issues or Issue Description Refs #13674 and #13676. Related environment and Agent Chat fixes landed in #13677 and #13678. Those changes do not cover these diagnostics. **What happened?** A malformed completion call received a general field list without the failed schema location. The everyday restart test hid later adapter errors on an intentionally interrupted run until its deadline. A clean pnpm install also broke the shutdown test because it resolved an undeclared Playwright package. **Expected behavior** Return enough schema information to repair completion calls without returning submitted values. Fail promptly on an unexpected recovery error. Resolve the declared test package's CLI. **Steps to reproduce** Run the new completion-validation and everyday lifecycle regressions against the parent commit. The new assertions fail there. Run the shutdown test in a clean workspace installation. **Paperclip version or commit** Based on master |
||
|
|
f589660ec0 |
feat(routines): add safe webhook setup and in-routine run management (#13637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Routines turn scheduled work and external events into tasks for an assigned agent. > - Webhook setup was disabled, and actor authentication rejected valid webhook bearer keys. > - Operators need to connect and test a sending app before events can start work. > - This pull request adds a guided setup with durable connection tests that cannot dispatch a task. > - It keeps trigger management, execution tasks, and activity within the routine. > - The benefit is a webhook that can be configured, verified, and operated from one place. ## Linked Issues or Issue Description Fixes #11937. Related: #13216 adds provider-specific Sentry support. This PR addresses general routine setup and ingress. #6841 addresses legacy secret bindings; this PR retains the existing secret service. **Current behavior** Webhook creation is disabled. Bearer deliveries can fail in agent authentication before the routine checks its key. Setup has no safe connection test. Runs and Activity send the operator away from the routine. **Proposed behavior** Choose a schedule or a webhook. Follow the setup steps, copy credentials or complete agent instructions, and test delivery without creating work. Finish setup to allow future events to start tasks. Edit or remove compact trigger cards, undo removal, and inspect tasks and activity inside the routine. **Reason and benefit** An operator can verify credentials and delivery before enabling automatic work. Durable setup state survives refreshes and restarts. Retry receipts prevent an old test event from starting work after activation. ## What Changed - Add a production trigger wizard using reusable Slack setup navigation and footer components. - Add schedule and webhook choices, one-time credentials, agent instructions, and live connection feedback. - Persist pending setup, test delivery receipts, connection status, and reversible trigger removal. - Keep setup checks free of routine runs, tasks, and agent wakeups. Preserve delivery idempotency after activation. - Add compact trigger cards, inline editing, key rotation, pause controls, removal, and Undo. - Keep Runs and Activity in the routine. Use the shared task list and compact activity rows. - Permit only exact public delivery POSTs through actor authentication. Retain webhook authentication, JSON-object validation, and log redaction. - Add production-backed Storybook states and focused server, database, and UI coverage. - Document signing modes, setup checks, retries, rotation, HTTPS ingress, and navigation. ## Verification - Full workspace typecheck, build, and token gates passed on the rebased branch. Storybook also builds. - Focused routine, middleware, logging, shared wizard, and UI coverage passes on the rebased branch: 195 tests across 14 files. The migration passed on a fresh PostgreSQL database and on two repeated applications. - Browser testing used the real app, database, and a deterministic process worker through Tailscale HTTPS and the current Cloud proxy code. - Verified rejected keys, safe setup deliveries, persisted state after restart, activation, retry deduplication, key rotation, schedule editing, removal, and Undo. - Fresh bearer and GitHub-signed deliveries created tasks that the worker checked out and completed. Runs and Activity stayed within the routine. - Current Cloud ingress tests passed. Public delivery POSTs passed through without a browser session; management routes remained gated. - All 54 current-head PR checks pass, including general and serialized tests, all eight browser E2E shards, typecheck, build, runner checks, security checks, and the canary dry run. Two optional Storybook jobs are skipped by workflow conditions. - Greptile is 5/5 on commit `7ea63a61e`, with no unresolved review threads. The stale connection-status finding is fixed and covered by a regression test. - No production deployment was performed. ## Risks - Migration 0281 adds three trigger columns and a test-receipt table. It is additive and safe to reapply. Apply it before running the new server. Existing triggers remain live by default. - Requests without delivery IDs are new events after activation. Senders must reuse an event's delivery ID for retries. - Completed webhooks keep normal dispatch behavior. Their management connection check can start work; the UI states this. - Removing a trigger archives it. Undo restores the URL and credentials. Permanent deletion remains available through the existing API. - Public ingress must remain restricted to the delivery POST route. The tenant verifies credentials. Cloud sleeping-stack behavior is unchanged. - Shared setup components also serve Slack. Existing setup contracts and navigation tests cover that integration. - Senders must use application/json with an object. Other media types receive 415. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
20d26117c9 |
fix(chat): use the claimed Cloud origin for connector URLs (#13680)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connectors publish callback URLs and links to the board. > - Cloud can assign a warm instance its final origin after the server starts. > - The signed runtime identity already tracks that change. > - The chat service kept a copy of the startup origin and continued to publish it. > - This pull request resolves the trusted origin when it creates each URL. > - New connector setup uses the claimed hostname without a server restart. ## Linked Issues or Issue Description **What happened?** Chat setup in a claimed warm instance used its old pool hostname in provider callbacks. Account confirmation and task links could also use the old hostname. **Expected behavior** Chat URLs follow the signed canonical origin after the claim. An explicit webhook ingress override still applies only to provider callbacks. Self-hosted URL precedence stays the same. **Steps to reproduce** 1. Construct the chat service with a pool origin. 2. Apply the Cloud claim without restarting the service. 3. Open Slack setup or create an account-linking intent. 4. Observe the startup hostname in the returned URL. **Paperclip version or commit** Reproduced on master at `9335b7db1`. **Deployment mode** Paperclip Cloud warm-instance claim. Related: #12766 introduced the signed canonical runtime identity. ## What Changed - Resolve the signed Cloud origin when building chat setup, account confirmation, and task URLs. - Use the same callback origin for Telegram registration and GitHub webhook recovery. - Preserve explicit webhook ingress and self-hosted configuration precedence. - Add regression coverage for existing and new endpoints across Slack, GitHub, Teams, and Telegram. - Document the origin precedence and the need to update callbacks already saved at a provider. ## Verification - Reproduced both new regression cases against the original code. - Full chat integration and signed Cloud identity suites: 1,015 tests passed after the production-code correction. - Five origin and ingress cases passed after review additions, including Telegram registration and GitHub webhook repair after a live claim. - Focused origin, ingress, callback, task-link safety, and tenant-isolation checks: 67 passed. - `pnpm -r typecheck` and `pnpm build` passed. Server typecheck and compilation passed again after the task-link validation correction. - A broad local `pnpm test:run` started before the correction was stopped after the final-commit CI suite passed. It is not counted as a passing local run. - Final-commit CI: 54 successful checks; two optional Storybook checks skipped. Greptile: 5/5 with all review threads resolved. - No live deployment or Slack app mutation was performed. ## Risks - Cloud chat URLs now follow the signed runtime identity. Request host headers cannot set this value. - An explicit webhook ingress override still takes precedence for callbacks. - Existing Slack app settings are external state. Operators must replace an old callback URL in Slack. - This change does not deploy the app or change gateway ingress policy. No schema migration is required. ## Model Used OpenAI Codex (GPT-6), with repository search, code execution, and automated tests. The exact model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c1f6c3310a |
fix(runner): repair catalog runtime and grading boundaries (#13676)
## Thinking Path > - Paperclip manages tasks across persistent agent sessions. > - The full Runner E2E catalog exposed failures in session restoration, tool validation, and test controls. > - These failures prevented valid work from resuming or made a valid interaction fail the test. > - Invalid completion reports also reached finalization before the provider received useful feedback. > - This pull request repairs those boundaries without changing production prompts or approval policy. > - Focused regressions and fresh paid cases verify each fix. ## Linked Issues or Issue Description Follow-up to #13655. Stacked on the trusted worker prerequisite fix in #13674. **What happened?** Read-only skill uploads failed in resumed Daytona sandboxes. Invalid criterion IDs escaped tool validation. A progress event could park a run before its tool response settled. Partial question forms hid required answers. Two test assumptions rejected valid plan keys or failed to navigate an optional question page. **What did you expect to happen?** Resume identical skill bundles, give repairable feedback for malformed completion calls, preserve in-flight tool responses, show all required questions, and test the rendered workflow accurately. **Steps to reproduce** Inspect the failed cases in https://github.com/paperclipai/paperclip/actions/runs/35417932353. Fresh campaigns: https://github.com/paperclipai/paperclip/actions/runs/35444497313 and https://github.com/paperclipai/paperclip/actions/runs/35445327618. The later backup cleanup is tested in https://github.com/paperclipai/paperclip/actions/runs/35446477285. Combined report: https://pages.paperclip.ing/runner-e2e-operational-35444497313/investigation.html. ## What Changed - Compare immutable archives before reusing read-only Daytona bundles. Reject corrupted content and preserve unrelated files. - Validate exact criterion IDs before accepting completion. OpenCode returns a tool error instead of emitting a result that terminates runnerd. - Complete the activity item for rejected OpenCode calls. - Remove retired read-only harness backups without altering live files or following symlinks. A fresh paid rerun exposed this later checkpoint-cleanup failure. - Exclude progress messages from the governed-wait completion boundary. - Reject newly created question forms that omit questions or contradict their stored answer semantics. Keep historical rows readable. - Navigate all rendered question pages and recognize revision-bound descriptive plan keys in the continuation suite. ## Verification - Harness unit suite: 383 tests pass. Harness typecheck passes. - Native session executor and status corpus: 381 tests pass. - Shared question and interaction-service tests: 42 pass; native question bridge and executor: 360 pass. Daytona sync: 21 pass, including foreign-owner archives and corrupted immutable content. - OpenCode driver: 29 tests pass, including wrong, missing, and duplicate criterion IDs followed by a valid retry. - Repository typecheck and build pass. The later OpenCode activity fix also passes its package build. - The latest commit passes all 52 PR checks and Greptile 5/5. The backup-cleanup fix also passes 351 related local tests and server typecheck. Local full-suite coverage completed across runs. adapter-auth-signal-routes and pipelines-routes encountered transient socket resets; both pass on retry, and all remaining 24 serialized files pass. Paid reruns are complete: 27 of 29 unique cases pass using the latest recording per case. Both Daytona controller-restart cases still fail with runner_state_identity_mismatch; the report describes this remaining runtime issue. Eight affected cells need #13674 on master before their rerun. ## Risks Creation rejects inconsistent dual question representations but does not change historical records. Immutable bundle comparison must verify bytes before skipping extraction. Completion feedback must use the contract bound to the current run. Durable suspension and approval checks remain enforced. Production prompts are unchanged. ## Model Used OpenAI GPT-6 via Codex, with repository inspection, code editing, and test execution. The exact API model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4cfc0e3f6c |
fix(runner-e2e): provision complete worker prerequisites (#13674)
## Thinking Path
> - Paperclip manages work performed by AI agents.
> - Runner E2E tests verify that work in local and remote environments.
> - The full catalog exposed setup failures before agents could execute
tasks.
> - ACPX Codex missed sandbox provisioning, and new artifact cases
missed Docker preparation.
> - A host Claude version probe also stopped native Claude stories that
use a packaged provider.
> - This pull request repairs trusted worker setup and keeps its
selectors covered by catalog tests.
> - The resulting reruns can measure behavior instead of missing
prerequisites.
## Linked Issues or Issue Description
Follow-up to #13655. Full-catalog campaign:
https://github.com/paperclipai/paperclip/actions/runs/35417932353.
**What happened?**
ACPX Codex could not start its sandbox. New artifact stories missed
Docker preparation. Native Claude stories tried to spawn an unrelated
host CLI. The report job failed on trusted lockfile drift.
**What did you expect to happen?**
Prepare each worker's required capabilities before paid execution and
publish the retained results from trusted code.
**Steps to reproduce**
Run local ACPX Codex, everyday agent-review-handoff, or native Claude
everyday cells in the full-stack workflow at
|
||
|
|
36dbb7ed1c |
fix: harden agent chat runner tools and recovery (#13678)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat turns discussion into plans, tasks, reviews, and hires. > - These workflows need reliable tool results and task context on the native runner. > - Live Claude and Codex tests exposed lost retry requests, invalid project inputs, and a child startup crash. > - Recovery also exposed a misleading retry action and missing child task context. > - This pull request fixes those paths and adds regression coverage. > - Agents can continue the original request and operators can inspect a stopped run. ## Linked Issues or Issue Description **What happened?** A failed Agent Chat retry could lose the user's question. Project creation accepted unsupported icons in its tool schema. Codex could stop when a helper's MCP startup event arrived before its thread lineage. A stopped task offered Retry even when the server required execution reconciliation. Resumed agents could miss existing delegated tasks. Hiring and review instructions did not describe the native runner's available tools and source requirements. **Expected behavior** Retries retain the selected request. Tool schemas match the API. Child startup information does not gain authority over the parent or stop it. Recovery actions match the server's requirements. Task context exposes existing child work. Handoffs contain the material the assignee needs. **Steps to reproduce** 1. Enable experimental Agent Chat in an isolated development instance. 2. Configure native Codex and ACPX Claude agents on Paperclip Runner. 3. Ask for a plan, revise it, approve task creation, and request a hire and status report. 4. Retry a failed chat turn and check that it answers the original request. 5. Start a Codex helper before its thread lineage arrives. 6. Resume a delegated task and inspect its existing children and saved output. **Paperclip version or commit** The live failures were found at `f2c5e54dc`. This branch is rebased onto `86b7ee992`. **Deployment mode** Isolated local development instance with native Codex and ACPX Claude. No database migration or default permission change. Related work: Refs #13284 for Agent Chat. Refs #13438 for the server-side API receipt fix, which this branch preserves. The transport also accepts the earlier HTTP receipt format. Refs #13655 for the current Codex continuation and helper lineage handling, which this branch also preserves. ## What Changed - Preserve failed Agent Chat wake-comment IDs and session generation from the authorized source run. Reject pre-reset retries. - Wrap API receipts with the correct semantic call identity. Test current and earlier receipt formats through real HTTP and runnerd. - Classify early child MCP startup notifications as information. Keep foreign completion and result events rejected. - Constrain project icons on both tool surfaces and regenerate the protocol contracts. - Include bounded, company-scoped visible direct child tasks in task context. Filter hidden tasks before applying the limit. - Replace the rejected Retry action with Inspect run for native continuation reconciliation. - Update hiring, review handoff, status reporting, and development guidance. ## Verification - Live tests covered Claude and Codex questions, plan revisions, approval, task creation, hiring, status, chat reset, failures, and recovery. - The recovered task produced its saved checklist and example. A later follow-up read the existing child tasks and document without creating more work. - Full build, repository type checks, token gates, 142 focused tests, 188 runner TypeScript tests, and the Rust notification/descendant regressions passed after rebase. The separate local full-suite run was stopped after the complete CI suite passed. - Review fixes passed the updated route, tool-authority, and icon regression tests plus server type checking. - Required commands: `pnpm build`, `pnpm -r typecheck`, `PAPERCLIP_IN_WORKTREE=false pnpm test:run`, and `pnpm check:token-gates`. - At `4ce8047b0`, all 55 applicable GitHub checks pass (two Storybook checks are intentionally skipped), including the complete general/serialized test matrix, runner tests, browser tests, build, type checks, Docker checks, and canary dry run. - Fresh Greptile review is 5/5 on `4ce8047b0`; all three findings were fixed with regressions and there are no unresolved review threads. - Two initial CI service-startup timeouts passed unchanged in local reproductions and in the latest CI run. ## Risks - The new event classification is limited to MCP startup information. It does not authorize foreign task completion, results, or tool requests. - Task context returns at most 100 direct child tasks and reports truncation. It excludes hidden tasks and other companies. This improves delegation context but does not enforce semantic duplicate detection. - Native reconciliation still requires an operator to inspect and record prior outcomes. The new link does not replace the recovery API. - API tools remain opt-in. Claude permission choices remain explicit. No default permission, schema, or workflow changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository editing, code execution, API tools, and browser testing. The exact deployment identifier and context-window size are not exposed in this session. Live acceptance agents used OpenAI `gpt-5.6-sol` and Anthropic `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9335b7db10 |
fix(runner): validate inherited environments and replace stale sandbox binaries (#13677)
Resolve the effective environment for account adoption and adapter tests. Preserve saved-agent overrides when the request omits environmentId, and treat explicit null as inheritance from the instance. Reject sandbox runners that lack unlimited-runtime and connection-lease-renewal capabilities. Stage the bundled runner before launch when the image binary is stale. Add regression coverage for environment precedence, fail-closed validation, adapter switches, API-key reverification, and runner artifact fallback. Document the operational workaround for older controllers. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c9e8677979 |
docs: document chat connector UX and make the runbook self-contained (#13675)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - Connections let agents work with external services. > - Contributors use the connection runbook to add and review providers. > - The runbook depended on private issue references and did not capture the chat setup UX conventions. > - This pull request adds a chat connector UX guide and puts the missing requirements in the runbook. > - Contributors can apply the guidance without access to the internal issue tracker or a personal skill installation. ## Linked Issues or Issue Description **Issue type** Missing documentation and unclear contributor instructions. **Where is the issue?** `doc/connections/CONNECTOR-PLAYBOOK.md` and the chat connector setup guidance. **What's wrong?** The runbook sent contributors to private issues for validation, OAuth ownership rules, and catalog review. The Slack setup work also established useful UX rules that other providers should share. **Suggested fix** Add a companion UX document. Link it from the runbook. Include validation and review requirements directly in the public documentation. Use portable company and instance examples. Related implementation: https://github.com/paperclipai/paperclip/pull/13638. A search of related PRs found no duplicate documentation change. ## What Changed - Add `CHAT-CONNECTOR-UX.md` with setup, credential, identity, footer, test, and management conventions. - Include adaptation examples for Discord, Telegram, and email providers. - Link the companion from the runbook introduction, contents, and UX section. - Replace private issue references with inline architecture boundaries, risk classification, validation evidence, and per-tool review requirements. - Replace personal deployment examples with sample company and instance addresses. - Mark the recorded Notion provider observations as a dated snapshot. ## Verification - Passed `git diff origin/master --check`. - Passed local validation of all 35 relative links and heading anchors in the two documents. - Passed code-fence and private-reference scans. - Reviewed the guide against the Slack setup decisions and the existing connection docs. - Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. They could not complete because this fresh worktree has no installed dependencies (`@types/node` and the CLI `tsx` entry point are missing). No application code changed. - GitHub CI passed on commit `070fefc8b2513cb9469bae76e66940c33c1465ea`, including typecheck, build, general and serialized tests, runner verification, browser tests, and canary dry run. - Security checks passed. Greptile scored the current commit 5/5 with no findings or unresolved review threads. ## Risks Low risk. This changes two documentation files only. Provider APIs and screens can change. The guide requires authors to verify provider capabilities, and the Notion example identifies its observation date. The documentation does not assert new runtime support. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context window are not exposed in this session. Used reasoning, repository inspection, shell execution, and documentation editing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes` / `Refs` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal issue id or instance-derived details - [x] I have run the applicable documentation checks locally and they pass; application checks were attempted and the environment limitation is recorded above - [x] I have added or updated tests where applicable (documentation validation; no runtime changes) - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1ef3b08714 |
feat(ui): integrate agent personas across the app (#13171)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A stable agent persona is useful only when the same identity appears across the app. > - Lists, task messages, selectors, and activity feeds need inexpensive static avatars. > - Onboarding and agent headers need a larger character with expressions and pointer tracking. > - This pull request connects the persona foundation to those existing views and preserves onboarding draft assignments. > - Full-page stories and Linux checks make the placements and performance contract reviewable. ## Linked Issues or Issue Description **Problem or motivation** Agents need a stable visual identity in lists, tasks, onboarding, and configuration. External tools also need an image URL for that identity. **Proposed solution** Assign each agent a permanent palette from a fixed ClipLab character library. Store the assignment on the agent. Render and cache preset PNG URLs on demand. Use static images in dense views and one animated character in larger placements. **Alternatives considered** A generated image bundle requires a separate asset build. A live renderer in every avatar adds unnecessary work in large lists. Arbitrary uploaded images do not provide the requested shared character system. **Roadmap alignment** This improves agent identity across existing control-plane views. It preserves agent permissions, company boundaries, and status labels. ROADMAP.md has no separate ClipLab persona milestone. Related approaches: #2422 adds configurable image URLs and DiceBear generation; #5578 adds optional uploaded avatars. This work uses a fixed, versioned character library and preset URLs. ## What Changed - Replace agent icons with static persona images across lists, the sidebar, org charts, tasks, comments, selectors, activity, and dashboard views. - Put one animated character in the agent header. Let it follow the pointer across the page, with reduced-motion and touch fallbacks. - Add larger padded characters to agent creation. Keep the palette stable across draft refreshes and connection retries, then reveal it after success. - Pass appearance through shared projections rather than fetching each agent separately. - Add real full-page Storybook examples for the agent list, overview, task, dashboard, new-agent dialog, and connection page. - Add Linux screenshot, clipping, density, and 500-avatar performance checks. ## Verification - `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased tree. Persona lifecycle tests pass. - The rebased feature passes 38 Linux screenshot/performance checks, including both display densities, corner pointer positions, and the no-WebGL/no-live-download contract for 500 avatars. - The final Linux persona suite passes all 38 visual, lifecycle, density, and full-page checks using the standard Storybook configuration and real on-demand avatar endpoint. - Final local focused verification: 45 avatar/native-recovery tests pass; UI identity/routine tests, typecheck/build, token gates, and Storybook build pass. - Current-head CI passes: full workspace/server tests, all serialized server groups, typecheck/release checks, build, canary validation, and end-to-end shards. The build passed after retrying a native-runner concurrency-test failure; its three targeted cases also pass locally. - Manual inspection covered stable identities in the app, header placement, full-page mouse tracking, onboarding size, and task/dashboard placements. ### Screenshots Linux captures use synthetic Storybook fixtures. Full-page captures use reduced motion. The live character, mouse tracking, and disposal are checked separately. <details> <summary>Agent overview with the character in its header</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png" width="900" alt="Agent overview with the character in its header" /> </details> <details> <summary>Task messages and assignee identity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png" width="900" alt="Task messages and assignee identity" /> </details> <details> <summary>Larger onboarding character with room for expressions</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png" width="900" alt="Larger onboarding character with room for expressions" /> </details> <details> <summary>Dashboard agent activity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png" width="900" alt="Dashboard agent activity" /> </details> ## Risks - This PR depends on #13170, the persona foundation. Merge the foundation first, then retarget this PR to master. - Many placements change from icons to character silhouettes. Human avatars and authoritative agent status labels retain their existing behavior. - Only one character can render live per view. Reduced motion, hidden/offscreen content, touch input, and renderer failures use the defined fallbacks. - The full-page stories use fixture data. They do not contact a real company or complete real provider sign-in. ## Model Used OpenAI Codex, GPT-6 family. The exact model identifier and context window are not exposed in this session. Used code editing, shell execution, browser inspection, and Linux visual testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Tonio <tonework@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
43acbcc398 |
fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner connects task state to provider sessions. > - Follow-up turns must retain provider memory and carry new user direction. > - Lost session IDs caused repeated context and extra input tokens. > - Native question answers and approval races could leave valid work blocked. > - This pull request repairs those paths and adds regression coverage. > - Agents can continue accepted work without repeating the conversation or losing the user's answer. ## Linked Issues or Issue Description Refs #13574. That merged PR shortened continuation prompts and moved question instructions into tool documentation. This change preserves sessions and fixes failures exposed by broader testing. Related runtime work: #13408 and #13410. **What happened?** Native follow-up turns could lose the provider session ID. Completion guidance could replace the original task with its latest comment. Claude native questions could remain pending after the user answered. Approval during a running tool call could suspend the run before the tool response arrived. Onboarding and chat handoff instructions also caused repeated planning or missing plan documents. **Expected behavior** Reuse a valid provider session. Send only new events when that session already has the history. Preserve the task requirements and apply later user direction. Store the question answer and deliver it to the waiting run. Finish governed tool responses before suspending. Execute the accepted plan without asking for the same approval again. **Steps to reproduce** Run the continuation, local-session-integrity, first-task, and agent-chat suites with native Codex and Claude. Include provider-question-bridge, accept-while-running, and plan-handoff. **Paperclip version or commit** This branch is based on master |
||
|
|
924f07be8c |
feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Chat connections let people start and continue that work from Slack. > - Setup mixed app creation, credentials, URL verification, account linking, and testing on the same screens. > - People also needed a safe way to link their own Slack identity after the first operator finished setup. > - This pull request gives each step a clear place and keeps membership approval separate from identity linking. > - It also makes connection details easier to use and fixes misleading callback health behind HTTPS proxies. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: chat routes and services, shared contracts, and the Apps board UI. **Problem or motivation** Slack onboarding made users find settings without enough guidance. A second user needed operator help to link their account. Activity stopped at 100 records, and TLS termination could mark working callbacks as stale. **Proposed solution** Use six setup steps with editable app names, a generated manifest, credential guidance, URL verification, account linking, and an optional message test. Send each Slack user a private, expiring confirmation link. Require company membership or an approved access request before linking. Add cursor pagination and tolerate the internal HTTP hop in callback diagnostics. **Roadmap alignment** This improves the existing connected-app surface and supports CEO Chat without changing the task-and-comments model. The maintainer requested and reviewed the flow during a live Slack test drive. **Additional context** Related work: #7, #3349, #13000, and #13620. Those cover broader chat capabilities, older webhook paths, or plugins. This PR improves the existing native connector's setup and account-linking flow. HTTPS documentation was published separately in paperclipai/paperclip-docs#128. ## What Changed - Split Slack onboarding into six clickable sidebar steps. Keep secondary and primary actions on one row. - Generate the Slack creation link and read-only manifest from editable app, bot, and command names. Add credential prefix validation and direct instructions. - Add live account-link status and an optional mention-based message test. - Add private, single-use Slack account invitations and membership access requests. Retain cloud authentication/bootstrap checks and enforce the chat rollout flag in all identity APIs. Default new Slack connections to linked users only. - Put Settings, Access, Conversations, and Activity in the sidebar. Simplify conversation rows and remove active header badges. - Add 25-item activity pages, stable timestamp/ID cursors, and replay safety across pages. Preserve the legacy array API for clients without pagination parameters. - Fix false callback warnings when HTTPS terminates at a proxy. Keep host, port, and path drift detection. - Document the setup flow, pagination, callback diagnostics, and shared wizard footer rule. ## Verification - Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`. - Passed: focused Slack callback and pagination integration tests; UI clipboard, wizard, pagination, and activity tests; OpenAPI route tests. The final access-gate fix also passes 27 focused tests covering cloud authentication/bootstrap, nonmember invitations, token validity, and the server-enforced rollout flag. - Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11 provider browser scenarios (including mobile light/dark navigation). After rebase, the identity route, sidebar, and 25 clipboard tests pass. - The full local `pnpm test:run` was attempted. The first run found 14 Slack fixtures that needed explicit guest access; those are fixed and the complete chat suite passes. Unrelated embedded PostgreSQL startup/resource failures and timeouts prevented a clean full local run. All CI checks pass on `2d858b036`, including the full chat, server, workspace, build, typecheck, and browser suites. - Live test drive: Slack app creation, credential setup, URL verification, private account confirmation, mention messages, and thread replies. Verified the callback warning clears for the existing proxied connection. - Review: create a Slack connection, follow the six steps, link a second user's account, and browse older activity with Next and Previous. ## Risks - Identity invitations carry a temporary capability. Tokens are hashed, expire after 15 minutes, work once, and require explicit confirmation by a company member. Access requests do not grant membership. - New Slack connections reject unlinked people by default. Existing connection settings remain intact. - Activity is a live ledger. Updated action rows can move forward in time. Older pages do not poll. - Proxy tolerance affects health display only. Slack signature checks and proxy authentication settings remain unchanged. - No database migration or package-lock changes. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser verification. The runtime does not expose an exact model build ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full local-run limitations documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
84fe89906d |
fix: complete native agent review handoffs (#13581)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Native execution uses durable runs, issue locks, wake requests, and typed tool authority > - A child can finish with a native agent review request while its original assignee stays responsible for the work > - The reviewer then needs a bounded execution path that can inspect the child, record one decision, and finish safely > - Before this change, assignee-only gates rejected the reviewer or left the parent waiting after the child review ended > - This pull request adds typed reviewer admission, scoped reviewer tools, durable wake and recovery handling, and parent continuation evidence > - The benefit is that native review handoffs complete without changing child ownership or granting broad mutation access ## Linked Issues or Issue Description Refs: #13314 Refs: #13574 **What happened?** A native child run could report `needs_review` for an agent reviewer. The reviewer wake then failed assignee and execution-lock checks. The child remained in review and the parent remained waiting. **Expected behavior** The named reviewer should receive one durable wake. The reviewer should inspect the child and resolve the exact review card. The child assignee should stay unchanged. The parent should receive the recorded review outcome after the child reaches its terminal state. **Steps to reproduce** 1. Run a native task with a different named agent reviewer. 2. Keep the child assigned to its original worker. 3. Let the worker finish with a native completion review request. 4. Start the durable reviewer wake. 5. Resolve the review and finish the reviewer run. 6. Observe the child and parent state. **Paperclip version or commit** Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification source: `eea171aae`. **Deployment mode** Built from source. **Installation method** Built from source (pnpm build). **Agent adapter(s) involved** Not adapter-specific (core bug). **Access context** Both. **Database mode** Embedded PostgreSQL in the isolated live test fixtures. ## What Changed - Add server-validated native review assignment facts. - Admit only the exact company, issue, source run, decision, revision, addressee, and resolver policy. - Give reviewer runs a narrow set of Paperclip read and resolve tools. File and shell access follow the configured agent and environment policy, so reviewers can run tests. - Separate server-owned reviewer instructions from untrusted persisted review data. Escape the data boundary; retain server-enforced authorization. - Keep the child assignee unchanged. Atomically claim the reviewer run, wake request, and issue execution lock. A competing lock prevents provider startup. - Require the exact running reviewer session and current issue lock to resolve its assigned card. Reject missing, unrelated, or terminal reviewer runs. - Add durable reviewer wake, lock, stale-card, and abandoned-run recovery handling. - Prevent duplicate native wake dispatches during deferred admission and recovery. - Carry accepted or rejected child review outcomes into parent task context and continuation evidence. - Add focused server, runner, and native protocol coverage. - Preserve upstream continuation rules. Add child review decisions as separate evidence, while keeping real human answers in their own field. - Return actionable completion validation feedback to both providers. Permit a corrected completion after rejection. Keep strict terminal acknowledgment validation. - Apply exclusive shared-workspace locks to sandbox environments. Local and SSH folders can run concurrently, including when old settings request serialization. - Repair test timing, native event parsing, and the review artifact assertion. Allow a valid reject, correct, and accept review sequence. Check the accepted card against its reviewer run and decision. Keep polling within the existing deadline when review acceptance precedes the parent wake projection; report a specific missing-continuation error at timeout. - Apply the ACPX pending-call limit to reserved finish/block calls, with capacity-release and cancellation tests. ## Verification - `pnpm build`: passed on `eea171aae`. - `pnpm -r typecheck`: passed on `eea171aae`. - `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on `b31ad9ab8`; runner E2E typecheck also passed. - `pnpm check:token-gates`: passed. - Focused DB review, reviewer authority, and prompt-boundary checks: 31 tests passed. They cover invalid reviewer runs, competing locks, atomic admission, duplicate claims, and valid resolution. - Heartbeat, workspace, and recovery checks: 30 tests passed. - ACPX sidecar suite: 27 tests passed. Moving the capacity guard back below reserved handling makes both new regression cases fail. - Four focused live continuation checks passed on their first attempt at `f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and question tool guidance (12/12 each). These cases do not use the reviewer prompt path changed afterward. - Fresh Codex and Claude review-handoff checks passed all 29 native checks each on their first attempt at `eea171aae`. Both runs received the expected fixed prompt and completed cleanup. Only the six selected live flows were tested; no full paid provider catalog run. - The final commit only extracts the existing test-harness timeout diagnostic into a shared helper and adds positive and negative coverage. Removing the accepted-review guard makes two regression assertions fail; restoring it passes all six timeout tests. Production runtime code, prompts, deadlines, and grading criteria are unchanged by this final commit. - Deadline regressions: a valid continuation delayed 20 seconds succeeds within its 30-second unit-test deadline; an absent wake returns a specific candidate-failure diagnostic at that same deadline. Both assertions failed before the fix. Production E2E deadlines remain unchanged. - Historical native failures remain recorded: Docker availability failures; a valid reject/correct/accept sequence that the first-card grader misread; and a test that rejected the gap between accepted child review and parent wake projection. No failed result was regraded. The latest tests use a protected reference to the pinned Docker image and the unchanged artifact oracle and time limits. - Full repository verification runs in GitHub CI. Local verification uses the focused suites above, full build, and full typecheck. An unchanged Codex shutdown timing test failed once in CI, passed in isolation, and its full shard passed on the final commit without changes to that test or its causal code path. The original failure is retained in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no outstanding actionable findings. All review threads are resolved. All current-head CI gates passed, including the isolated native runner Docker build (55 successful checks; two skipped by the workflow). ## Risks - Reviewer admission depends on exact persisted decision and interaction bindings. A stale or changed card is rejected. - Paperclip control-plane tools are limited to inspection and review resolution. This is not a filesystem permission boundary; provider file and shell access retain the configured policy. - Deferred wake recovery changes dispatch receipt coalescing. A scheduler regression could delay a continuation if the receipt state is wrong. - Parent review outcomes are evidence for the model. They do not grant tool authority or change issue ownership. - This change does not address legacy lease-hold handoff behavior. > Roadmap review: native execution, review gates, and durable recovery are existing roadmap capabilities. This PR completes a narrow reliability path for those capabilities. ## Model Used OpenAI `gpt-6-astra` with reasoning, tool use, and code execution. OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and journal work. Context window size is not exposed by this session. Live test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the PR authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e26d787928 |
Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must continue tasks using user answers without losing earlier requirements or approval gates. > - The wake prompt mixed human decisions with prior tool evidence and repeated detailed question instructions. > - Those instructions belong with the question tool, with a short routing hint in the wake. > - The Runner evals need to prove that answers, approvals, and completed work survive later turns. > - This PR shortens the prompts, separates authenticated answers, and adds continuation tests with useful screenshots. ## Linked Issues or Issue Description Refs #13517. This is a follow-up to the merged onboarding skill and Runner E2E work. Related #13539 covers responses received while a run is active; this PR preserves its cases and adds continuation coverage. Existing continuation/recovery and question PRs were searched; none covers this prompt/documentation and eval change. **What existing behavior does this improve?** The instructions sent when an agent continues a task, the native human-input tool documentation, and the evidence captured by Runner full-stack E2E. **Current behavior** The wake repeats a long question-tool guide. Human answers appear alongside untrusted prior results. Screenshot capture can finish at DOM load while the task still shows a spinner, even when backend behavior checks pass. **Proposed behavior** Keep earlier requirements unless the user changes them. Treat clarification as distinct from approval. Give authenticated human responses a scoped field. Keep tool and agent results as evidence. Put detailed question behavior in the tool descriptor and retain one routing sentence in the native wake. Wait for the correct task and loaded conversation before taking screenshots. **Reason and benefit** Reduce repeated prompt text and make authority boundaries clear. Test that real question cards, later answers, approval gates, and completed child tasks still work. Make screenshots useful for human review. ## What Changed - Shorten shared continuation instructions for legacy and native runners. Separate authenticated user responses from tool results and agent summaries. - Remove the detailed question guide from native wake prompts. Keep its behavior in the canonical `request_human_input` descriptor and existing payload schema. Regenerate semantic contracts and fixture hashes. - Add five continuation cases across four local profiles. Add a dedicated choice-then-text case for native Codex and native Claude. All 22 cells join the shared full E2E campaign. - Cover revised scope, clarification without approval, hostile instructions in a handoff file, and reuse of a completed child after restart. Keep production instructions and fixed user facts. - Capture continuation screenshots only when the intended task and conversation have rendered. Add provider-free browser regressions for loaders and wrong-task capture. - Preserve current master’s extra tool and onboarding cases. The default campaign now contains 166 cells; 35 manual everyday cells remain separate. ## Verification - `pnpm -r typecheck`: passed after replay on current master. - `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed. - `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm test:e2e:runner:browser-support`: 4 passed. These tests failed against immediate screenshot capture and passed after the fix. - Focused continuation and native-input tests: 36 passed locally. The tool-authority suite could not initialize embedded PostgreSQL locally, including one isolated retry; its 17 assertions did not run locally. The full remote server shards passed on this PR commit. - `pnpm build`: passed after replay on current master. `pnpm test:run` was attempted locally but hit the same embedded PostgreSQL initialization failure; the remaining local run was stopped after complete remote CI passed. This is not claimed as a full local test pass. - [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685): passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All server/chat/workspace/serialized shards, browser shards, Runner checks, typecheck, build, canary and policy checks passed. The isolated native Runner build and security checks also passed: 57 successful checks, with two expected Storybook skips. - Greptile reviewed the exact PR head at 5/5, with no findings or unresolved review threads. The PR has no merge conflicts. - [Live question-docs report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/): 3/3 passed at source `83dd132f2` before replay on master. Native Codex and Claude each asked a choice, waited, asked a text question, and saved both answers. Claude also passed a completed-child restart case. All three native turns are checked for absence of the old question block. - [Earlier continuation report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/): all five continuation cases passed on native Claude. The report retains campaign and revision provenance and separately shows two unresolved onboarding behavior failures. - [Before/after prompt report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/): full text, current recorded Claude inputs, and reproducible reference-token counts. The controlled wake comparison removes 401 reference tokens; the net counted input reduction is 339 after charging the larger tool description. These are text-size estimates, not measured billing savings. ## Risks - Prompt wording affects model behavior. Live results cover the stated cases, not every provider or conversation. Legacy profiles are registered but were not rerun for this change. - The optional continuation field changes prompt data only; there is no database migration or new production API. - Authenticated answer projection excludes generated summaries and agent-resolved interactions. It preserves the answer’s question or approval scope. - The screenshot guard can expose UI loading failures that earlier runs hid. Backend grading alone no longer makes those captures valid. - The two prior onboarding failures remain separate product issues: work before acceptance and a missing saved plan. This PR does not claim the entire onboarding suite passes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, code execution, and browser verification. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted tests above; the full local database-startup limit is documented - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d0b67bfe71 |
feat: queue approvals and answers during active runs (#13539)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users guide running agents through messages, questions, and approval cards. > - Messages already wait in a queue when an agent is running. > - Card responses did not appear in that queue. Some question answers also steered a later run without a user click. > - A fast approval could invalidate the agent's review handoff and cause it to stop its own run. > - This pull request gives card responses the same queue controls and preserves the exact response during delivery. > - Users can wait for completion or explicitly send the response with Interrupt or Steer. ## Linked Issues or Issue Description Refs #13517, which is merged. This PR targets master and adds queued interaction responses on top of the onboarding changes. Related continuation work: #10519 and #12866. **What happened?** Accepting a proposal while its source run was active left a saved response outside the message queue. The agent could then lose its review path, reassign the task, and cancel itself. Answers to older questions could also steer another active turn without a click. **Expected behavior** Save the response immediately. Queue its continuation behind the active run. Deliver it after completion, or when the user explicitly chooses Interrupt or Steer. Preserve approval revisions and answer choices. **Steps to reproduce** 1. Let an agent publish a confirmation card while its run is still active. 2. Accept the card before the agent finishes its review handoff. 3. Inspect the message queue and the task's next run. **Paperclip version or commit** Reproduced on da8a3876c with the onboarding changes from #13517. **Deployment mode** Local development from source. The fix covers legacy adapters and native Runner turns. ## What Changed - Project resolved cards into the existing queue as immutable responses. Keep answers and exact approval revisions. - Require an explicit click to steer a response into a compatible native turn. Use Interrupt when a fresh session is required. - Preserve typed response context through interruption, cleanup waits, and normal queue promotion. Keep the direct answer channel for a provider blocked on its original question request. - Accept the source run's review handoff after its card resolves. Reject stale agent reassignment that would orphan a queued response. - Add deterministic regression tests and an `accept-while-running` case to the first-task suite. Require recorded timestamp overlap before that case can pass. - Keep the first-task skill name out of user-facing messages. ## Verification - Red-green: the original route failed the queue regression; the changed route passes it. - Focused server/UI tests: 139 passed, including 64 queue-route tests. - Runner harness unit tests: 314 passed. - Server, UI, and Runner E2E typechecks passed. UI token gates passed. - Full repository typecheck and build passed. Server typecheck passed again after review fixes. - Review regressions: 165 queue/reopen route tests, 53 wake admission tests, and 18 run identity tests passed. Approval acknowledgement recovery and both message/approval arrival orders are covered. - Full local test run: 12,401 passed; three new admission regressions ran against a cached pre-fix module. A fresh run of that entire suite passed (53 tests). The complete CI suite passed on the final commit. - Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional Storybook checks skipped. Every server/workspace/browser shard, Runner verification, build, typecheck/release registry, canary, policy, and security check passed. Greptile: 5/5, no unresolved threads. Earlier interrupted CI workers were replaced by this fresh complete run. - After integrating the updated parent: 314 harness tests, 119 queue/admission tests, 44 onboarding/question-delivery tests, and 13 native recovery tests passed locally. Full repository typecheck and build passed. - Clarified the skill wording preference: routine replies describe the action without announcing the internal skill; direct questions and permission/security/execution disclosures remain truthful. - The paid `accept-while-running` scenario is registered for all four local first-task profiles. It has not been run against a model in this change. - Rebased onto the merged parent at `11921075a`; the resulting tree exactly matches the locally verified integration tree. Final-head CI on `b53054807` passed: 54 successful checks, 2 optional Storybook checks skipped, no failed checks. Every new server/browser shard, aggregate verify/e2e gate, Runner, typecheck, build, canary, and security check passed on the first attempt. Greptile reviewed this exact head at 5/5 with no unresolved threads. ## Risks - Responses now wait instead of implicitly steering another active turn. A provider blocked on the original question still receives its answer directly. - Approval receipts cannot be edited, discarded, or reordered as comments. This preserves the recorded decision. - Interruption must still prove that the prior execution stopped. The tests cover cleanup waits and duplicate delivery. - The new paid overlap case can be unexercised if the model finishes before the click lands. It cannot pass without evidence of overlap. - No database migration is required. This repairs the existing approvals and execution controls; it does not implement the roadmap's work-stream queues. ## Model Used OpenAI GPT-6 through Codex. The exact deployed model ID and context-window size were not exposed in this session. Capabilities used: agentic reasoning, repository inspection, code editing, terminal commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9fd2e50310 |
feat: create company skills from runner tasks (#13538)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Runner gives agents tools to change company resources. > - Users need agents to save reusable skills during a task. > - A saved skill needs a visible result that users can inspect and edit. > - This pull request adds `create_skill` and a task feed card linked to Skill Studio. > - Users can open the saved skill from the task and edit the same resource. ## Linked Issues or Issue Description **Subsystem affected** Runner tools, company skill storage, task feed, and Skill Studio. **Problem or motivation** The Runner has no dedicated tool to create a company skill. A user cannot follow a creation result from the task feed to the saved skill. **Proposed solution** Add a company-scoped `create_skill` tool. Save the skill with the existing company policy. Add one creation card to the task. Open a named sidebar tab from that card. Let the user open the same skill in Skill Studio. **Alternatives considered** An agent can write a local file, but that file is not a company skill. A second document copy in the task would become stale after a Studio edit. The sidebar therefore reads the saved skill directly. **Roadmap alignment** This extends the shipped Skills Manager, Skill Studio, and Skills Store milestone. The maintainer requested and approved this scope. Search found no duplicate `create_skill` PR or issue. Related UI validation work: #8715. This PR does not change that validation display. ## What Changed - Add the real Runner tool, its contract, and its mock implementation. - Validate the complete SKILL.md and derive company, task, agent, and run identity from authentication. - Apply the existing company skill policy. Do not assign the skill to an agent. - Make keyed retries return one skill and one creation event. Reject conflicting retries. - Make concurrent file creation safe. Never replace an existing published skill during creation. - Add a creation card, a named sidebar tab, and an Open in Skill Studio action. - Show saved Studio edits when the user returns to the task. - Add storage, policy, mode, retry, UI, and Product E2E tests. Document the tool. - Fix deleted-name reuse, onboarding panel persistence, immediate feed refresh, and mock validation parity from review. - Serialize Studio file edits and renames with skill deletion and recreation. Reject stale editor requests before they can change a replacement skill. - Generate the standalone mock parser and validator from the production contract. Use portable UUIDs so the browser scenario bundle builds. ## Verification - All latest-head PR checks pass on `145dd76a5`, including all server shards, browser E2E, Runner verification, build, typecheck, and release dry run. Greptile: 5/5 with no open findings. An interrupted CI runner was retried successfully. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Review regressions: 73 storage tests, 6 real API tests, 63 UI tests, and 61 semantic runtime tests passed. Parser synchronization passed. - CI exposed existing fire-and-forget Sentry test races. Reproduced the resumption race locally, then synchronized the related sweep and finalizer assertions on the actual report; all 27 tests across the three affected files pass. - Runner scenario browser build and strict content-security-policy check: passed. - Runner suite: 2,012 tests passed; 10 skipped. - `pnpm test:run`: the general-server batch had 12,416 passes and two failures. The old tool-count assertion was fixed; all 16 authority tests then passed. The chat webhook test had a socket error; it passed four isolated reruns. - Both workspace test groups passed. The isolated route suites completed. Two socket failures in the initial route batches passed on individual reruns; all remaining 61 files passed. - Product E2E `create-skill-studio`: passed with local Codex and local ACPX Claude. - Manual browser test: submit a task, observe the real tool call and creation card, open the sidebar, edit in Studio, save, and return. The task reached Done. The saved second revision and sidebar tab survived a server restart. - The new companion headless Runner Eval passed. Companion coverage PR: https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not run because no immutable runner image was configured. ## Risks - Database writes and local file writes cannot share one transaction. Recovery accepts only an exact file-for-file retry after a database rollback. Conflicting files remain untouched. - The sidebar displays the current skill. The feed card remains the historical creation receipt. - No database migration, dependency, or workflow change is included. - Remote Daytona behavior still needs a run with a configured immutable image. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and browser verification. OpenAI `gpt-5.6-luna` assisted with bounded implementation and eval work. Both used code execution and tool access. The host did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18989a9e73 |
docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its evaluations test both the Runner and complete product workflows. > - The guides and run histories are in separate places. > - The shared Evalbook viewer can make the test boundary unclear. > - This pull request names the two families and adds a guide, authoring skills, and a public hub. > - Contributors can choose the correct test and inspect its history. ## Linked Issues or Issue Description **Issue type** Missing documentation. **Where is the issue?** Runner and Product E2E evaluation guides, case-authoring procedures, and public result navigation. **What's wrong?** There is no single entry point. A report format can be mistaken for an execution boundary. There are no dedicated case-authoring skills for these two families. **Suggested fix** Add a guide and three skills. Link both existing histories from a public hub. Keep existing campaign URLs and grading unchanged. ## What Changed - Add `doc/evals.md` and links from existing guides. - Add the `paperclip-evals`, `add-runner-eval`, and `add-product-e2e-eval` skills. Install copies in `~/paperclipai/.agents/skills`. - Add a static hub builder that reads the existing public history feeds. - Show a dated snapshot for each family. Label partial campaigns and preserve measurement dates across report refreshes. - Document publication and refresh commands for https://pages.paperclip.ing/evals/. ## Verification - Seven Python summary tests pass: `python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR does not modify package scripts. - All three skills pass the skill-creator `quick_validate.py` check with `/usr/bin/python3`. - Build tested with saved history fixtures and the live public feeds. - Desktop and mobile browser checks pass. The mobile page has no horizontal overflow. - Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200, no page errors, all eight links return HTTP 200, no mobile overflow. - Independent skill exercises found the existing Notion-decline case and a direct Runner permission-denial case. Roster validation with an explicit run ID passes. - Missing refresh measurement date: regression fails before the fix and passes after it. - `git diff --check` passes. - No paid evals were run for this documentation and reporting change. The preceding head passed typecheck, build, server/workspace tests, runner verification, browser E2E, and the canary dry run. Checks for the latest commit are pending. Local repository-wide typecheck, test, and build were not repeated because no product code changed. ## Risks The hub is a dated static snapshot. It can lag behind the linked histories until an operator refreshes it. A changed history schema stops the build. Existing archives and grades are not modified. The published guide link is pinned to the reviewed commit so branch deletion cannot break it. Later builds can use master. ## Model Used OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna for documentation and independent skill checks. Both used repository tools and code execution. Context window sizes are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (latest commit pending) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (preceding head was 5/5; latest commit pending) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4577d10029 |
fix: prepare everyday artifact and Codex sandbox prerequisites in CI (#13516)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The runner E2E workflow executes paid everyday workflow stories on disposable CI hosts > - The everyday artifact oracle requires a pinned Python image and fails closed when it is absent > - Fresh CI hosts did not prepare this image before the paid cell, so project stories failed during preflight > - This pull request prepares and verifies the pinned image before the affected everyday cells > - The benefit is reliable artifact isolation checks on fresh trusted CI hosts ## Linked Issues or Issue Description **What happened?** Fresh trusted CI runners did not have the pinned Python artifact oracle image. **Expected behavior:** The workflow prepares and verifies the pinned image before an everyday project story starts. **Steps to reproduce:** Run an everyday project story on a fresh CI host without the image cached. The `everyday-artifact.py --preflight` check fails before task creation. **Paperclip version or commit:** `master` at `bd51f157e`. **Deployment mode:** Other: GitHub Actions trusted paid workflow. ## What Changed - Add a matrix-gated CI step for everyday project and recovery cells. - Check Docker, pull the fixed digest with bounded timeouts, and verify the exact repo digest. - Apply the existing Codex sandbox preparation to both native Codex profiles, including the mini profile. - Add workflow security assertions for ordering, condition, digest, timeouts, and secret isolation. - Document that CI prepares the pinned oracle image. - Check provisioning eligibility against every catalog cell, and scope the Daytona registry inspection assertion to the Daytona image job. ## Verification - `pnpm test:e2e:runner:unit` — 24 files and 222 tests passed. - `pnpm test:e2e:runner:typecheck` — passed. - `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts workflow-security.test.ts` — 10 tests passed. - Python artifact oracle calibration — 12/12 passed. - `git diff --check` — passed. ## Risks Low risk. The image step runs only for everyday cells that execute the artifact preflight. The Codex setup now covers both native Codex profiles. It uses a fixed public image digest and has no provider credentials. ## Model Used OpenAI gpt-5.6-luna (implementation subagent) and gpt-6-astra (review fixes and orchestration), using code execution and repository tools. Context window size is not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Paperclip <paperclip@paperclip.ing> |
||
|
|
669bd0e7a9 |
fix(ui): show progress and verify subscriptions during onboarding (#13499)
Automatically verify detected Claude and Codex subscriptions during initial onboarding. Show connection progress, support retries, and ignore stale results after navigation. Prefer the personal default subscription while preserving the later create-agent chooser. Allow Enter to advance from the agent name field. Add regression tests and production-component Storybook coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9adeabb590 |
fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections let people give agents access to external tools. > - A personal connection needs the current user's authorization. > - A new MCP URL must be probed before Paperclip can discover its sign-in method. > - Requiring a personal grant before that probe prevents sign-in from starting. > - This pull request permits the initial probe for a creator-owned draft with no credentials. > - People can complete personal setup while later requests retain authorization checks. ## Linked Issues or Issue Description **What happened?** Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and "This connection needs the current user's authorization". Paperclip checked for a personal grant before contacting the provider. The health wrapper also changed the expected authorization error into a server error. **Expected behavior** Discover OAuth and start browser sign-in. Create a personal grant after consent. For a public endpoint, discover its tools and create the empty personal grant after a successful probe. Keep missing authorization on later health checks as HTTP 422. **Steps to reproduce** 1. Add an unknown remote MCP URL with no saved credentials. 2. Select "Just me". 3. Check the link. Before this fix, the request fails before sign-in or tool discovery. **Paperclip version or commit** The three original regressions fail against `6cfe4acff` with the service fix removed and pass with it restored. **Deployment mode** The defect was reported in production and reproduced in local server tests with isolated PostgreSQL. Related work: Refs #11831. Refs #11144. Searches found no duplicate fix. ## What Changed - Allow an initial credential-free probe only for the creating user's personal draft with unknown authentication and no supplied credentials. - Leave OAuth grant creation to the callback. Create an empty personal grant only after a public probe succeeds. - Preserve the personal grant for URLs that already contain a credential. - Make empty personal grant creation conflict-safe without overwriting a concurrent grant or duplicating its creation audit. - Commit the empty grant and audit atomically. Retain the established public draft identity after a later catalog failure, so a failed retry cannot remove a successful retry's grant. - Run the following catalog/default-profile step in its own transaction, so failures discard partial catalog, profile, binding, and audit changes without deleting the established identity. - Verify archived personal connections retain their owner: another user is rejected before probing, while the original owner can resume setup. - Preserve `user_authorization_required` and HTTP 422 in health failures. - Cover the connect and OAuth callback routes, real loopback HTTP, credential-bearing URLs, and later health checks. - Document personal setup and the test fixtures. ## Verification - Red/green: the original three tests failed with the exact reported error before the fix and passed after it. - The credential-bearing personal URL regression also failed before its guard was added. - Focused suite: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/generic-mcp-connection.test.ts src/__tests__/tool-access-service.test.ts` passed all 385 tests across the two suites. An earlier run had a socket hang-up in an existing agent-permissions test; the unchanged suite passed on rerun. - The concurrent rollback regression failed before its fix because the successful retry's grant was deleted. It now verifies the grant and draft survive and a later normal health check succeeds. - Database fault injection during profile-entry insertion reproduced partial catalog writes before the transaction fix. The regression now verifies unchanged catalog rows, no partial profile/bindings, a retained grant, and successful retry. - CI's first serialized-server shard 3 attempt failed an existing peer-agent mutation test (the real run-context guard ran despite the test's mock). The test passed in isolation and all 108 tests in that suite passed unchanged locally. The single failed-shard rerun passed without code changes. - Final-commit CI: all 32 applicable checks passed on `639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected tests, the full test matrix, browser suite, build, typecheck, release checks, and security checks. The two Storybook-only checks were not applicable and skipped. Greptile is 5/5 with all review threads resolved. [Successful CI run](https://github.com/paperclipai/paperclip/actions/runs/35027478353). - `pnpm -r typecheck` passed. - `pnpm smoke:mcp-fixtures -- --require-paperclip` passed. - `pnpm build` passed. - Full local `pnpm test:run` was attempted: its general-server group finished with 12,372 passed, 4 failed, and 70 skipped tests. The run started before review edits; its two MCP failures used the old cached service (including an insert without the new conflict clause). All 385 focused tests pass on the final code. The other failures were existing workspace-cleanup and runtime-port tests; their unchanged suites passed on rerun (66 passed, and 25 passed/3 skipped). The local command stopped before later groups. The final-commit CI matrix is the full-suite merge gate; this local run is not claimed as green. ## Risks - The initial probe must not become a general authorization bypass. It is restricted to the creating user's draft. Normal health checks retain authorization enforcement. - Public endpoints get a personal grant with no secrets only after they answer successfully. Credential-bearing URLs keep their existing grant. - The concurrent-probe regression seeds the catalog and default profile to isolate grant creation. Existing first-time catalog/profile creation races are outside this change; this does not claim to make the entire setup flow concurrency-safe. - No database migration or UI change is required. OAuth tests use a simulated provider; the public endpoint test uses real loopback HTTP. ## Model Used OpenAI Codex, GPT-6-based assistant for regression tests and PR preparation; a GPT-5-based Codex assistant assisted with the initial implementation. Exact runtime model IDs and context-window sizes are not exposed in this session. Both used reasoning, repository tools, and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9cfa7fd2d1 |
fix(ui): share compact live and saved activity across runners (#13421)
## Thinking Path > - Paperclip helps people manage AI agents and inspect their work. > - Task threads show live activity and saved transcripts from several runners. > - Legacy runs used a separate activity renderer that expanded tool cards as commands arrived. > - A compact live row makes ongoing work easier to follow. > - This PR shares the native runner activity group across live and saved legacy turns. > - Both paths now use the same labels, icon gutter, animation, and history controls. ## Linked Issues or Issue Description Related: Refs #13255, which introduced rolling native-runner activity groups. **What happened?** During a legacy CLI run, each new command added an expanded tool card under an activity heading. The original legacy parity story rendered a completed turn, so it did not exercise this live path. **Expected behavior** Show one current activity line per commentary group. Roll that line forward when a new activity starts. Keep tool icons aligned on the left, use friendly labels, and show history only when expanded. **Steps to reproduce** 1. Start a task with the Codex local adapter using the CLI engine. 2. Ask it to read two files in separate tool calls and run a test command. 3. Watch the activity feed while the run is active, then expand its history. **Paperclip version or commit** The live behavior was reproduced on bdee5ebb21b2d09280e49c88bc329435da1b07c8. This update includes both live and saved rendering fixes. **Deployment mode** Built from source with the local test-drive command and a real Codex CLI agent. ## What Changed - Use `TaskChatRunnerActivityGroup` for both saved legacy phases and `TaskChatLiveTail`. - Align the legacy Working spinner with the shared activity icon gutter. - Cover streamed reasoning updates, successive commands, image labels, hidden details, and explicit history expansion in tests. - Feed raw legacy transcript events through the real adapter and live renderer in Storybook. Add live, expanded, narrow, light, and completed stories with the final reply preserved. ## Verification - Passed 195 targeted tests covering the live tail, status pill, shared activity group, native turn, and full task thread. - Passed UI typechecking, token gates, UI build, and Storybook build before submission. - Observed a real Codex CLI run while it read separate files and ran tests. The current activity stayed at one 32-pixel row without an accumulated tool list; the spinner and activity icon centers aligned. - Compared native and legacy Storybooks: identical rolling animation, fixed height, persistent expansion, truncated long labels, and correct light and completed states. - Full repository `pnpm -r typecheck` and `pnpm build` passed after replaying the branch on current master, including the Rust runner. The full `pnpm test:run` suite is still running. ## Risks - Legacy activity now starts collapsed. Users can expand each group and each row to inspect the same transcript details. - Live runtime request cards must retain their timeline positions. Existing task-thread and native-runner regression tests cover this boundary. - No API, database, or transcript format changes. ## Model Used - OpenAI GPT-6 through Codex for the live-path fix, tests, and browser verification. Capabilities used: reasoning, tool use, local code execution, and image inspection. The exact deployment ID and context-window size are not exposed in this session. - The original saved-turn change recorded OpenAI Codex with GPT-5.6; its exact variant and context size were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d49f168381 |
fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents must publish generated files so users can inspect their results after a sandbox stops. > - Legacy sandbox bridges blocked attachment listing and could not carry multipart binary uploads through the queue transport. > - The native runner has a separate verified file registration path that needs the same durable result. > - This pull request repairs legacy binary transport, makes native download receipts explicit, and reveals new outputs in the task Artifacts tab. > - Users can open generated files from either runner without a transport flag change. ## Linked Issues or Issue Description Refs #13355 for the existing native file publication path. Related filename fixes: #2615 and #4788. Related sandbox persistence work: #13376. This change repairs attachment delivery through the existing API; it does not add workspace persistence. **What happened?** The upload helper first lists task attachments to avoid duplicates. Both legacy bridge allowlists rejected that GET request with 403. A direct multipart upload also failed: the queue bridge accepted only JSON, excluded attachment uploads, and converted bytes to UTF-8 text. Enabling HTTP/2 alone did not fix the missing listing route. These failures occurred before attachment storage. **Expected behavior** Both runners can publish a workspace file, register its work product, bind it to a response, and return a working download. The file stays accessible after sandbox deletion. A new output opens the task Artifacts tab. The agent receives accurate errors and decides how to retry or report a failure. **Steps to reproduce** 1. Run a legacy agent in Daytona with the duplex bridge disabled. 2. Invoke the bundled upload helper with Bash on a PNG or PDF. 3. Repeat with the duplex bridge enabled. 4. Register the same file through the native runner with generic API tools disabled. 5. Retry registration, delete the sandbox, and compare the downloaded bytes with the original file. **Paperclip version or commit** The failing baseline was `f2c5e54dc`. This branch is rebased onto `6cfe4acff`. **Deployment mode** Source checkout with a local API and real isolated Daytona sandboxes. ## What Changed - Allow authenticated attachment listing, upload, and content download through both legacy bridge transports. - Add optional base64 body encoding to queue envelopes. Preserve the existing UTF-8 contract when the encoding field is absent. Decode binary bodies before forwarding them. - Preserve multipart headers. Bound raw bytes, encoded envelopes, and in-flight reservations. Retain timeout and uncertain-write behavior. - Preserve helper deduplication and return structured uncertain-write failures. Document explicit Bash invocation in live skills. - Add attachment IDs and content/download paths to native registration receipts. Reuse verified local and remote file reads, attachment storage, work-product registration, and response binding. - Preserve Unicode upload filenames and provide a valid Content-Disposition header. - Open the task Artifacts tab when new stored outputs arrive, including a closed desktop panel or mobile drawer. Deduplicate upload and registration events by object ID. Preserve manual selection on refetches, edits, and panel remounts. - Remove task artifact filters, the company Artifacts footer link, and the unassigned group heading and timestamp. ### Screenshot  This is the local display fixture. The image was generated separately and published through the attachment and work-product APIs. ## Verification - Post-rebase `pnpm -r typecheck` and `pnpm build` pass. - The post-rebase local `pnpm test:run` passed 12,369 tests before one existing conversation reset test timed out; all 33 tests in that suite pass when rerun with isolated test configuration. The aggregate command stopped before its remaining groups. GitHub runs the complete suite in separate shards. - All [GitHub verification checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350) pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks and two configured skips. The native-session recovery assertion initially raced its fire-and-forget Sentry report; all 13 tests pass locally, and the same-commit CI rerun passes all 170 suites (3,079 tests). - Live post-rebase Daytona: all three file-delivery tests pass. They cover the real Bash helper with the queue bridge, the helper with HTTP/2, and native `register_deliverable` with generic API tools disabled. - Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate registration, response binding, authorization controls, and byte-for-byte downloads after sandbox deletion. - Local focused coverage includes transfer bounds, malformed encoding, interrupted transfers, remote path containment, and native file verification. The attachment route suite passes all 32 tests, including an eight-case filename-header matrix for Unicode and special characters, inline and forced downloads, and full and partial responses. - Browser verification confirms image previews, persisted downloads, automatic Artifacts selection, and preserved manual selection after edits and reloads. Desktop/mobile component coverage passes. The latest UI cleanup passes its 10 affected tests and token gates. - Coverage limit: the Daytona tests call the real helper and native registration path directly. They do not replay a complete model-led image-generation task through the browser. Live command (requires a configured Daytona credential): ```sh PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts ``` ## Risks - Binary queue bodies use more memory because base64 adds encoding overhead. Transfer and process limits must remain aligned. - An interrupted write can have an unknown result. The bridge reports this state and preserves stable retry identities. - New artifacts intentionally change the active task tab. Existing history and repeated updates must not take focus again. - Transport flag defaults, server authorization, frozen skill snapshots, and completion policies remain unchanged. No schema migration is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository tools, code execution, and browser testing. The runtime does not expose a more specific model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6cfe4acff7 |
fix: preserve README images in npm package (#13488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its CLI is published as the paperclipai npm package, with the root README shown on the package page > - The root README uses repository-relative image paths so images render correctly on GitHub > - npm resolves those paths under the package repository directory, which is cli, so the image requests point to missing cli/doc/assets files > - This pull request prepares the generated npm README by converting only image src and srcset asset paths to stable raw GitHub URLs > - The benefit is that the same source README remains correct on GitHub and the published npm README displays its images ## Linked Issues or Issue Description **Where is the issue?** The issue is in the root README image assets and the npm packaging step in scripts/build-npm.sh. The affected public page is https://www.npmjs.com/package/paperclipai. **What's wrong?** The npm build copies README.md into cli/ before publishing. npm resolves relative image paths beneath the package repository directory, so doc/assets/banner.jpg becomes cli/doc/assets/banner.jpg. Those files do not exist, and the images render as broken on npm. **Suggested fix** Keep the root README paths relative for GitHub. Rewrite repository-relative image paths only in the generated npm README copy to absolute raw.githubusercontent.com URLs. ## What Changed - Added a small npm README preparation script that rewrites relative image src and srcset asset paths. - Updated scripts/build-npm.sh to use the preparation step when generating the npm package README. - Added a regression test for src, srcset, immutable refs, absolute URLs, and non-image Markdown links. - Pinned release-build image URLs to the source commit, while preserving tarball builds by passing their known source refs. ## Verification - Passed: node --test scripts/prepare-npm-readme.test.mjs - Passed: bash -n scripts/build-npm.sh scripts/e2e-install-lifecycle.sh scripts/e2e-update-migrations.sh - Passed: focused README and E2E migration harness tests - Passed: git diff origin/master...HEAD --check - Generated README asset URLs were checked against raw GitHub and all seven returned HTTP 200. ## Risks Low risk. The change affects only the temporary README generated for npm packaging. It does not change the GitHub README or runtime code. The generated npm README depends on the public raw GitHub asset URLs remaining available. ## Model Used OpenAI Codex, GPT-5. Tool-enabled repository inspection, code execution, browser verification, and git/GitHub operations were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes: # / Refs: # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f4cdc7b231 |
fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The control plane prepares task workspaces before it starts an agent. > - Workspace preparation reads Git state so it can preserve edits and exclude private files. > - A failed scan was treated as a non-Git folder and lost its actual failure code. > - The resulting generic setup failure could not recover, even when the cause was temporary. > - This pull request keeps the cause and uses the existing bounded retry schedule before provider startup. > - Tasks can recover without human intervention, while permanent failures and exhausted retries stop with useful guidance. ## Linked Issues or Issue Description **What happened?** A Git scan error during managed repository preparation became `Configured repository folder is not a Git checkout`, followed by generic `setup_failed`. The agent never started. Generic recovery could not distinguish a temporary timeout from a bad workspace configuration. **Expected behavior** Keep the closed scan error code. Retry temporary timeouts and queue saturation under the existing shared budget. Preserve edits, exclusions, ownership, and pause gates. Stop permanent failures and exhausted retries with a specific explanation. Do not replay historical generic setup failures. **Steps to reproduce** 1. Configure a task project with a local Git source that must be copied into its managed repositories. 2. Make the ignored-file scan exceed its timeout before the agent starts. 3. Before this fix, the snapshot returns null and the run ends as non-retryable `setup_failed`. 4. Use the disposable browser fixture in `tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and test the full recovery path. **Paperclip version or commit** Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`. **Deployment mode** Built from source. The defect is in core workspace setup, not a specific model provider. Related work: Refs #13442 (managed repository preparation), Refs #11572 (bounded Git scheduler), Refs #12997 (separate adapter startup retry work), Refs #13469 (separate terminal-workspace scan performance work). ## What Changed - Return the non-Git fallback only for repository discovery. Propagate failed scans of a confirmed repository. - Replace full ignored status output with an ignored-only directory listing. Preserve NUL-delimited paths and exclusions. - Preserve typed, sanitized scan errors through workspace preparation and persist pre-provider failure details. - Retry only timeouts and queue saturation, using the existing durable two-retry budget and issue gates. Prevent generic recovery from adding another budget. - Show workspace-specific failure copy and actionable exhausted-recovery notices. - Add red-green unit tests, real-database restart and retry-boundary tests, and opt-in browser acceptance fixtures with real Git subprocess timeouts. - Document the recovery contract and browser verification procedure. ## Verification - Red: injected scan failures returned null instead of rejecting; setup lost the timeout code; task-thread and recovery notices had generic copy. - Green: 119 focused adapter/backend tests, 20 recovery-boundary tests, and 136 task-thread tests. - `pnpm -r typecheck` — passed. - `pnpm build` — passed on the final production code. - `pnpm check:token-gates` — passed. - The initial local `pnpm test:run` overlapped source edits and was interrupted after two late-added assertions saw pre-fix behavior; it is not counted as a green full run. A fresh final-head run passed all 229 tests across the six affected adapter/backend/UI suites. The clean latest-head CI full test matrix passed: all five general-server shards, all five serialized-server shards, and all three general-workspace shards. - Latest-head CI also passed all three browser shards and their aggregate gate, typecheck and release registry, build, runner verification, canary dry run, policy, Docker context integrity, and security gates. Greptile: 5/5, with the review thread resolved. - Browser: created a task in a disposable instance. A real Git timeout scheduled recovery, the next run completed through the run-scoped API without manual Retry, and Done survived reload. The deterministic process worker checked preserved source edits and excluded private files; no model calls were made. - `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6 minutes). The persistent case made exactly three failed attempts, never started the worker, showed the cause-specific notice, stayed stopped for another scheduler tick, and retained Blocked after reload. - Extra red-green coverage: 50 recovery tests passed after fixing an exhausted-bootstrap classification that incorrectly implied unknown provider actions. Missing or uncertain evidence still retains the safety hold. - Verified the documented Git executable override during repository seeding. ## Risks - A confirmed repository scan failure now fails closed instead of falling back to directory sync. This prevents unfiltered copying but makes previously hidden errors visible. - Temporary host problems can create up to two additional setup attempts, 30 seconds apart. Permanent scan errors do not auto-retry. Generic recovery cannot reset this budget. - The durable retry path still enforces ownership, pause, and work eligibility. Integration tests cover restart, duplicate promotion, pause, exhaustion, and non-retryable categories. - No schema migration, new runtime setting, new retry budget, production deployment, or historical task replay. ## Model Used OpenAI Codex, GPT-5-based coding agent, with reasoning, repository tools, shell execution, and browser testing. The exact deployment model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
cceeb0aa66 |
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4510bf7c9e |
ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI uses a trusted workflow on the AWS runner fleet. > - The caller used a fixed SHA that also needed runner-group admission. > - A mainline pin update left CI queued because the group still allowed older SHAs. > - This pull request calls the trusted workflow on master, which requires code-owner review. > - New merged workflow versions can use the existing master runner-group entry. ## Linked Issues or Issue Description Related: #12968. That Dependabot PR proposes another SHA rotation. This change keeps this first-party workflow on master instead. **What happened?** CI run 34975562974 stayed queued because its trusted workflow SHA was absent from the runner-group allowlist. The fleet itself was healthy. **Expected behavior** New reviewed versions of the trusted workflow on master should receive runner access without a separate SHA allowlist update. **Steps to reproduce** Change the caller to a new trusted workflow SHA without adding that SHA to the restricted runner group. Its jobs remain queued. The master reference removes that recurring synchronization step. ## What Changed - Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`. - Exclude this exact first-party workflow from Dependabot updates. - Update the existing E2E shard workflow tests for the master caller contract. - Document the runner-group entry, required code-owner review, and old-reference retention. ## Verification - `actionlint .github/workflows/pr.yml` passed. - `node --test scripts/__tests__/e2e-shard.test.mjs .github/scripts/tests/cloud-runner-routing.test.mjs .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests. - Parsed Dependabot YAML and checked the exact workflow exclusion. - `git diff --check` passed. - Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS has no errors, and the active master ruleset requires code-owner review. This covers the trusted workflow, caller, and CODEOWNERS itself. - The approved organization setting now allows `paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`. All previous references and other runner-group settings remain intact. - Full application typecheck, tests, and build were not repeated locally for this workflow-only change. This PR's CI and review are pending. ## Risks New versions of the trusted workflow take effect for new callers after merge to master. Keep code-owner review and master protection enabled. Existing administrator pull-request bypasses remain unchanged. Third-party action pins, runner routing, and infrastructure are unchanged. Older callers still use their SHA pins; their allowed refs remain in place. ## Model Used OpenAI GPT-6 through Codex performed implementation and orchestration; its exact runtime variant and context-window size were not exposed. OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions and applied the approved runner-group setting. Both used code and tool access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a2e7ffdc34 |
fix(runtime): validate sandbox paths and preserve live controller leases (#13432)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents can run in remote sandboxes. > - Connection checks must use the selected execution target. > - Recovery must respect the controller that owns an active run. > - A host path or PID does not describe a remote sandbox. > - This pull request checks sandbox paths on the target and preserves live controller leases. ## Linked Issues or Issue Description **What happened?** Selecting an AI account for a sandbox agent could fail because the Claude ACP environment check tried to create the sandbox directory on the Paperclip host. The recovery sweep could also interrupt a sandbox run while its controller lease was still valid. It treated a PID absent from the local host as proof that the run had stopped. **Expected behavior** ACP checks directories on the selected execution target. Recovery leaves a run with a live controller lease alone. Its final database write rejects a stale snapshot after renewal, a claim, a controller change, or a runtime change. **Steps to reproduce** 1. Test a Claude ACP sandbox agent with a directory that cannot be created on the host. The check fails before this fix. 2. Give a running sandbox task a valid controller lease and a PID absent from the host. Run the stale-lock sweep without an in-memory handle. The sweep interrupts the run before this fix. 3. Renew or replace the controller between the sweep's read and write. The old snapshot must not end that controller's run. **Paperclip version or commit** Rebased onto `origin/master` at `0e9b24c8216171c26c8358ba387d77858e02c7a9`. All seven regression cases still fail against this base. Refs #13438, which supplies the managed hiring and task-connection behavior, and #13433, which preserves non-assignee subscription comment wakes. This PR preserves both upstream changes and addresses the two remaining sandbox failures. ## What Changed - Resolve and create Claude ACP test directories through the execution-target helpers. - Preserve active legacy controller leases during stale-lock recovery, including finalization after a task becomes terminal. - Recheck the controller, lease, runtime mode, and native ownership in the terminal database write. - Add two sandbox-directory cases and five database-backed controller-lease cases. - Document the target used for ACP directory checks. ## Verification - Red: all seven new cases fail against `0e9b24c82` without these two implementation changes. - Before the final upstream sync, 192 focused tests passed. Full `pnpm test:run` coverage completed using the repository's group/shard runner: all general server and workspace groups passed, and all 147 serialized server suites passed across the initial run and isolated continuations. Five cold-import timeout suites passed with `--experimental.fsModuleCache`; their assertions and deadlines were unchanged. - The hiring routes, default-selection service, and upstream hiring tests match `origin/master` exactly. The two remaining fixes are unchanged by the final rebase. - On final head `bd28d5cefbdf7084acc3759199cd9661907e7a26`, all 372 focused tests pass across 16 suites covering both upstream changes and these fixes. Two timeouts in the combined run (database setup and an existing ACP case) pass in isolated reruns with fresh test homes and temporary directories. `pnpm -r typecheck` and `pnpm build` also pass on this head. - All 32 active checks pass on final head `bd28d5cef`, including the full test matrix and browser shards, in [CI run 34904204849](https://github.com/paperclipai/paperclip/actions/runs/34904204849). Two Storybook checks are skipped by path filters. - Greptile reviewed final head `bd28d5cef` at 5/5 with no findings or unresolved review threads. ## Risks - A failed remote directory check still blocks connection adoption. - A live controller retains finalization authority after its task becomes terminal. Cleanup waits for ownership to expire and must pass the final ownership check. - No schema or credential-storage changes. ## Model Used OpenAI GPT-6 in Codex, with reasoning, repository inspection, code execution, and API tools. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8f1905d34d |
fix: provision all project repositories for local and sandbox tasks (#13442)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Projects can now attach several source repositories. > - Task preparation still treated these sources as alternative workspaces. > - Sandbox sync preserved Git history only for the selected repository. > - A task needs every attached repository to complete work across the project. > - This pull request prepares all distinct project repositories and preserves their separate Git histories through sandbox restore. ## Linked Issues or Issue Description **What happened?** A user reported that a project with two repositories received only the first repository in Daytona. Repository-only project rows also reached the agent with null local paths. Managed checkouts with matching repository names could resolve to the same directory. **Expected behavior** Local and sandbox tasks receive every distinct repository attached to their project. Repository-only sources work without preconfigured local folders. Each repository keeps its own Git history and working files. **Steps to reproduce** 1. Create a project with two repository sources and no local folder paths. 2. Assign a task to the project and run it in Daytona. 3. Inspect the task workspace and the repository paths exposed to the agent. 4. Observe that the original implementation supplies only the selected checkout. Related change: #13010 added multiple repository selection. The open repository-catalog proposals #11234 and #11228 cover a different data model. This fix uses the existing project workspaces. ## What Changed - Materialize each additional distinct repository as an editable checkout inside the task root. Seed configured local sources with their current working files and retain task edits across runs. - Pass materialized repository paths to local agents and native sandbox task prompts. Apply existing run-scoped Git credentials to each remote clone. - Preserve each repository's Git history, dirty files, and restore baseline during sandbox staging and durable recovery. Apply each repository's ignore rules and the operator's workspace exclusions. - Keep same-name managed repositories in separate directories. Report additional clone failures before the task starts. - Add task-level, checkout, sandbox round-trip, environment-hint, and recovery-descriptor regression coverage. Document checkout and restore behavior. ## Verification - Red: the original implementation fails the sandbox test because the second repository has no Git directory. It also fails the same-name checkout test and both real-database task tests because repository hints have no local path. - Green: focused tests pass for one and two repository-only sources, local source edits, clone failures, per-repository credentials, separate Git histories, ignored files, and recovery from remote or durable seed state. - Live Daytona smoke passed with two disposable repositories through the production provider sync functions. Both repositories arrived with Git history. Commits from both restored locally. Ignored files stayed excluded. The disposable sandbox was deleted. - Passed on final commit `93ab76763`: `pnpm -r typecheck` and `pnpm build`. - Final focused coverage: 254 assertions across the six changed test areas passed across the serial run and an isolated rerun of the existing process-kill timing test. The live Daytona smoke also passed. - The local `pnpm test:run` overlapped source edits and retained stale transformed code. Its first phase reported 12,240 passed assertions, nine failed assertions, three hook failures, and one worker error; later phases did not run locally. This run is not claimed as green. Fresh focused tests verify the changes, and every general/workspace and serialized-server CI shard passes on the final commit. - Final CI is green on `93ab76763`: all test shards, all three browser shards, typecheck, build, runner verification, canary dry run, and security checks. The initial unrelated chat-delivery browser timing failure passed in the final CI run. Optional Storybook visual checks were skipped. - Greptile is 5/5 on the final commit with no unresolved review threads. Its checkout-race finding was reproduced with a failing test, fixed, and rechecked. ## Risks - Additional repositories need disk space and clone time. Access failure for an attached repository stops preparation. - Additional checkouts live under `.paperclip-repositories/` and keep independent histories. Changes stay in those task copies; they do not overwrite configured source folders. - Detached or reconfigured repository copies are retained under `.paperclip-runtime/detached-repositories/`. Sandbox recovery retains per-repository merge baselines. - No database migration, UI contract change, or new credential delegation is required. Referenced projects retain their separate read-only behavior. ## Model Used OpenAI Codex, based on GPT-6, with repository inspection, code execution, and tool use. The runtime does not expose a more specific model deployment ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e9b24c821 |
fix: resume subscription comment wakes and preserve retry status (#13433)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Concurrent task runs can share a managed AI subscription with one credential lease. > - #13438 added durable retries for busy subscriptions and task-lock checks. > - A comment wake can run for an agent who is not the task assignee. Such a run never owns the task lock, so the new checks suppress its retry. > - This pull request preserves those comment wakes while keeping the lock checks for assignee runs. > - It also shows the subscription wait in task status and preserves the failure count through the database projection. ## Linked Issues or Issue Description Refs #13438. **What happened?** When a subscription is busy, a non-assignee comment wake is cancelled without a successor. Repeated subscription waits also display an inflated attempt count because the execution query omits the preserved failure count. **Expected behavior** An eligible comment wake waits and resumes without claiming the assignee's task lock. The task shows “Waiting for AI subscription”. Waiting does not consume provider-failure retries. Assignee retries still stop when the task lock is cleared or transferred. **Steps to reproduce** 1. Configure an agent with a managed subscription and concurrent runs. 2. Hold its credential lease in one run. 3. Mention the agent on a task assigned to another actor. 4. Release the lease and inspect whether the comment wake has a scheduled retry. The test uses a real embedded Postgres database and a real credential lease. Provider execution uses a fixture. The reassignment test applies a database mutation after the real checkout. No live provider account is needed. **Paperclip version or commit** Based on `f912ecaac` from #13438. Before the reconciliation fix, the rebased regression at `2063cfd13` failed because the comment wake had no scheduled retry. The original configuration failure was reproduced on `5282cabde` before #13438 merged. **Deployment mode** Built from source with embedded Postgres for local verification. ## What Changed - Record non-assignee comment-wake authority at admission while holding the task and run locks. A later reassignment cannot grant this exception. Preserve it through repeated subscription waits. - Keep master’s lock checks for assignee retries, including the check inside the scheduling transaction. - Show the subscription wait and preserve its failure count in the execution projection query. - Limit pre-provider wait receipts to fresh executions. A persisted native execution input must retain its recovery path. - Add lease, retry, projection, cancellation, reassignment, pause, revocation, service-recreation, and ownership regression coverage. - Use master’s 60–120 second retry interval and document the resulting behavior. ## Verification - Red: after rebase, the comment-wake test failed with no scheduled retry; the other 50 contention and retry-scheduling tests passed. - Red/green: reassigning the task immediately after real checkout reproduced an unwanted successor for both assignment and comment wakes. Both tests pass after recording authority at admission. - Green: all 285 targeted tests pass across 12 suites, including #13438’s four cancellation-race phases, its hiring/connection suite, run-dispatch integration tests, and both reassignment regressions. - `pnpm -r typecheck` passes after rebase. - `pnpm build` passes on the final revision. - The full general and serialized test suites pass in [CI run 34899514681](https://github.com/paperclipai/paperclip/actions/runs/34899514681) on `99f1e4c1d`. Full-suite verification ran in CI; local verification used the 285 targeted tests. - All 32 active PR checks pass, including build, typecheck, native runner verification, canary dry run, and all browser shards. Two Storybook checks are skipped by path filters. - Greptile reviewed `99f1e4c1d` at 5/5 with no outstanding findings. All review threads are resolved. - `git diff --check` and `node scripts/check-module-boundaries.mjs` pass. ## Risks - The non-assignee exception requires authority recorded under admission locks or its server-created subscription retry. Tests cover caller-supplied flags, assignment and comment reassignment races, repeated comment waits, and assignee lock protections. - A wait uses master’s 60–120 second interval. A long-held lease can produce multiple wait records. - Existing native sessions can have provider effects. They must not receive a fresh-execution receipt that permits replay. - No schema migration or credential permission change is required. ## Model Used OpenAI Codex, model `gpt-6-astra`, with repository search, code editing, shell execution, and test tools. The session does not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f912ecaacf |
fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents hire other agents and assign tasks to them. > - Managed AI connections must follow those hires across legacy and native runners. > - Missing accounts should pause task execution and let the user connect from the task. > - Subscription contention must wait without asking for new credentials. > - This pull request fixes these paths and the native tool and Daytona staging failures found during live tests. > - The result is a working hire, subtask, and connection setup flow on local and remote runners. ## Linked Issues or Issue Description **What happened?** A managed Claude or Codex agent could hire a teammate without a usable AI binding. Cross-provider hiring could fail before the user had a chance to connect the new provider. First-time task setup did not show the existing AI credential form inline. A busy subscription could request a new connection. Native API replies could stop the parent after a hire had already committed. Fresh Daytona sandboxes could fail to extract read-only skill directories created on macOS. **Expected behavior** Compatible hires inherit the managed connection choice. A hire for another provider uses the responsible user's default. If that account is missing, the hire succeeds and the task asks for a connection. Completing setup in the task resumes work automatically. Explicit child auth settings and existing unmanaged login paths keep precedence. Shared-account access checks remain in force. **Steps to reproduce** 1. Connect a Claude or Codex parent with a managed AI account. 2. Ask it to hire one agent of each provider and create a self-assigned subtask. 3. Assign work to both hires without connecting the second provider first. 4. Connect the missing provider from its task card. 5. Check that all tasks finish and same-provider work uses the original account. 6. Repeat with native runners and fresh Daytona sandboxes. The opt-in browser suite in `tests/hiring-ai-connections/README.md` performs these steps. **Paperclip version or commit** The live failures were reproduced from `f2c5e54dc`. The branch is rebased onto `5282cabde`. **Deployment mode** Isolated local development instance. Legacy CLI and native runners. Local execution and ephemeral Daytona sandboxes. Related work: Refs #13247 for managed AI connections. Refs #13268 for legacy credential-reference inheritance, which this branch preserves. Refs #13432 for a concurrent managed-inheritance fix. This PR also covers cross-provider task setup, subscription waits, native API replies, and Daytona extraction. It permits missing responsible-user defaults at hire time; restricted shared selections still fail. ## What Changed - Apply managed connection defaults to both agent creation routes. Preserve explicit auth choices and legacy credential-reference inheritance. - Allow hires before their responsible user connects the provider. Keep approval gates, company boundaries, and shared-account access checks. - Reuse the production AI credential form inside the pending task card. Resume the task after setup. - Retry subscription lease contention without consuming the provider-failure allowance or creating a connection request. - Require task execution-lock ownership when scheduling, promoting, and dispatching subscription retries. Recheck ownership under the issue row lock. - Rename the HTTP operation identity at the native tool boundary so it cannot override the runner's operation identity. - Delay directory permission restoration during Daytona extraction. Preserve the final read-only modes. - Add database-backed regressions, real browser acceptance tests, and Storybook states. Document setup and run-log behavior. ## Verification - Six real browser scenarios passed: both parent providers on legacy local and legacy Daytona; native Codex locally; native Claude on Daytona. Each scenario hires both providers, completes a self-subtask and assigned work, and connects the missing provider inline with automatic continuation. - Successful runs verify the account, responsible user, runner mode, and Daytona lease. All 18 test sandboxes were deleted. - Live authentication used API keys. Subscription inheritance, lease contention, and retry have integration coverage. Fresh subscription OAuth sign-in was not automated. - Red/green tests reproduced missing bindings, missing inline forms, subscription contention, native API reply failure, and GNU tar permission failure. - Seven Storybook browser checks passed. They cover both providers, method selection, narrow layout, completion, cancellation, and invalid credentials. - Full local suite coverage completed before rebase. Initial timing and fixture startup failures passed unchanged on isolated reruns. The first full command did not exit cleanly; the remaining workspace and serialized groups were completed separately. - After rebase, 107 hiring/auth/retry tests and 59 native API, task-card, and Daytona tests passed. The full workspace typecheck, production build, and token gates passed again. Storybook build passed before rebase. - Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and four explicit cancellation-race cases passed. Eight cross-provider cases cover stale auth keys on both creation routes and both runner types. Server typecheck and build passed. - Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks passed. Two optional Storybook jobs were skipped. The full server, workspace, browser, native runner, build, typecheck, and release checks passed. - Three unchanged tests initially failed on a busy port, a chat row-lock race, and preview-server readiness. Each affected job passed after one CI rerun. Isolated local checks also passed: 41 credential tests, the chat-concurrency case, and 25 preview-runtime tests. - Greptile reviewed the final commit at 5/5. Both review threads are resolved. GitHub reports no merge conflicts. ## Risks - A missing personal account now defers authentication to the first task. Explicit incompatible bindings and restricted shared accounts still fail at hire time. - An inherited personal default uses the responsible user's existing authorization to install access for the new agent. It never copies credentials or another user's identity. - Subscription contention retries after a delay and rechecks task eligibility. It does not consume the provider-failure budget. - Native hiring uses the existing managed API-tools opt-in. Remote native runners require a matching Linux binary and provider pack, as documented in the acceptance README. - No schema changes. Live tests make paid provider calls and remain opt-in. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository inspection, code execution, browser automation, and API tools. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |