mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
5c69b69ef2aeab8d8a367b25a8d894cc5308befa
1965
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5c69b69ef2 |
Integrate Cursor child counter provenance
Co-Authored-By: Paperclip <noreply@paperclip.ing> * commit '986f8cc5344107cc199c69e0581cba5401c54282': Verify Cursor child usage lifecycle and fence profile v9 # Conflicts: # doc/architecture/runner-cursor-capabilities.md |
||
|
|
986f8cc534 |
Verify Cursor child usage lifecycle and fence profile v9
Mark missing and invalid child run ordinals as partial observations. Exercise the pinned vendor child construction and reuse methods on all three platforms and regenerate the Cursor-only execution identity. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3518902733 |
Integrate bounded warm retirement recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing> * commit '84a50124a997f4256932eab15fe79dd9ce875940': fix(runner): bound instruction collection and preserve pending retirement |
||
|
|
84a50124a9 |
fix(runner): bound instruction collection and preserve pending retirement
Stop immediate collection retries when ownership is deferred or attempts stop advancing. Keep unconfirmed agent-file retirement recoverable through generic cleanup, with later-owner and restart-proof regressions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6e0f54f565 |
Include the checked Cursor profile fixture in the combined candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing> * commit '4d6e2c00b88af9902facb20ad9177fb70a3a26e0': Keep the next Cursor profile fixture within the typed contract |
||
|
|
4d6e2c00b8 |
Keep the next Cursor profile fixture within the typed contract
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b09e0a1967 |
Combine warm recovery and versioned ACP usage candidates
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e95e28d9c8 |
Preserve partial Cursor native usage diagnostics and version ACP profiles
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3be4f778c6 |
fix(runner): fence warm directory successor ownership and recovery
Retain stable projectless scope, transfer collection to the exact successor lease, and preserve stopped remote files without restarting their sandbox. Validate complete remote directories and cover the composed DB/session lifecycle, retirement, and fresh preparation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
5d4b3ac6c4 |
Finish stop-proved recovery of retained agent directories
Mark lost remote unchanged-turn copies unavailable only after exact termination proof, then perform owned cleanup. Preserve the live unchanged-turn guard and document warm ownership and crash preservation boundaries. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
6c6c05dccb |
Retain unchanged agent files with their warm native owner
Probe complete materializations in bounded children, atomically transfer current-run collection authority, and preserve the exact agent home across warm turns. Retire before collecting changed or uncertain copies, fence stale claims and cleanup, and retain unresolved retirement ownership. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c58a8881f2 |
fix: validate Cursor plan waits against canonical contract policy
Reuse the production completion-contract envelope hash and exercise real contract creation and reuse in accepted-plan persistence tests. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
a2984a0a9d |
Preserve historical Cursor plan proof query bounds
Retain the original Cursor6 event filter when verifying an exact committed wait. Progress rows remain part of Cursor7 lifecycle proof without consuming the historical query budget. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
770dc88a11 |
Bind Cursor plan waits to the native tool lifecycle
Carry the admitted native plan parent tool into canonical request identity and require its exact successful durable lifecycle before passive settlement. Preserve historical committed Cursor6 waits while versioning new admission to Cursor7. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4b94270a0b |
Await clean Stop continuation settlement in recovery test
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ade7c83dbb |
Preserve committed Cursor plan waits across profile upgrades
Keep current profile qualification mandatory when first creating a wait. Recovery uses its scoped committed receipt to verify the entire original proof, so future catalog/settings changes cannot authorize task work. Cover actual persisted finalization, catalog drift, original-profile tampering, and document explicit user continuation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
72a88e124d |
Keep accepted Cursor plans waiting for explicit continuation
Bind normal native plan acceptance to its durable request, delivered answer, admitted run and completion contract. Commit a passive in-progress result with a visible next-message summary, preserve semantic finish priority, and suppress recovery until task-specific continuation without changing modes. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f401b831f2 |
Stabilize route setup and attachment announcement handling
Load route modules during fixture setup so cold transforms cannot leave a timed-out request running against the next test mock state. Dismiss delayed product announcements through the normal Board UI in the attachment receipt fixture. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
41503eb38f |
fix: avoid fresh attachment staging paths for empty wakes
Keep authorized residue scrubbing and unsafe-path checks while leaving untouched workspaces unchanged when no wake attachments were selected. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4597967f89 |
Handle durable cancellation before workspace failure recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
951b06d08a |
fix(runner): fence warm ACPX owners on runtime contract changes
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e154e85394 |
test: scope Telegram lease fault to its fixture company
Fence the credential-reclamation hook used by the global recovery sweep so it cannot leave synthetic lease owners on earlier companies. Exercise a paused foreign receipt after the fault is armed and verify it settles without a leaked lease. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ceb28bdba4 |
fix(server): bind Cursor mode at recovery identity boundaries
Reject missing, foreign, or changed observed modes before accepting suspended checkpoints or interaction-driven identity transitions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
17c6da5e32 |
feat(runner): carry explicit Cursor mode through configuration and transport
Preserve candidate models, reject foreign mode settings, and bind Cursor mode into session identity plumbing. Native mode admission and Rust validation follow in coordinated commits. Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4f2709a611 |
test(runner): check App icon parity at the server boundary
Keep both live and catalog schema assertions through public runner exports while removing the standalone suite dependency on App source. Paperclip-Task: rich-acp-production Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
d2734a81cb |
Merge ACP qualification test context and runtime readiness fixes
Paperclip-Task: rich-acp-production * commit 'ed3a0df1c': test: await connection setup and persisted runtime identity |
||
|
|
ed3a0df1c1 |
test: await connection setup and persisted runtime identity
Complete the dialog fixture context and wait for the actual setup render. Observe the post-spawn PID update before checking the starting service identity, preserving provisioning and readiness assertions. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
482471f0c1 |
Merge reviewed shared ACP input and continuation fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing> * codex/runner-acp-inputs: fix(runner): refresh settled ACPX run grants during authenticated attachment fix(runner): align editable draft bounds with Unicode schema lengths fix(ui): preserve canonical drafts in classic question cards Keep the legacy question fixture payload type-safe |
||
|
|
bb56af7fbd |
Keep the legacy question fixture payload type-safe
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c5a4714add |
feat(runner): preserve editable question defaults through explicit response
Carry bounded text-only initialText through ACP normalization, Rust durable validation, shared contracts and the question UI. Saved edits and explicit clears survive reload. Preserve canonical submitted text while retaining legacy custom-answer normalization and existing validation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
012aa52c77 |
feat(runner): preserve editable question defaults through explicit response
Carry bounded text-only initialText through ACP normalization, Rust durable validation, shared contracts and the question UI. Saved edits and explicit clears survive reload. Preserve canonical submitted text while retaining legacy custom-answer normalization and existing validation. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
97217c5313 |
Integrate reviewed rich ACP instruction and startup fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing> * codex/runner-pi-acp: Bind Cursor instructions to native rules and gate every ACP connection fix(copilot): refresh instructions under provider lifetime ownership fix(copilot): deliver native instructions with v4 session identity docs(runner): retain Pi native prompt delivery proof fix(runner): refresh trusted system instructions before ACPX restore fix(runner): refresh runtime context when restoring ACPX sessions test(copilot): record native system instruction delivery evidence fix(runner): refresh trusted system instructions before ACPX restore perf(runner): hash Pi runtime inventory with bounded workers perf(runner): bound native snapshot copy concurrency at 32 perf(runner): bound native snapshot copy concurrency at 32 test(copilot): preserve final question failure and offline recovery proof fix(runner): refresh runtime context when restoring ACPX sessions feat: return completed handoffs to Agent Chat (#14408) fix(adapters): preserve ACP terminal failure diagnostics (#14573) fix(runner): honor Codex effort selected in composer (#14568) fix: grade Codex clarification and refusal outcomes from evidence (#14570) fix(inbox): keep other users’ failed runs out of Mine (#14572) |
||
|
|
ad472a7755 |
Retry bounded company prefix collisions in chat fixtures
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
1d02ab77b2 |
fix(runner): refresh trusted system instructions before ACPX restore
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
83076d7e7c |
feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage. Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3b4b270650 |
fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The shared ACP adapter engine records agent failures for operators. > - ACP providers can report a failure category, title, and detailed cause. > - Our patch kept only the category in the saved error, so an operator could not diagnose a failure when tracing was off. > - This pull request preserves redacted provider diagnostics in the run error, transcript, and structured run result. > - Operators can now inspect the provider message and any supplied request ID or stack trace after the run ends. ## Linked Issues or Issue Description Refs #13889 (the diagnostic gap; this PR does not update the bundled Claude version). Refs #14484 (related model-refusal classification; this PR retains diagnostics for all terminal failure categories). **What happened?** An ACP turn failed with only `ACP agent reported a terminal service failure.` The provider's title and details were available in memory but absent from the saved error and transcript. **Expected behavior** The run retains useful provider diagnostics even when raw tracing is disabled. Credentials remain redacted. A size limit must report truncation instead of silently removing the cause. **Steps to reproduce** 1. Run an ACP agent that returns an error-severity typed session failure. 2. Include an HTTP error, request ID, and stack text in its title and details. 3. Inspect the failed run with tracing disabled. Before this change, only the category survives. ## What Changed - Both pinned ACPX patches pass complete error text to the in-memory callback, so redaction happens before truncation. - The shared engine retains the sanitized category, title, and details in `resultJson.terminalSessionFailure` and includes the text in the run error and error transcript. - Diagnostics redact configured environment values even under arbitrary names, unknown launch-environment values, connection URL passwords, run credentials, and common credential syntax. Known boolean settings remain readable, while credential values are redacted even when embedded in other text. Diagnostics remove control characters and invalid Unicode. - Title and detail limits keep escaped transcript JSON below the server's chunk limit. Truncated fields include an omission count. The safe run-result projection preserves a byte-bounded diagnostic preview when the result exceeds its byte budget, with an explicit pointer to the full adapter-bounded run error and transcript. - The existing UI and CLI display the error. Diagnostics do not become assistant output. Issue continuation summaries and session-compaction prompts receive only the generic category, preventing provider text from becoming handoff instructions. Existing quota classification, warnings, timeout precedence, and control-channel failure precedence remain in place. - Regression tests cover real ACP child processes with both pinned versions in one-shot and persistent modes, credential redaction, request IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and database retrieval of oversized multibyte diagnostics. ## Verification - Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2 intentionally skipped, no pending or failing checks**. Includes typechecking, build, all Vitest shards, Runner checks, browser E2E, and the canary packaging/public-install dry run. - Greptile: **5/5** on this commit. Superagent security scan passes. All review threads are resolved. - Local verification passed: shared ACP engine suite (395 tests); real Claude ACP child-process and diagnostic regressions across both pinned runtimes and both execution modes; run retrieval and model-handoff regressions (59 tests); ACPX patch packaging (16 tests); full typecheck and build. Affected package typechecks and focused tests were rerun after review fixes. - The broad local `pnpm test:run` was stopped after review edits made its cached imports stale. Fresh targeted runs pass, including both affected server suites. Cold-build import failures were also rerun after dependency builds: chat integration (1,063 tests) and tool access (351 tests) pass. The final commit's complete CI matrix is green. ## Risks - Provider diagnostic text is untrusted. This change retains more of it in company-scoped run records. Redaction and size bounds apply before persistence. - Diagnostics are limited to fields the provider supplies. Old runs cannot recover discarded error text. - No schema migration, recovery-policy change, or new Telemetry or OpenTelemetry export. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
da887ea3e9 |
fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path > - Paperclip manages AI agents that work on assigned tasks. > - The task composer lets a person choose an assignee, model, and effort for the next run. > - A Paperclip Runner agent can use Codex as its provider. > - The composer hid Codex effort for that agent because it checked only the older Codex adapter. > - The native Runner input also did not carry an effort choice to Codex. > - This pull request carries the chosen effort from the composer to each Codex turn. > - People can now select a supported effort and get the effort they selected. ## Linked Issues or Issue Description Refs #14322 **What happened?** The composer showed a model but no effort slider when the assignee used Paperclip Runner with the Codex provider. A task-level model override also did not reach the native Runner input. **Expected behavior** The composer shows effort choices for a known Codex model. The next native Codex turn uses the selected model and effort. **Steps to reproduce** 1. Open a task composer. 2. Select an agent that uses Paperclip Runner with the Codex provider. 3. Select a known Codex model such as `gpt-6-astra`. 4. Open the assignee and model picker. The effort slider is missing before this change. ## What Changed - Show known Codex effort levels for Paperclip Runner Codex assignees. - Save the task effort override in the native run input and send it to Codex on each turn. - Apply the task's merged model and effort overrides when the native run starts. - Apply a task model override for OpenCode Runner without changing the agent's provider. - Add Runner effort tests and desktop and mobile Storybook cases. ## Verification - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm build-storybook` passed. - `pnpm check:token-gates` passed. - Focused UI, server, Runner contract, and Codex driver tests passed. - The full CI test matrix, build, typecheck, and canary dry run passed on the latest head. ## Risks - Native Runner inputs add an optional Codex effort field to the current v5 input. Older inputs keep their previous behavior. - A known model rejects an effort that its catalog does not support. Unknown models do not show a slider. > This fixes an existing composer bug. I checked `ROADMAP.md`; it does not describe this bug as planned work. ## Model Used OpenAI Codex, GPT-6. The exact deployment ID and context window are not exposed in this session. The model used reasoning, code execution, and repository tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com> |
||
|
|
7636966452 |
fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Mine inbox shows work that needs the current user. > - Failed-run rows used the latest run for every agent in the company. > - A failure from another user therefore appeared in Mine and its badge. > - Run list responses also omitted the responsible user needed to filter these rows. > - This pull request uses run ownership for personal failure routing. > - Users see their own failures and can still inspect company failures in All. ## Linked Issues or Issue Description **What happened?** An agent run started for one user failed. Its row and failure badge appeared in another user's Mine inbox. **Expected behavior** Mine and its badge include failed runs for the current responsible user. Other users' failures remain available in All and run details. **Steps to reproduce** 1. Use a company with two human users. 2. Create a failed or timed-out run attributed to the first user. 3. Open Mine as the second user. Before this fix, the failed run appears there and increases the badge. **Paperclip version or commit** Reproduced in regression tests on master at `24beb0057`. **Deployment mode** Authenticated deployment with multiple users. Tests also cover the local single-user board. Related prior work: #933 addressed inbox dismissal and badge consistency. No duplicate ownership fix was found. ## What Changed - Return `responsibleUserId` in normal and summary run lists. - Share one ownership rule across both inbox versions and client/server badges. - Select the latest run per agent before applying the ownership filter. This prevents old failures from resurfacing on shared agents. - Keep unattributed historical failures in the local board's Mine view. Hide them from authenticated users with no matching owner. - Keep company health alerts outside the personal badge, consistent with the client. - Document the routing contract and add page, badge, and database regression coverage. ## Verification - Red: the new badge cases failed with three company failures instead of one personal failure; eight Mine page cases failed across both inbox versions. - Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`, `ui/src/pages/Inbox.test.tsx`, `server/src/__tests__/heartbeat-list.test.ts`, and `server/src/__tests__/inbox-dismissals.test.ts`. - `pnpm check:token-gates` passes. - Agent calls on behalf of a user have two additional red-to-green API regressions. - Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also passes after the agent-call fix. - All CI test shards and browser tests pass on `243bfa681`. The duplicate local `pnpm test:run` was stopped after the CI test lanes completed; it did not finish locally. ## Risks - Authenticated users no longer receive unattributed legacy failures in Mine. Those failures remain visible in All. - The server badge no longer counts company health alerts, matching the existing client badge. - No migration, run state, retry behavior, or company access rules change. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, terminal execution, and browser tools. The exact deployment variant and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
c9b93d7e8c |
fix: preserve terminal task owners during release (#14561)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Tasks record an assigned owner and separate checkout and execution locks. > - A completed task must retain its owner after execution ends. > - The release endpoint currently clears that owner when it clears the locks. > - This pull request preserves the assignee of Done and Cancelled tasks during release. > - Unfinished tasks keep the existing relinquishment behavior. > - The benefit is stable task attribution without retaining execution locks. ## Linked Issues or Issue Description **What happened?** An agent completed an assigned task, then called the release endpoint. The task stayed Done, but its assignee became null. The same defect affects Cancelled tasks. It caused the legacy Claude clarification/reuse and multiple-repository handoff E2E assertions to fail. **Expected behavior** Release must clear execution locks on terminal tasks and preserve their assignee, final status, and disposition timestamps. Release of unfinished tasks must still clear the agent assignee. Only In Progress work returns to Todo. **Steps to reproduce** 1. Create an assigned task with checkout and execution locks. 2. Complete or cancel the task. 3. Call `POST /api/issues/:id/release` as the assigned agent. 4. Read the saved task. Before this fix, its assignee is null. **Paperclip version or commit** Reproduced on master commit `d172197117a14b80a1eb2d2835a0e7cce2679656`. **Deployment mode** Local tests against real PostgreSQL through the production issue routes and services. This is a core lifecycle defect, independent of the agent adapter. Refs: #11689, #6899, #7769. These are related open release proposals. This is an independent fix limited to terminal task ownership. It does not include timer scheduling changes. ## What Changed - Preserve the current assignee when releasing Done or Cancelled tasks. - Keep all execution-lock cleanup and existing unfinished-task behavior. - Cover all seven task statuses through the release API and read back saved state. - Check disposition timestamps, activity attribution, and repeated board cleanup. - Update the API contract, agent reference, and CLI help. ## Verification - Red commit `ddaabb754`: the two terminal-owner regressions failed with `assigneeAgentId: null`; 12 other route tests passed. - Green: all 14 route tests pass, plus the existing successor-checkout race test (15 selected tests total). - Command: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-stale-execution-lock-routes.test.ts src/__tests__/issues-service.test.ts -t 'stale issue execution lock routes|does not let stale release clobber a successor checkout lock'`. - The local host has exhausted its SysV semaphore pool. The red/green runs used the existing test-provider hook to start disposable Docker PostgreSQL 17 instances. Routes, services, migrations, and assertions were unchanged. No database tests in the selected set were skipped. The other 134 tests were excluded by the name filter. - Capability contract and inventory drift checks pass. - `pnpm build` and `pnpm -r typecheck` pass. - The local `pnpm test:run` was interrupted after environment failures while the complete sharded CI suite ran in parallel: native PostgreSQL bootstrap fails under the host semaphore limit, and the large Git fixture hits macOS `ENAMETOOLONG`. A focused rerun confirmed these happen before the relevant assertions. The interrupted local run is not counted as a full pass. - Greptile completed on `1caeeb827e9cb658ddb71f16c2421ec20f80634e` with **5/5**, a successful check, and no review threads. - All CI gates pass on the current head: typecheck, build, general and serialized tests, Runner checks, browser E2E, release packaging, and security checks. Server shard 11 passed on one targeted retry; the first attempt had 836 passing tests but an unhandled workspace-runtime startup rejection caused by an existing timing window. All other successful jobs were reused. - No paid provider evaluations were run. ## Risks - A caller that used release to erase ownership from terminal work will now retain that owner. An explicit assignment update or the board force-release option with `clearAssignee=true` can still clear it. - No schema or migration changes. The transaction, company access, assignee/run checks, and activity log remain in place. ## Model Used - OpenAI GPT-6 through Codex, with tool use, code execution, and test debugging. The session does not expose a more specific backend model version or context-window size. - OpenAI `gpt-6-luna` assisted with read-only test discovery and review. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ca196b0a6 |
feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing> |
||
|
|
53aad90b9e |
fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path > - Paperclip coordinates agent work through execution adapters. > - Sandbox ACP sessions send ordered input through a remote file queue. > - A temporary provider 502 currently closes the session during an input upload. > - A lost response can occur after the sandbox has consumed the message, so a blind retry can duplicate input. > - This pull request retries gateway failures with the same sequence and drops consumed sequences at the receiver. > - The session can continue through a brief provider failure without repeating a tool call. ## Linked Issues or Issue Description **What happened?** A sandbox ACP run can fail with `ACP agent disconnected during request (connection_close, exit=null, signal=null)` when a provider input upload returns HTTP 502. The bridge destroys its local socket on the first failure and can discard the diagnostic before the proxy reads it. **Expected behavior** A temporary gateway failure should get a bounded retry. A lost response after successful delivery must not duplicate input or reorder later messages. Permanent failures must still close the session. **Steps to reproduce** 1. Run the real sandbox process bridge with an echo child and a local test runner. 2. Inject a provider 502 before preparation, after a chunk upload, or after final publication and consumption. 3. Send the next input message. Before this change, the connection closes instead of delivering it. Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and `502 sandbox`. Related work: #13287 covers shutdown after bridge loss; #13793 covers large launch envelopes. This change covers ordered input delivery within a running legacy ACP session. ## What Changed - Retry input uploads up to three times for recognized Daytona and Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms delays. - Give each upload separate temporary paths and discard already-consumed input sequences, including late publication from an earlier attempt. Clean failed attempts in the background without removing a published message or another attempt’s files. Cleanup cannot delay retries or shutdown. - Keep later input behind the retry. Stop queued input on permanent failure and flush a fixed diagnostic before closing the socket. Neither failure-diagnostic persistence nor shutdown-warning persistence can block teardown. - Add real-process regression tests for lost responses, late publication, retry exhaustion, immediate permanent failure, and diagnostic redaction. - Give accepted run-log file appends up to three seconds to drain before finalization computes the size, hash, and durable copy. Close the run handle to later appends. This waits only for file writes, independently of later DB progress or live-event persistence. If writes remain stalled, return null size/hash metadata and skip the final durable copy so the run can settle. Late writes cannot restart mirroring. - Preserve legacy comment attribution when final log size is unknown by reading existing entries within the unchanged 2 MB scan limit. Storage errors or a three-second read deadline return the evidence already read instead of failing the comment listing; pagination stops at the deadline. The deadline requests cancellation of the underlying local stream or S3 HEAD, GET, and response stream. A separate response timeout returns partial evidence even when filesystem I/O delays cancellation; late reads cannot append evidence or start another page. Each listing retains its existing batches of eight reads, without a shared admission cap that skips readable logs under contention. - Document the retry and log-finalization boundaries in the development guide. ## Verification - Final commit `347daa564b`: [Linux CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2) passed. Greptile Apex review 13 scored this commit 5/5 with no new findings; all 12 review threads are resolved. - The final CI run initially hit a Cursor test timeout and four Discord credential-lock contention failures. All five cases passed in isolation. The two failed shards and their aggregate gate passed on retry without a code change. Those intermittent failures are not claimed fixed by this PR. - `pnpm --filter @paperclipai/adapter-utils typecheck` passed. - `pnpm exec vitest run packages/adapter-utils/src/execution-target-stdin-race.test.ts packages/adapter-utils/src/execution-target-sandbox.test.ts packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests passed on the final implementation, including 21 new regressions. The original three fault-injection cases failed before the fix. - The regressions cover failed and indefinitely stalled cleanup, Cloudflare gateway responses and retry exhaustion, permanent errors that must not retry, and teardown while failure logging remains indefinitely stalled. Seven Apex regression cases failed before the review fixes. Adapter-utils typecheck and build passed again after the final review change. - `pnpm exec vitest run server/src/services/run-log-store.test.ts server/src/services/run-log-store-cancellation.test.ts`: all 25 tests passed, including four new regressions that failed before the finalization fix. They cover delayed and failed appends, late-write admission, agreement between the local bytes/summary/durable copy, and a stalled append that exhausts the three-second budget. The timeout case verifies unknown metadata, no final upload, and no mirror restart after late completion. New cancellation tests use the real AWS SDK against a local HTTP server. They verify that stalled HEAD, GET, and response-body connections close on abort and that a subsequent read succeeds. Local range and already-aborted read cases also pass. - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t 'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14 targeted tests passed. The null-size reader case, both storage-error cases, the stalled-read case, and the cancellation/concurrent-listing cases failed before their fixes. The new regressions verify that timed-out reads are cancelled, subsequent listings recover, and two concurrent listings both retain their attribution markers. A read that ignores cancellation still returns partial evidence at three seconds and cannot resume pagination when it finishes; this regression failed before the response-timeout fix. - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter @paperclipai/server build` passed after the response-timeout change. - Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR; the affected packages were rechecked after review fixes. - Full local `pnpm test:run` failed in the general-server group: 511 files passed, 40 failed, and 158 were skipped. Failures include embedded PostgreSQL initialization, read-only cache directory renames, a macOS long-path fixture, and a workspace exposure assertion. The PostgreSQL, cache-permission, and long-path failures also reproduce with both changed implementation files restored to baseline commit `24c58e479a`. The exposure suite passes in isolation both on baseline and the fixed branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later local test groups were not reached. - An earlier CI run hit the Telegram retry-timing failure fixed upstream in #14501. The branch includes that master fix. The selected recovery test passed against a fresh, migrated PostgreSQL 16 database. The embedded PostgreSQL runner is unavailable on this Mac; the isolated database was stopped and removed afterward. - No live agent turn was replayed. The tests use local child processes and injected provider failures. ## Risks Retries are restricted to recognized Daytona SDK and Cloudflare bridge gateway-error messages, which survive plugin RPC serialization. Other errors fail immediately. Temporary upload paths are now unique for all command-managed queue writes. Receiver sequence checks prevent duplicate input; retries do not restart an agent turn. Cleanup and failure logging are nonblocking and best effort; session teardown remains the final cleanup boundary. Log finalization now drains accepted local file writes for at most three seconds and ignores later appends on the closed run handle. A timeout leaves final size/hash unknown and skips the final durable upload; an existing partial mirror may remain available, but it is not claimed as a verified final snapshot. It does not wait for later DB progress or live-event persistence. Optional attribution keeps partial evidence when a read fails or times out. Cancellation closes S3 requests and response streams. Local filesystem I/O may finish after the caller deadline, but a late read cannot change the returned evidence or continue pagination. Later listings can retry after storage recovers. There is no schema, authentication, or permission change. Revert this commit to restore the previous behavior. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and local test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally; targeted tests pass and full-suite limitations are documented above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ea371b9684 |
fix: retain project defaults in partial workspace overrides (#14502)
## Thinking Path > - Paperclip manages AI agents and their work. > - Project workspace policies define how isolated worktrees are set up. > - Tasks can override a branch without providing every setup field. > - The resolver currently replaces the entire project strategy with that partial override. > - Losing an explicit setup command can run the repository fallback script and block the task. > - This pull request keeps enabled project defaults when the task uses the same strategy type. ## Linked Issues or Issue Description **What happened?** A project uses `git_worktree` with `provisionCommand: "true"`. A task overrides only `baseRef`. The resolver drops the command. Worktree creation then invokes `scripts/provision-worktree.sh`, which can fail because its required setup is absent. **Expected behavior** A branch override keeps the project's provision, runtime provision, and teardown commands unless the task explicitly overrides them. A different strategy type must not inherit those commands. **Steps to reproduce** Configure the project with an enabled `git_worktree` strategy and `provisionCommand: "true"`. Give the task an isolated workspace with a `git_worktree` strategy and a different `baseRef`. Add a failing repository fallback provisioner. Before this change, worktree creation invokes that script. After this change, it uses the project's explicit command and succeeds. Related: #4968 concerns agent strategy and working-directory fallback. #13903 concerns gated API fields and reusable-workspace updates. #11091 concerns provision hooks on workspace reuse. None fixes partial task overrides discarding project defaults. ## What Changed - Merge a partial task strategy over the enabled project's strategy only when their types match. - Preserve explicit null values when parsing nullable strategy fields, so they can clear project values. - Keep explicit empty-string overrides and agent fallback behavior. - Exclude disabled project strategies and avoid an inherited branch template when a task pins an existing branch. - Add policy regression coverage and a real Git worktree test with a failing fallback script. - Document inheritance, explicit clearing, and no-op provisioning in the development guide. ## Verification - Policy regression: eight failures before the fix; all 41 policy tests pass after it. - Real worktree regression: passes and creates a worktree using the task's base branch without invoking the failing fallback provisioner. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm test:run`: general-server phase completed with 13,873 passed, 86 skipped, and 14 failures in unchanged macOS skills-cache and Git long-path tests. The same failures reproduce on unmodified base code. The command stops at that phase, so no full local pass is claimed. - CI initially failed the existing Telegram subscription recovery test on a 15-second timeout. The separate fix and investigation are in #14501. A serialized job also lost its runner; GitHub reported lost communication, and that job was rerun without source changes. All 52 final-commit checks pass, with two intentional skips. Greptile is 5/5, with no unresolved comments or merge conflicts. The chat shard passed on one unchanged rerun. The timeout cause remains unproven; #14501 adds phase diagnostics for a recurrence. ## Risks Tasks that specify a partial strategy now retain the project's omitted fields, including setup and teardown hooks. This is the intended behavior change. Inheritance requires an enabled project policy and matching strategy types. Explicit task values still win. Null and empty commands restore existing runtime defaults; they do not guarantee that no script runs. Use `"true"` for an explicit no-op provision command. No migration, live configuration change, or task replay is included. ## Model Used OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes` / `Refs` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full-suite status is recorded above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
90119181e7 |
test: control Telegram subscription retry timing (#14501)
## Thinking Path > - Paperclip manages AI agents and their work. > - Telegram delivery recovers subscription changes after a restart. > - The recovery test leaves a failed action on the real one-second retry timer. > - A slow runner can cross that deadline before the test checks that no retry occurred. > - This pull request holds the fixture deadline until the explicit restart transition. > - The test still checks that recovery uses fresh provider options. ## Linked Issues or Issue Description **What happened?** The Telegram subscription recovery test expected one `setWebhook` request but saw two. A 1.5-second delay after the first failed request reproduces the failure. **Expected behavior** The test controls when the failed request becomes eligible for retry. Host speed does not change its result. **Steps to reproduce** Run the test named `retries an unknown subscription mutation after restart` with a 1.5-second delay after the first failed-action assertion. The old fixture retries too early. The updated fixture passes with the same delay. Related: #13952 fixes a separate Telegram fixture cleanup problem. This change addresses retry timing. ## What Changed - Set the stored retry deadline to 2099 before the pre-restart assertions. - Keep the existing explicit epoch deadline after restart and all provider request assertions. - Report the current test phase only when this test fails, to diagnose an observed intermittent CI timeout. - Leave production retry code unchanged. Temporary delay and per-step console tracing are not included. ## Verification - Delayed regression: failed before the change with two requests instead of one; passed after the change. - Focused Telegram durable private draft Stop group: 34 tests passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - Full local chat shard: 355 tests passed twice. - `pnpm test:run`: general-server phase completed with 13,861 passed, 86 skipped, and 14 failures in unchanged macOS skills-cache and Git long-path tests. The same failures reproduce on unmodified base code. The command stops at that phase, so no full local pass is claimed. - CI exposed a separate 15-second timeout. A diagnostic run passed all 355 shard tests, with the affected test completing in under one second. Its cause remains unproven. Normal step logging is removed; a failure-only phase report remains for a recurrence. The final commit also passes the 355-test chat shard, and Greptile rates it 5/5. All 52 final-commit checks pass, with two intentional skips. There are no unresolved review comments or merge conflicts. ## Risks Low risk. This changes only the fixture deadline. It does not disable a test, extend a timeout, or change production retry behavior. Existing assertions still verify the failed action, the pending state, restart recovery, and fresh provider options. The intermittent CI timeout is not claimed fixed; phase diagnostics narrow the next occurrence without changing the timeout. No documentation change is needed for a test fixture correction. ## Model Used OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes` / `Refs` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run focused tests locally and they pass; full-suite status is recorded above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes, or explained why none is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad1f7e98ea |
fix: retain diagnostic reasons for native runner failures (#14481)
Retain bounded reasons for runner identity, harness recovery, and provider-pack read failures. Preserve existing ownership and cleanup proofs and compatibility with receipt-gated chat recovery. Verified with executor, recovery, diagnostic privacy, typecheck, build, and full CI checks. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
4f6cf5b3ff |
fix: prevent run identity locks from blocking audit checks (#14478)
Use NO KEY UPDATE for identity locks so audit foreign-key checks can proceed while identity writers remain serialized. Preserve task-before-run ordering, company scoping, and foreign keys. Verified with PostgreSQL concurrency regressions, focused tests, typecheck, build, and full CI. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
992f720262 |
fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task descriptions, comments, continuation data, skills, and execution rules enter several agent adapters. > - The same source can be rendered by more than one automatic input carrier. > - Failed resumes can also rebuild input from stale or compact context. > - This pull request gives each Paperclip-owned source one delivery owner and preserves the required transport boundaries. > - It adds deterministic adapter, interaction, runner, and browser tests for these boundaries. > - The benefit is more predictable context delivery with explicit evidence for later live qualification. ## Linked Issues or Issue Description Related: #13144 removes a duplicate environment payload and bounds wake lists. Related: #11360 addresses Hermes resume behavior. This pull request preserves compatible active-session formats while repairing context ownership and stale question creation. **What happened?** Task descriptions and comments could enter more than one automatic context block. Native transports could wrap a complete model input in a second task envelope. Some legacy and gateway adapters could omit the owned assignment on ordinary tasks or rebuild a failed resume with stale compact context. A continuation could also request a question after newer human comments had arrived. **Expected behavior** Each task or comment source has one automatic model-facing owner. Distinct comment IDs and repeated wording remain distinct. Fresh fallback attempts rebuild the required full context. A question request is rejected when newer queued human direction makes it stale. Harness access policy remains owned by execution configuration. **Steps to reproduce** 1. Build a task with a description and current comments. 2. Capture the actual adapter or runner input. 3. Compare source ownership and task-envelope nesting. 4. Queue a human comment before a continuation requests a question. 5. Trigger a failed resume and inspect the fresh retry input. 6. Run the focused adapter, interaction, runner, and browser checks. ## What Changed - Add shared prompt-section selection at the provider-attempt boundary. - Deliver owned assignment context through native, legacy CLI, ACP, gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and Hermes paths. - Rebuild full or compact context after resume recovery changes the attempt. Add native and Claude ACP tests of actual recovery requests. - Preserve custom templates, loaded instruction files, execution policies, and older active-session formats. - Record continuation source metadata and reject stale question creation under the issue-row lock. - Add explicit Product E2E context-integrity profiles, prerequisite gates, credential-isolation checks, and report fixtures. - Bypass service-worker forwarding for same-origin Vite development modules. A real Chromium test fails with resource exhaustion before the repair and passes after it. Production asset caching keeps its existing policy. - Add browser diagnostics and service-worker module-loading regressions. - Add an explicit zero-retry eval option. The default retry behavior remains unchanged. Each campaign records its effective policy. - Remove the model-facing working-directory sentence from four prompt builders. Existing workspace, sandbox, permission, and custom-template configuration remains unchanged. - Align the everyday workflow assertion with the current 47-entry catalog. Compared with current upstream master, the branch carries the context-ownership implementation and its tests, the explicit context-integrity catalog and evidence harness, and the focused browser regression checks. ## Verification **Merge assessment:** focused regression evidence supports merge. This is not full completion of the original broad qualification matrix. The maintainer has authorized merge after fresh verification of the master integration. - Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`. All 14 conflicts are resolved. Cancellation checks, workspace finalization, native Grok support, and both sets of tests are retained. - Current-head Greptile: **5/5**, with no blocking findings. The review names this exact commit. All **59 reported checks are terminal: 55 successful, 4 skipped, zero pending or failing**. This includes the full root general and serialized suites, separate runner checks, typecheck, build, canary, browser E2E, Docker, and security checks. The successful legacy security status is included in that total. - After integration: workspace typecheck and full build passed. Separate runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust tests, and 39 preparation checks**. Other passing checks include 621 Product E2E harness units, 376 focused shared/adapter tests, 160 real-database/API tests, 86 Hermes tests, 18 browser-support checks, and Product E2E typechecking. The complete root suite passed in CI. The duplicate local monolithic root run was stopped after that CI result; it is not counted as a completed local pass. - New native recovery coverage retains full assignment, completion contract, and explicit skill selection after safe replacement, for old and prepared input formats. Full native session test file: **136/136 passed**. - New Claude ACP coverage captures actual fresh, resumed, and missing-session fallback requests. It verifies one assignment copy, comment order, identical text under distinct comment IDs, and full fallback context. Full file: **33/33 passed**. Both affected TypeScript checks passed. - Existing deterministic tests cover source revisions, approval and trust boundaries, completion validation, custom templates, compatible sessions, standalone driver wrapping, and maintained adapter transport requests. - Provider-free browser support: **17/17 passed** after the master merge. Service-worker unit tests: **33/33 passed**. The module-overload regression failed before the repair and passed after it in real Chromium. ### Fresh live comparisons The new batch ran exactly four Product E2E attempts. **All four passed on the first attempt; no retries.** Each has six terminal matchers plus the existing browser lifecycle and invariant checks. | Exact case ID | Control | Candidate | |---|---|---| | `core-compatibility.runner-codex.local.plan-revise-accept` | Passed | Passed | | `local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume` | Passed | Passed | The plan case checks a revised canonical plan and revision-bound approval before completion. The question case restarts the server before submitting the answer, then verifies the continuation completes. Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical frozen definitions and provider versions: Codex `0.156.0` with `gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with `claude-sonnet-5`. The September 24 head added master browser recovery and test-only changes. The September 28 head also integrates newer master changes, including cancellation, workspace finalization, and native Grok. These are frozen-source live results, not exact-head live runs. The candidate received one description copy where the control initially received three. The submitted initial plan envelopes were 7,969 versus 19,097 characters. Question envelopes were 7,592 versus 18,919. These are structural measurements, not whole-provider token or dollar savings. ### Earlier evidence and failed attempts - The preceding fresh batch has four effective passing pairs: OpenCode comment continuation and assigned skill, native Codex comment continuation, and native Claude comment continuation. It retains **11 attempts: eight passed and three failed**. - Original failures remain recorded: missing local PostgreSQL library links before task creation; host-sleep cleanup after task/page checks passed; and a Claude **control** session-open rejection before a model turn. Setup was repaired identically on both worktrees. The permitted unchanged infrastructure retries passed. The underlying Claude provider startup error was not retained and remains unknown. - Older R2 retains **17 passes and one failure** across 18 attempts, including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode blank-page failure led to the service-worker repair. R2 is historical evidence: master changed the native fixed prompt and removed duplicate wake environment data afterward. - The September 24 CI run initially failed one unrelated preview readiness test (`ECONNREFUSED` on its local fixture). Its test and production code match master. Isolated local verification passed **28 tests, 3 skipped**. One unchanged CI retry passed the full shard: **831 passed, 1 skipped**, including all **31 preview-exposure tests**. The aggregate CI gate passed afterward. The precise startup cause remains unknown; a port race is a hypothesis, not a proved cause. ### Limits The original wider profile/workflow matrix, repeated trials, and remote Daytona qualification are incomplete. These results support a focused merge recommendation, not statistical equivalence or universal harness qualification. Some usage receipts are missing in both variants, so no token or dollar savings are claimed. The $500 ceiling was preserved using conservative allowances; failed attempts and unknown charges remain in the ledger. Reproduce the focused additions with `pnpm exec vitest run packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm --filter @paperclipai/paperclip-runner exec vitest run src/native-session-runtime.test.ts`. Full checks use `pnpm -r typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner checks. Paid evals require the frozen definitions, profiles, and credentials; do not use `--all` as a substitute for the selected cases. ## Risks - Context placement changes can affect model behavior. Deterministic checks cover the selected paths, but live qualification remains incomplete. - The stale-question guard can reject a request when queued human comments arrived during the run. This is intended. - New stored inputs and model envelopes retain compatibility readers for older active sessions. - Custom templates may intentionally repeat content. - Removing a model-facing working-directory sentence does not change filesystem, command, sandbox, or permission configuration. - The worker bypass applies only to same-origin development module paths. Cache-policy tests preserve private-response handling and production asset caching. Mounted HTTP fixture changes remain test-only. - This PR does not claim measured token savings or statistical equivalence across every harness. ## Model Used OpenAI Codex, exact model gpt-6-astra, with repository tools and code execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The serving context-window size is not exposed in this task. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR using the required issue fields - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused local checks and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect these changes - [x] I have considered and documented risks above - [x] All current-head Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups for the current head - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1a394bd30 |
feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |